A method for detecting and removing spatter on the surface of welded workpieces based on light field deep learning.

CN122156214BActive Publication Date: 2026-09-01HUAZHONG UNIV OF SCI & TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610638354.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-11
Publication Date
2026-09-01
Estimated Expiration
2046-05-11

AI Technical Summary

Technical Problem

[0002]焊接过程中常会产生表面飞溅,针对焊接工件表面的飞溅检测及去除,现有方法以人工视觉检测与手工打磨去除为主,但该方式面临着稳定性及人工成本等问题,而基于传统视觉检测的方法通常也面临识别不准确以及算法鲁棒性不足等问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122156214B_ABST
    Figure CN122156214B_ABST
Patent Text Reader

Abstract

This application relates to a method for detecting and removing spatter on the surface of welded workpieces based on light field deep learning, comprising: acquiring an original lens white image of the surface of the welded workpiece using a light field camera, and correcting the distortion of the original lens white image; extracting sub-apertures from the corrected lens white image to obtain a sub-aperture image array; calculating a light field epiplanet feature map based on the sub-aperture image array; constructing a depth estimation network and inputting the light field epiplanet feature map into the depth estimation network to generate a scene depth map; selecting the optimal sub-aperture image from the sub-aperture image array and performing weighted fusion of the optimal sub-aperture image and the scene depth map to obtain an imaging image; inputting the imaging image into a U-Net network to identify the spatter region in the original lens white image; transforming the pixel coordinate system of the original lens white image to the coordinate system of the end effector and driving the end effector to move to the spatter region to remove the spatter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of spatter detection and removal technology on the surface of welded workpieces, and in particular to a method for spatter detection and removal on the surface of welded workpieces based on optical field deep learning. Background Technology

[0002] Surface spatter is often generated during welding. Current methods for spatter detection and removal on welded workpiece surfaces primarily rely on manual visual inspection and manual grinding. However, this approach faces challenges related to stability and labor costs. Traditional visual inspection methods also typically suffer from inaccurate recognition and insufficient algorithm robustness. Furthermore, traditional visual inspection methods usually rely on 2D grayscale images, providing limited scene information and failing to capture the 3D information of the processed surface. Traditional methods for obtaining 3D scene information often employ structured light-based approaches, which have two main problems: 1) High structural complexity: Structured light methods typically require a combination of a camera and a projector, with specific requirements for the placement and angle of these devices; 2) High image acquisition complexity: Structured light methods usually capture multiple sets of images, and the projection of stripes and multiple camera shots require significant time. Additionally, the stability of the equipment is crucial during image acquisition, resulting in excessively high image acquisition complexity in practical applications. Similarly, at the algorithm level, traditional algorithms also exhibit insufficient robustness when facing different processing scenarios. Summary of the Invention

[0003] Therefore, it is necessary to provide a method for detecting and removing spatter on the surface of welded workpieces based on optical field deep learning, including: S1: Obtain the original lens white image of the surface of the welded workpiece through a light field camera, and perform distortion correction on the original lens white image; S2: Extract sub-apertures from the corrected lens white image to obtain a sub-aperture image array; calculate the light field epiplanet feature map based on the sub-aperture image array; build a depth estimation network based on multi-scale feature extraction and attention enhancement, and input the light field epiplanet feature map into the depth estimation network to generate a scene depth map; S3: Select the optimal sub-aperture image from the sub-aperture image array, and perform weighted fusion of the optimal sub-aperture image with the scene depth map to obtain the imaging image; S4: Input the imaging image into the U-Net network to identify the splatter area in the original lens white image; S5: Transform the pixel coordinate system of the original lens white image to the coordinate system of the end effector, and drive the end effector to the splatter area to remove the splatter.

[0004] Preferably, distortion correction is performed on the original lens white image, including: ; ; in, This represents the coordinates of the microlens image after correction of the original lens white image. This represents the coordinates of the microlens image before correction of the original lens white image. Indicates the first radial distortion coefficient. Indicates the second radial distortion coefficient. Indicates the third radial distortion coefficient. This represents the distance from the original microlens image to the center of distortion. Indicates the first tangential distortion coefficient. This represents the second tangential distortion coefficient.

[0005] Preferably, sub-aperture extraction is performed on the corrected lens white image, including: ; ; in, This represents the extracted sub-aperture image array. This indicates a traversal operation. This represents the total number of rows or columns of microlenses in the microlens array corresponding to the corrected lens white image. Indicates the position as Images of the stitched area under microlenses. This indicates that the microlens is located at the first OK, This indicates that the microlens is located at the first OK, This indicates the location where the spliced ​​region image is stored. Indicates the position as The center coordinates of the microlens This represents the corrected white image of the lens. This represents the image of the stitched region. This indicates the diameter of the microlens.

[0006] Preferably, the polar plane feature map of the optical field is calculated based on the sub-aperture image array, including: The polar plane feature map of the optical field includes a horizontal feature map and a vertical feature map; The formula for calculating the horizontal feature map is: ; in, Represents the horizontal feature map. This indicates an operation that iterates through the data before concatenating it. This indicates the row number of the sub-aperture image in the sub-aperture image array. This represents the four-dimensional representation of a sub-aperture image array. This represents the row pixel information of any sub-aperture image. Indicates the first Sub-aperture image, This indicates the column containing the center sub-aperture image in the sub-aperture image array; The formula for calculating the vertical feature map is: ; in, Represents a vertical feature map. This indicates an operation that iterates through the data before concatenating it. This indicates the column number of the sub-aperture images in the sub-aperture image array. This indicates the row containing the center sub-aperture image in the sub-aperture image array. This represents the column pixel information of any sub-aperture image. Indicates the first Example of aperture image.

[0007] Preferred architectures for depth estimation networks include: Three first convolutional layers, three second convolutional layers, one third convolutional layer, and one fourth convolutional layer are used for multi-scale feature extraction and are connected in sequence. A residual module is passed through each first convolutional layer and each second convolutional layer; An attention module is passed after the third first convolutional layer and the third second convolutional layer.

[0008] Preferably, the workflow of a depth estimation network includes: The polar plane feature map of the light field is passed through three first convolutional layers and three corresponding residual modules in sequence to obtain the first preliminary feature, the second preliminary feature, and the third preliminary feature, respectively. The third preliminary feature is processed through the first attention module to obtain the fourth preliminary feature; The first advanced feature, obtained by element-wise addition of the first preliminary feature and the fourth preliminary feature, is passed through the first second convolutional layer and the corresponding residual module to obtain the second advanced feature; The third advanced feature, obtained by element-wise addition of the second preliminary feature and the second advanced feature, is passed through the second convolutional layer and the corresponding residual module to obtain the fourth advanced feature; The fifth advanced feature, obtained by element-wise addition of the third preliminary feature and the fourth advanced feature, is passed through the third second convolutional layer and the corresponding residual module, as well as the second attention module, to obtain the sixth advanced feature; The second and sixth high-level features are added element-wise to obtain the final features; The final features are passed through the third and fourth convolutional layers in sequence to output the scene depth map.

[0009] Preferably, the kernel size of each first convolutional layer is 5×5, and the number of channels of the first first convolutional layer is 32, while the number of channels of the second and third first convolutional layers is 64. The kernel size of each second convolutional layer is 3×3, and the number of channels in each second convolutional layer is 64. The kernel size of the third convolutional layer is 2×2, and the number of channels in the third convolutional layer is 64. The kernel size of the fourth convolutional layer is 2×2, and the number of channels in the fourth convolutional layer is 1.

[0010] Preferably, selecting the optimal sub-aperture image from the sub-aperture image array includes: Extract three monochrome channels of each sub-aperture image from a multi-view sub-aperture image array; For any sub-aperture image, filter the different gray values ​​corresponding to any pixel in the three monochrome channel images, retain the gray value in the middle of the three gray values ​​as the first gray value of the corresponding pixel, and combine the first gray values ​​of all pixels to obtain the second sub-aperture image. Dark areas, intermediate areas, and overly bright areas are defined as clusters, and the pixels of all second sub-aperture images are classified according to each cluster based on the K-means clustering algorithm; For any pixel, assign a gray value according to the corresponding category to obtain the second gray value of the pixel. Combine the second gray values ​​of all pixels to obtain the optimal sub-aperture image.

[0011] Preferably, for any pixel, a grayscale value is assigned according to the corresponding category, including: If a pixel is classified as a dark region, then the maximum gray value of all pixels belonging to the dark region is used as the second gray value of the corresponding pixel. If a pixel belongs to the middle region, then the pixel's own gray value is used as the second gray value of the corresponding pixel. If a pixel is classified as an overbright region, then the maximum gray value of all pixels belonging to the overbright region is used as the second gray value of the corresponding pixel.

[0012] Preferably, S5 includes: Perform light field camera calibration to obtain camera intrinsic parameters; Transform the pixel coordinate system of the original lens white image to the camera coordinate system based on the camera intrinsic parameters; The camera coordinate system is transformed to the coordinate system of the end flange of the laser processing robot arm in the end effector using the first transformation matrix; The coordinate system of the end flange of the robotic arm is transformed to the coordinate system of the laser processing robotic arm base in the end effector through the second transformation matrix; The coordinate system of the laser processing robot arm base is transformed to the coordinate system of the cleaning robot base in the end effector using the third transformation matrix; The cleaning robot's robotic arm, mounted on its base, moves to the area of ​​the splashes to remove them.

[0013] Beneficial effects: This method outputs an accurate scene depth map by performing distortion correction and sub-aperture extraction on the original lens white image and passing it through a depth estimation network based on multi-scale feature extraction and attention enhancement. This enriches the dimensions of image acquisition information and reduces the spatial complexity of image acquisition. At the same time, it utilizes the excellent image processing capabilities of U-Net to identify the splash area, improving the accuracy and timeliness of image processing. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a flowchart of a method for detecting and removing spatter on the surface of a welded workpiece based on light field deep learning, as described in this application.

[0016] Figure 2 This is a schematic diagram of distortion correction in an embodiment of this application.

[0017] Figure 3 This is a schematic diagram of sub-aperture extraction in an embodiment of this application.

[0018] Figure 4 This is a schematic diagram of the architecture of the depth estimation network in the embodiments of this application.

[0019] Figure 5 This is a schematic diagram showing the positional relationship between the various structures and their coordinate systems in the end effector of this application embodiment. Detailed Implementation

[0020] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the specific embodiments of this application are described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.

[0021] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0022] like Figure 1 As shown, this embodiment provides a method for detecting and removing spatter on the surface of welded workpieces based on optical field deep learning, including: S1: Obtain the original lens white image of the surface of the welded workpiece using a light field camera, and perform distortion correction on the original lens white image.

[0023] In this embodiment, the least squares method is used to perform nonlinear fitting on the line connecting each center point of the microlens to correct the tilt error of the image. A schematic diagram of the correction is shown below. Figure 2 As shown. Specifically, the correction formula includes: ; ; in, This represents the coordinates of the microlens image after correction of the original lens white image. This represents the coordinates of the microlens image before correction of the original lens white image. Indicates the first radial distortion coefficient. Indicates the second radial distortion coefficient. Indicates the third radial distortion coefficient. This represents the distance from the original microlens image to the center of distortion. Indicates the first tangential distortion coefficient. This represents the second tangential distortion coefficient.

[0024] S2: Extract sub-apertures from the corrected lens white image to obtain a sub-aperture image array; calculate the light field epiplanet feature map based on the sub-aperture image array; build a depth estimation network based on multi-scale feature extraction and attention enhancement, and input the light field epiplanet feature map into the depth estimation network to generate a scene depth map.

[0025] Specifically, such as Figure 3 As shown, sub-aperture extraction is performed on the corrected lens white image, including: ; ; in, This represents the extracted sub-aperture image array. This indicates a traversal operation. This represents the total number of rows or columns of microlenses in the microlens array corresponding to the corrected lens white image. Indicates the position as Images of the stitched area under microlenses. This indicates that the microlens is located at the first OK, This indicates that the microlens is located at the first OK, This indicates the location where the spliced ​​region image is stored. Indicates the position as The center coordinates of the microlens This represents the corrected white image of the lens. This represents the image of the stitched region. This indicates the diameter of the microlens.

[0026] Furthermore, the polar plane feature map of the optical field is calculated based on the sub-aperture image array, including: The polar plane feature map of the optical field includes a horizontal feature map and a vertical feature map; The formula for calculating the horizontal feature map is: ; in, Represents the horizontal feature map. This indicates an operation that iterates through the data before concatenating it. This indicates the row number of the sub-aperture image in the sub-aperture image array. This represents the four-dimensional representation of a sub-aperture image array. This represents the row pixel information of any sub-aperture image. Indicates the first Sub-aperture image, This represents the column containing the center sub-aperture image in the sub-aperture image array. The column containing the center sub-aperture image is extracted, and the row pixel information of each row of sub-aperture images in that column is stitched together to obtain the horizontal feature map.

[0027] The formula for calculating the vertical feature map is: ; in, Represents a vertical feature map. This indicates an operation that iterates through the data before concatenating it. This indicates the column number of the sub-aperture images in the sub-aperture image array. This indicates the row containing the center sub-aperture image in the sub-aperture image array. This represents the column pixel information of any sub-aperture image. Indicates the first Sub-aperture images. Extract the row containing the center sub-aperture image from the sub-aperture image array, and stitch together the column pixel information of each row of sub-aperture images in that row to obtain a vertical feature map.

[0028] In this embodiment, as Figure 4 As shown, the architecture of the depth estimation network includes: Three first convolutional layers, three second convolutional layers, one third convolutional layer, and one fourth convolutional layer are used for multi-scale feature extraction and are connected in sequence. A residual module is passed through each first convolutional layer and each second convolutional layer; An attention module is passed after the third first convolutional layer and the third second convolutional layer.

[0029] It is worth further explaining that the workflow of a depth estimation network includes: The polar plane feature map of the light field is passed through three first convolutional layers and three corresponding residual modules in sequence to obtain the first preliminary feature, the second preliminary feature, and the third preliminary feature, respectively. The third preliminary feature is processed through the first attention module to obtain the fourth preliminary feature; The first advanced feature, obtained by element-wise addition of the first preliminary feature and the fourth preliminary feature, is passed through the first second convolutional layer and the corresponding residual module to obtain the second advanced feature; The third advanced feature, obtained by element-wise addition of the second preliminary feature and the second advanced feature, is passed through the second convolutional layer and the corresponding residual module to obtain the fourth advanced feature; The fifth advanced feature, obtained by element-wise addition of the third preliminary feature and the fourth advanced feature, is passed through the third second convolutional layer and the corresponding residual module, as well as the second attention module, to obtain the sixth advanced feature; The second and sixth high-level features are added element by element to obtain the final features; The final features are passed through the third and fourth convolutional layers in sequence to output the scene depth map.

[0030] In this embodiment, as Figure 4 As shown, the kernel size of each first convolutional layer is 5×5, and the number of channels of the first first convolutional layer is 32, while the number of channels of the second and third first convolutional layers is 64. The kernel size of each second convolutional layer is 3×3, and the number of channels in each second convolutional layer is 64. The kernel size of the third convolutional layer is 2×2, and the number of channels in the third convolutional layer is 64. The kernel size of the fourth convolutional layer is 2×2, and the number of channels in the fourth convolutional layer is 1.

[0031] In this embodiment, as Figure 4 As shown, the residual module includes: The sequence consists of the fifth convolutional layer, the batch normalization layer, the LeakyReLU activation function, the fifth convolutional layer, and the batch normalization layer.

[0032] In this example, the attention module works as follows: ; ; in, This represents the features after passing through the attention module. This represents the features input before the attention module. This represents the attention weight coefficient. This represents the Sigmoid function. This represents a multilayer perceptron. This indicates global average pooling.

[0033] This embodiment also provides the training process for the depth estimation network, including: The loss function is calculated based on the real depth map and the scene depth map predicted by the network. The expression of the loss function is as follows: ; ; in, Represents the loss function. Indicates the number of pixels. Indicates the first The logarithmic difference between the true depth value and the predicted depth value at each pixel Indicates the first The true depth value at each pixel. Indicates the first The predicted depth value at each pixel. The scale invariance coefficient is typically set to 0.5. The network parameters of the depth estimation network are optimized using the loss function and gradient descent until the loss function converges, resulting in the trained depth estimation network.

[0034] S3: Select the optimal sub-aperture image from the sub-aperture image array, and perform weighted fusion of the optimal sub-aperture image with the scene depth map to obtain the imaging image.

[0035] Specifically, the optimal sub-aperture image is selected from the sub-aperture image array, including: Extract the three monochrome channels of each sub-aperture image from the multi-view sub-aperture image array. The expression for extracting the grayscale value of the monochrome channel of a pixel is: ; ; ; in, Indicating sub-aperture images pixels at The grayscale value of the red channel, Indicating sub-aperture images pixel Red channel information, Indicating sub-aperture images pixels at The grayscale value of the green channel, Indicating sub-aperture images pixel Green channel information, Indicating sub-aperture images pixels at The grayscale value of the blue channel. Indicating sub-aperture images pixel The blue channel information.

[0036] For any sub-aperture image, filter the different gray values ​​corresponding to any pixel in the three monochrome channels, retaining the middle gray value among the three as the first gray value of the corresponding pixel. Combine the first gray values ​​of all pixels to obtain the second sub-aperture image; the filtering expression is: ; in, Represents pixels in a sub-aperture image , This indicates a centering filter formula used to filter the grayscale value that is centered among three grayscale values, in order to avoid information loss caused by an image that is too bright or too dark.

[0037] Dark areas, intermediate areas, and overly bright areas are defined as clusters, and pixels in all second sub-aperture images are classified according to their respective clusters based on the K-means clustering algorithm. The classification process specifically includes: Step 1: Randomly select 3 samples as the sample centers of the corresponding 3 clusters. Step 2: Calculate the Euclidean distance between each sample and the gray values ​​of each pixel at all sample centers, and assign it to the cluster with the nearest sample center; Step 3: Recalculate the new sample centers in the three clusters. The calculation formula includes: ; in, Indicates the first During the nth iteration New sample centers in each cluster Indicates the first During the nth iteration The total number of samples in each cluster express The first in One sample; Step 4: Repeat steps 2-3 until the Euclidean distance between any two samples in each cluster is less than the preset threshold, and the classification is complete.

[0038] For any pixel, assign a gray value according to the corresponding category to obtain the second gray value of the pixel. Combine the second gray values ​​of all pixels to obtain the optimal sub-aperture image.

[0039] Furthermore, for any given pixel, a grayscale value is assigned according to the corresponding category, including: If a pixel is classified as a dark region, then the maximum gray value of all pixels belonging to the dark region is used as the second gray value of the corresponding pixel. If a pixel belongs to the middle region, then the pixel's own gray value is used as the second gray value of the corresponding pixel. If a pixel is classified as an overbright region, then the maximum gray value of all pixels belonging to the overbright region is used as the second gray value of the corresponding pixel.

[0040] Furthermore, the weighted fusion formula for the optimal sub-aperture image and the scene depth map is expressed as: ; in, Represents the image being captured. Indicates the fusion weight. Represents the optimal sub-aperture image. The scene depth map represents an image with more obvious surface morphology features obtained by combining the optimal sub-aperture image with the scene depth map with the 3D topography information.

[0041] S4: Input the imaging image into the U-Net network to identify the splatter area in the original lens white image.

[0042] Specifically, the U-Net network is a conventional U-Net network, in which: The contraction path consists of 8 convolutional layers with a kernel size of 3x3, divided into 4 layers, each with two convolutional kernels. The number of channels doubles in successive increments of 64, 128, 256, and 512. The bottleneck layer consists of 2 convolutional layers with a kernel size of 3x3 and a number of channels of 1024. The expansion path consists of four upper convolutional layers with a kernel size of 2x2 and eight reconstructed convolutional layers with a kernel size of 3x3, divided into four layers, each with two convolutional kernels. The upper convolutional layers are used to halve the number of channels of the feature map of the previous layer step by step, i.e., decreasing in the sequence of 512, 256, 128, 64. The reconstructed convolutional layers receive the feature stream after concatenating the output of the upper convolutional layers with the features of the corresponding layers of the contraction path, and remap its channel number to the dimensions of 512, 256, 128, 64, which are symmetrical to the contraction path. The final output layer uses a classification and annotation tool to classify and annotate the detected targets in the imaging image through a convolutional layer with a kernel size of 1x1, and obtains a dataset. The dataset is divided into training set, validation set and test set in a ratio of 7:2:1. The obtained training dataset is input into the U-Net network. The feature map first enters the encoding path, and the input feature is defined as X. Each layer contains two 3x3 convolutional operations C and one 3x3 max pooling operation P. In the w-th downsampling level, the feature mapping relationship can be expressed as: ; in, Indicates the first Features after layer downsampling This indicates a max pooling operation. Represents the ReLU activation function. This represents the convolution operation. Indicates the first Features after layer downsampling; The network performs upsampling transformation through path expansion. The upsampling operation is defined as `up`, and the reconstructed features are: ; in, Indicates the first Reconstructed features after layer upsampling Indicates the first Reconstructed features after layer upsampling; Each layer's shrinkage characteristics and corresponding expansion features The fusion can be represented as: ; in, This represents the output features after feature fusion. This indicates the concatenation of channel dimensions. This indicates a jump connection used to enable cross-scale information transfer; In the last layer, a 1x1 convolutional layer maps the multi-channel feature vectors to the corresponding analog numbers, thus completing image segmentation.

[0043] S5: Transform the pixel coordinate system of the original lens white image to the coordinate system of the end effector, and drive the end effector to the splatter area to remove the splatter.

[0044] The positional relationships between the various structures and their coordinate systems in the end effector are as follows: Figure 5 As shown, specifically, step S5 includes: Perform light field camera calibration to obtain camera intrinsic parameters; The pixel coordinate system of the original lens white image is transformed to the camera coordinate system based on the camera intrinsic parameters. The transformation formula is as follows: ; in, This represents the camera coordinates in the camera coordinate system. This represents the actual physical length of a single pixel in the x and y directions. This indicates the intersection of the camera's optical axis and the imaging plane. The coordinates of the microlens image after correction of the original lens white image are represented; the converted pixel coordinates are... ,in The vertical displacement of the welded workpiece surface from the optical center of the light field camera is expressed by the following formula: ; in, The baseline of the light field camera represents the distance between two adjacent viewpoints. The focal length of the light field camera representing the main viewpoint. Indicates parallax.

[0045] The camera coordinate system is transformed to the coordinate system of the end flange of the laser processing robotic arm in the end effector using a first transformation matrix; the first transformation matrix is ​​expressed as: ; in, Denotes the first transformation matrix. The first rotation matrix (3×3) represents the coordinate system between the camera coordinate system and the coordinate system of the end flange of the laser processing robot arm. The first translation matrix (3×1) represents the coordinate system between the camera coordinate system and the coordinate system of the end flange of the laser processing robot arm.

[0046] The coordinate system of the end effector flange of the robotic arm is transformed to the coordinate system of the laser processing robotic arm base in the end effector mechanism using the second transformation matrix; the second transformation matrix is ​​expressed as: ; in, This represents the second transformation matrix. The second rotation matrix (3×3) represents the coordinate system between the end flange of the robotic arm and the coordinate system of the laser processing robotic arm base. The second translation matrix (3×1) represents the coordinate system between the end flange of the robotic arm and the coordinate system of the laser processing robotic arm base.

[0047] The coordinate system of the laser processing robotic arm base is transformed to the coordinate system of the cleaning robot base in the end effector using a third transformation matrix; the third transformation matrix is ​​expressed as: ; in, This represents the third transformation matrix. The third rotation matrix (3×3) represents the coordinate system between the laser processing robotic arm base and the cleaning robot base. The third translation matrix (3×1) represents the coordinate system between the laser processing robotic arm base and the cleaning robot base.

[0048] The cleaning robot's robotic arm, mounted on its base, moves to the area of ​​the splashes to remove them.

[0049] The method for detecting and removing spatter on the surface of welded workpieces based on optical field deep learning provided in this embodiment has the following beneficial effects: This method uses a light field camera to acquire an original lens white image containing multi-dimensional scene information as input. Then, it performs distortion correction and sub-aperture extraction on the original lens white image and passes it through a depth estimation network based on multi-scale feature extraction and attention enhancement to output an accurate scene depth map. This enriches the dimensions of the image acquisition information and reduces the spatial complexity of image acquisition. At the same time, it utilizes the excellent image processing capabilities of U-Net to identify the splash area, improving the accuracy and timeliness of image processing. This method has practical significance and good application prospects.

[0050] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0051] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for detecting and removing spatter on the surface of a welding workpiece based on light field deep learning, characterized in that, include: S1: Obtain the original lens white image of the surface of the welded workpiece through a light field camera, and perform distortion correction on the original lens white image; S2: Extract sub-apertures from the corrected lens white image to obtain a sub-aperture image array; calculate the light field epiplanet feature map based on the sub-aperture image array; build a depth estimation network based on multi-scale feature extraction and attention enhancement, and input the light field epiplanet feature map into the depth estimation network to generate a scene depth map; S3: Extract the three monochrome channels of each sub-aperture image from the multi-view sub-aperture image array; For any sub-aperture image, filter the different gray values ​​corresponding to any pixel in the three monochrome channel images, retain the gray value in the middle of the three gray values ​​as the first gray value of the corresponding pixel, and combine the first gray values ​​of all pixels to obtain the second sub-aperture image. Dark areas, intermediate areas, and overly bright areas are defined as clusters, and the pixels of all second sub-aperture images are classified according to each cluster based on the K-means clustering algorithm; For any pixel, assign a gray value according to the corresponding category to obtain the second gray value of the pixel, and combine the second gray values ​​of all pixels to obtain the optimal sub-aperture image. The optimal sub-aperture image is weighted and fused with the scene depth map to obtain the imaging image; S4: Input the imaging image into the U-Net network to identify the splatter area in the original lens white image; S5: Transform the pixel coordinate system of the original lens white image to the coordinate system of the end effector, and drive the end effector to the splatter area to remove the splatter.

2. The welding workpiece surface spatter detection and removal method based on light field deep learning according to claim 1, characterized in that, Distortion correction is performed on the original lens white image, including: ; ; in, This represents the coordinates of the microlens image after correction of the original lens white image. This represents the coordinates of the microlens image before correction of the original lens white image. Indicates the first radial distortion coefficient. Indicates the second radial distortion coefficient. Indicates the third radial distortion coefficient. This represents the distance from the original microlens image to the center of distortion. Indicates the first tangential distortion coefficient. This represents the second tangential distortion coefficient.

3. The method for detecting and removing spatter on the surface of welded workpieces based on optical field deep learning according to claim 1, characterized in that, The polar plane feature map of the optical field is calculated based on the sub-aperture image array, including: The polar plane feature map of the optical field includes a horizontal feature map and a vertical feature map; The formula for calculating the horizontal feature map is: ; in, Represents the horizontal feature map. This indicates an operation that iterates through the data before concatenating it. This indicates the row number of the sub-aperture image in the sub-aperture image array. This represents the four-dimensional representation of a sub-aperture image array. This represents the row pixel information of any sub-aperture image. Indicates the first Sub-aperture image, This indicates the column containing the center sub-aperture image in the sub-aperture image array; The formula for calculating the vertical feature map is: ; in, Represents a vertical feature map. This indicates an operation that iterates through the data before concatenating it. This indicates the column number of the sub-aperture images in the sub-aperture image array. This indicates the row containing the center sub-aperture image in the sub-aperture image array. This represents the column pixel information of any sub-aperture image. Indicates the first Example of aperture image.

4. The method for detecting and removing spatter on the surface of welded workpieces based on optical field deep learning according to claim 1, characterized in that, The architecture of depth estimation networks includes: Three first convolutional layers, three second convolutional layers, one third convolutional layer, and one fourth convolutional layer are used for multi-scale feature extraction and are connected in sequence. A residual module is passed through each first convolutional layer and each second convolutional layer; An attention module is passed after the third first convolutional layer and the third second convolutional layer.

5. The method for detecting and removing spatter on the surface of welded workpieces based on optical field deep learning according to claim 4, characterized in that, The workflow of a depth estimation network includes: The polar plane feature map of the light field is passed through three first convolutional layers and three corresponding residual modules in sequence to obtain the first preliminary feature, the second preliminary feature, and the third preliminary feature, respectively. The third preliminary feature is processed through the first attention module to obtain the fourth preliminary feature; The first advanced feature, obtained by element-wise addition of the first preliminary feature and the fourth preliminary feature, is passed through the first second convolutional layer and the corresponding residual module to obtain the second advanced feature; The third advanced feature, obtained by element-wise addition of the second preliminary feature and the second advanced feature, is passed through the second convolutional layer and the corresponding residual module to obtain the fourth advanced feature; The fifth advanced feature, obtained by element-wise addition of the third preliminary feature and the fourth advanced feature, is passed through the third second convolutional layer and the corresponding residual module, as well as the second attention module, to obtain the sixth advanced feature; The second and sixth high-level features are added element-wise to obtain the final features; The final features are passed through the third and fourth convolutional layers in sequence to output the scene depth map.

6. The method for detecting and removing spatter on the surface of welded workpieces based on optical field deep learning according to claim 5, characterized in that, The kernel size of each first convolutional layer is 5×5, and the number of channels in the first first convolutional layer is 32, while the number of channels in the second and third first convolutional layers is 64. The kernel size of each second convolutional layer is 3×3, and the number of channels in each second convolutional layer is 64. The kernel size of the third convolutional layer is 2×2, and the number of channels in the third convolutional layer is 64. The kernel size of the fourth convolutional layer is 2×2, and the number of channels in the fourth convolutional layer is 1.

7. The method for detecting and removing spatter on the surface of welded workpieces based on optical field deep learning according to claim 1, characterized in that, For any pixel, assign a grayscale value according to the corresponding category, including: If a pixel is classified as a dark region, then the maximum gray value of all pixels belonging to the dark region is used as the second gray value of the corresponding pixel. If a pixel belongs to the middle region, then the pixel's own gray value is used as the second gray value of the corresponding pixel. If a pixel is classified as an overbright region, then the maximum gray value of all pixels belonging to the overbright region is used as the second gray value of the corresponding pixel.

8. The method for detecting and removing spatter on the surface of welded workpieces based on optical field deep learning according to claim 1, characterized in that, S5 include: Perform light field camera calibration to obtain camera intrinsic parameters; Transform the pixel coordinate system of the original lens white image to the camera coordinate system based on the camera intrinsic parameters; The camera coordinate system is transformed to the coordinate system of the end flange of the laser processing robot arm in the end effector using the first transformation matrix; The coordinate system of the end flange of the robotic arm is transformed to the coordinate system of the laser processing robotic arm base in the end effector through the second transformation matrix; The coordinate system of the laser processing robot arm base is transformed to the coordinate system of the cleaning robot base in the end effector using the third transformation matrix; The cleaning robot's robotic arm, mounted on its base, moves to the area of ​​the splashes to remove them.

Citation Information

Patent Citations

  • Accurate high-reflection removing method based on light field iteration

    CN112419185A

  • PCB detection system and method

    CN113483655A