Structured light and line laser fused welding seam thickness grading three-dimensional positioning method and system

By integrating structured light with line laser cameras, combined with deep learning and classic image processing technology, high-precision and rapid three-dimensional positioning of the shield machine cutter head weld is achieved, solving the problems of insufficient weld recognition accuracy and environmental interference in existing technologies, and improving welding efficiency and stability.

CN120612367AActive Publication Date: 2025-09-09CHINA RAILWAY 14TH BUREAU GRP LARGE SHIELD ENG CO LTD +2

Patent Information

Application Number
CN202511109163.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-09-09
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

Existing weld recognition methods have problems such as insufficient measurement accuracy, significant environmental influence, and point cloud loss in large-scale, high-dynamic shield machine cutter head welding projects, making it difficult to meet the high precision and high stability requirements of robots.

Method used

A method of fusing structured light and a line laser camera is adopted. The structured light camera is used for coarse recognition to obtain the weld position and project a coded stripe image. The Center Net convolutional neural network is then used to extract the three-dimensional information of the weld position. The line laser camera is then used to scan and the Marr-Hildreth edge detection algorithm is used to identify the three-dimensional coordinates of the weld.

Benefits of technology

It achieves high-precision and rapid positioning of welds in complex environments, can maintain stable detection performance under dynamic thermal deformation and strong light interference, and improves the consistency and automation level of the welding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612367A_ABST
    Figure CN120612367A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of welding seam positioning, and discloses a structured light and line laser fused welding seam thickness grading three-dimensional positioning method and system, and the key point of the technical scheme is that the method comprises the following steps: S1, shooting a workpiece through a structured light camera, and obtaining a workpiece image; s2, based on the workpiece image, through a channel space attention mechanism and a Center Net convolutional neural network, welding seam position coarse identification is carried out, and the welding seam position is determined; s3, according to the identified position of the welding seam, scanning the position of the welding seam through a line laser camera to obtain three-dimensional point cloud data of the welding seam; and S4, based on the three-dimensional point cloud data, through a Marr-Hildreth edge detection algorithm, identifying to obtain three-dimensional coordinates of the welding seam, and rapidly and accurately positioning the three-dimensional data of the welding seam of the workpiece.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of weld positioning, and more specifically, to a weld coarse and fine graded three-dimensional positioning method and system by integrating structured light and line laser. Background Art

[0002] In the shield machine cutterhead welding project, the cutterhead is the core component of the shield equipment. Its weld trajectory often presents large curvature and multi-degree-of-freedom spatial distribution characteristics, and the welding working range can reach tens of square meters, which puts extremely high demands on the robot's positioning and recognition system.

[0003] Traditional single-sensor recognition solutions face significant challenges in such large-scale, highly dynamic operating scenarios: Taking structured light cameras as an example, although they can quickly acquire a 3D point cloud of the weld area through array projection, their measurement accuracy drops sharply as the working distance increases. Especially when the cutterhead surface has complex rust, oil stains, or highly reflective areas, the point cloud loss rate can reach over 10%, resulting in failure in weld feature extraction. Taking the line laser camera as an example, although it has micron-level measurement accuracy, its line-by-line scanning working mode requires the high-frequency movement of the robotic arm to cover the entire cutter head surface. A single full-size scan takes tens of minutes, which is difficult to meet the real-time requirements of the welding process.

[0004] What is more serious is that the shield machine cutterhead will produce millimeter-level deformation displacement due to heat deformation during the welding process. A single sensor system cannot simultaneously meet the dual needs of global deformation compensation and local weld tracking.

[0005] The traditional methods for weld identification are: One method involves directly using a line laser camera to scan and extract weld features. This requires first establishing a spatial relationship model between the laser plane and the camera. This is typically accomplished through calibration to obtain the camera's intrinsic parameters and the laser plane equation. When laser stripes strike the surface, the 3D coordinates corresponding to each pixel in the stripe image captured by the camera can be calculated using triangulation. Key recognition methods include extracting the centerline of the stripe and employing sub-pixel positioning techniques. Once the center coordinates of the stripes are obtained, the 2D pixels are mapped into 3D point cloud data using a combination of calibration parameters and a triangulation model.

[0006] Although this method can improve the resolution of local features, it has problems such as narrow field of view of single imaging, low global reconstruction efficiency, and difficulty in adapting to the rapid positioning of large-size workpieces.

[0007] Another weld feature extraction method uses a structured light camera. First, a structured light vision sensor is used to scan the workpiece to obtain 3D point cloud data. The weld's principal direction is determined by calculating the minimum directed bounding box of the point cloud. Equally spaced cutting planes are generated along this direction to slice the point cloud. A local coordinate system is established to project the 3D sliced ​​point cloud onto a 2D plane. A gridding process is then used to generate a binary image, mitigating the complexity of 3D data processing.

[0008] This method relies solely on surface structured light and has obvious limitations: surface structured light is easily interfered by metal reflections during dynamic scanning, resulting in large point cloud noise. The reconstruction accuracy of complex geometric contour features, such as steep groove edges and small curvature welds, is insufficient, and the reconstruction speed is slow during high-resolution scanning.

[0009] In summary, existing weld seam recognition methods have many limitations in practical applications, such as insufficient accuracy, significant environmental influences, and point cloud loss. These limitations make them unable to meet the high precision and stability requirements of welding robots in industrial settings. Therefore, it is particularly necessary to explore and adopt other more advanced weld seam recognition and positioning methods to improve welding efficiency and accuracy, overcome the shortcomings of existing technologies, and ensure the stable performance and precise operation of robots in complex working conditions. Summary of the Invention

[0010] The purpose of the present invention is to provide a weld coarse and fine graded three-dimensional positioning method and system that integrates structured light and line laser, which can quickly and accurately locate the three-dimensional data of the workpiece weld.

[0011] The above technical objectives of the present invention are achieved through the following technical solutions: a method for three-dimensional positioning of weld seam coarse and fine grading by fusion of structured light and line laser, comprising the following steps: S1. Use a structured light camera to shoot the workpiece and obtain an image of the workpiece; S2, based on the workpiece image, the channel space attention mechanism and Center Net convolutional neural network are used to perform rough recognition of the weld position and determine the weld position; S3. Scan the weld position using a line laser camera based on the identified weld position to obtain three-dimensional point cloud data of the weld; S4. Based on the 3D point cloud data, the 3D coordinates of the weld are identified through edge detection algorithm.

[0012] As a preferred technical solution of the present invention, S1 includes: photographing the workpiece with a structured light camera to obtain an RGB image, projecting coded stripes onto the workpiece, and capturing a deformed stripe image; S2 includes: S21. Calculate the depth map based on the deformed stripe image, align and fuse the depth map with the RGB image, and generate a four-channel RGB-D tensor. S22. Input the four-channel RGB-D tensor into the Center Net convolutional neural network that integrates the channel spatial attention mechanism for processing, and output the three-dimensional information of the weld.

[0013] As a preferred technical solution of the present invention, in S21, the wrapped phase of the deformed fringe image is calculated by a phase shift method, and a depth map is calculated based on the calibration parameters of the camera; S22 includes: S221. Through the feature encoding network of Center Net convolutional neural network, feature extraction is performed on the four-channel RGB-D tensor to generate a multi-scale feature map; S222. Through the channel attention mechanism, channel weight processing is performed on the multi-scale feature map to obtain channel enhanced features; S223. Perform spatial weight processing on the channel enhancement features through the spatial attention mechanism to obtain channel spatial enhancement features; S224. Gaussian kernel diffusion is applied to the channel space enhancement features to obtain the weld center probability heat map; at the same time, weld size, center point offset, and depth correction are regressed in parallel to output the three-dimensional information of the weld.

[0014] As a preferred technical solution of the present invention, S222 includes: S2221. Perform global maximum pooling and average pooling on the multi-size feature maps to obtain a maximum pooling channel vector and an average pooling channel vector; S2222: Input the maximum pooled channel vector and the average pooled channel vector into the shared MLP to generate channel attention weights; S2223. Apply the channel attention weights to the multi-scale feature maps to obtain channel enhanced features.

[0015] S223 includes: S2231. Perform maximum pooling and average pooling along the channel on the channel enhancement features to obtain a maximum pooling spatial feature map and an average pooling spatial feature map; S2232, concatenate the maximum pooling spatial feature map and the average pooling spatial feature map, and generate spatial attention weights through a large convolution kernel; S2233. Apply the spatial attention weight to the channel enhancement feature to obtain the channel space enhancement feature.

[0016] As a preferred technical solution of the present invention, S4 includes: S41, preprocessing the 3D point cloud data, including sequentially performing straight-through filtering, dimension conversion, coordinate mapping, and region segmentation and repair, to obtain a target depth grayscale image and a dimension mapping relationship; S42, using the Marr-Hildreth edge detection algorithm, sequentially performing Gaussian filtering, Laplace operation, and zero-crossing detection on the target depth grayscale image to obtain the two-dimensional coordinates of the weld edge; S43. Obtain the three-dimensional coordinates corresponding to the two-dimensional coordinate data of the weld edge according to the dimensional mapping relationship.

[0017] As a preferred technical solution of the present invention, S41 includes: S411, performing through-filter processing on the three-dimensional point cloud data according to a preset weld buffer expansion amount, and retaining point cloud data of key weld areas; S412, projecting the retained three-dimensional point cloud data onto the laser scanning plane, and quantizing along the depth direction to generate a depth mapping grayscale image; S413, establishing a dimension mapping relationship and constructing a bidirectional lookup table based on the three-dimensional point cloud data and the depth mapping grayscale image; S414, separating the groove area in the depth mapping grayscale image from the parent material background to generate a binary mask image, and then performing adaptive threshold segmentation. After segmentation, performing morphological repair on the binary mask image to obtain a target depth grayscale image.

[0018] A weld seam coarse and fine grading 3D positioning system integrating structured light and line laser, comprising: A structured light camera is used to photograph the workpiece and obtain an image of the workpiece; The coarse recognition module is used to perform coarse recognition of the weld position and determine the weld position through the channel space attention mechanism and the Center Net convolutional neural network; Line laser camera, used to scan the weld position and obtain 3D point cloud data of the weld; The fine recognition module is used to identify the three-dimensional coordinates of the weld using the Marr-Hildreth edge detection algorithm; The time-sharing trigger control module controls the application timing of the structured light camera and line laser camera through hardware synchronization signals.

[0019] As a preferred technical solution of the present invention, the structured light camera and the line laser camera are further respectively equipped with filtering devices, and the filtering devices eliminate the mutual crosstalk between light sources of different wavelength bands through spectral isolation technology.

[0020] In summary, the present invention has the following beneficial effects: Through the deep collaboration of structured light cameras and line laser cameras, the core contradiction of incompatible efficiency and precision in large-scale welding scenarios has been resolved. The structured light camera, with its wide field of view, quickly locks onto the weld area, significantly reducing the scanning range of the line laser and avoiding the redundant and time-consuming full-width scanning of traditional line lasers. The line laser, guided by coarse positioning, focuses on the target area and captures the microscopic morphological features of the weld through high-precision line-by-line scanning. The two achieve seamless data integration through a dynamic triggering mechanism, breaking through the physical limitations of a single sensor in terms of measurement range and accuracy. Spatial registration technology eliminates multi-source data deviations, forming a closed-loop collaborative system of global coarse positioning and local fine detection, providing highly robust spatial perception capabilities for complex curved surface welding.

[0021] An innovative framework integrating deep learning and classical image processing has significantly improved reliability in harsh industrial environments. A neural network architecture based on multimodal data (RGB-D) uses an intelligent feature screening mechanism to enhance the perception of key weld features, effectively suppressing the impact of interference factors such as metal reflections and oil obscuration on recognition accuracy. Combined with an improved edge detection algorithm, it accurately extracts groove geometry parameters even under abnormal conditions such as missing point clouds and deformation and displacement. This layered processing mechanism retains the physical interpretability of traditional algorithms while incorporating the adaptive advantages of deep learning, maintaining stable detection performance in complex scenarios such as dynamic thermal deformation and strong light interference.

[0022] Through the coordinated optimization of hardware architecture and software algorithms, a complete welding positioning solution has been constructed. The rigid connection and shock-absorbing design of the dual sensors effectively suppress mechanical vibration interference, and the time-sharing trigger control module eliminates optical crosstalk, ensuring the temporal and spatial consistency of multi-source data acquisition. Innovative three-dimensional calibration technology enables precise mapping across sensor coordinate systems, supporting rapid changeovers and process switching. A dynamic compensation mechanism enables real-time tracking of welding thermal deformation, and combined with a path planning algorithm, automatically generates the optimal welding trajectory, forming a closed-loop control process of "perception-decision-execution", significantly improving the consistency and automation level of the welding process, and providing reliable technical support for large-scale equipment manufacturing. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a flow chart of the method of the present invention; Figure 2 It is a preliminary flow chart of rough identification of the present invention; Figure 3 Schematic diagram of the application of the channel and spatial attention mechanism of the present invention; Figure 4 This is a specific flow chart of the rough identification of the present invention; Figure 5 It is a flow chart of line laser scanning preprocessing of the present invention; Figure 6It is a detailed identification flow chart of the present invention. DETAILED DESCRIPTION

[0024] The present invention will be further described in detail below with reference to the accompanying drawings.

[0025] like Figure 1 As shown, the present invention provides a weld coarse and fine graded three-dimensional positioning method by combining structured light and line laser, comprising the following steps: Preparation step S0: Select the equipment to be used to ensure that the selected equipment can meet the dual requirements of global perception efficiency and local detection accuracy.

[0026] The structured light camera is selected with strong resistance to ambient light interference and high stripe decoding robustness to ensure that the basic point cloud can still be stably output in industrial scenarios such as metal reflections and smoke; the line laser camera is selected with high-speed scanning and high signal-to-noise ratio characteristics, which can penetrate the interference of welding spatter and accurately capture key features such as the root gap of the groove and the inclination angle of the blunt edge.

[0027] The parameters and specifications of the selected structured light camera are shown in Table 1, and the parameters and specifications of the line laser camera are shown in Table 2: Table 1. Main specifications of structured light cameras ; Table 2. Main specifications of line laser cameras ; In terms of spatial layout, the structured light camera is installed at the end of the robotic arm at an inclined angle to form a panoramic coverage of the workpiece, while the line laser camera needs to choose a vertical or lateral installation posture according to the geometric characteristics of the target area. It is installed through a multi-degree-of-freedom adjustment bracket to ensure that the fields of view of the two form an effective complement in space, avoiding optical occlusion and retaining the necessary overlapping area for data alignment.

[0028] To suppress vibration interference caused by the robotic arm's motion, a lightweight composite mounting base is used, along with a multi-stage vibration damping structure, to mount both the structured light camera and the line laser camera. This reduces the impact of high-frequency vibration on point cloud quality. A time-sharing trigger control module is also designed to coordinate the exposure timing of the two cameras via hardware synchronization signals to avoid cross-interference from active light sources. When using two cameras, differentiated filtering devices are configured for each camera. Spectral isolation technology eliminates crosstalk between light sources of different wavelengths, and a mapping relationship is established in the overlapping field of view, providing a spatial alignment foundation for subsequent multi-source point cloud fusion.

[0029] After the structured light camera and line laser camera are installed, a unified measurement benchmark is established through joint calibration. This process utilizes a three-dimensional calibration target with stable geometric features and a planar target to collect multiple sets of spatial feature points within the field of view of the structured light camera and line laser camera, respectively. A nonlinear optimization algorithm is then used to solve the overdetermined system of equations AX=XB composed of the measurement data. Ultimately, the hand-eye calibration coordinate transformation matrix between the structured light camera and line laser camera is solved to ensure data consistency and accuracy in the global coordinate system.

[0030] S1, such as Figure 2 and 4 As shown, the workpiece is photographed by a structured light camera to obtain an image of the workpiece; Specifically, a structured light camera is used to capture the workpiece and acquire an RGB image. This RGB image can be directly acquired through the API provided by the structured light camera manufacturer. Structured light cameras provide complete RGB image acquisition capabilities through standardized SDK development interfaces, which are available in multiple programming languages.

[0031] The structured light camera also projects a specifically coded fringe optical pattern onto the workpiece and quickly captures the deformed fringe image. This allows for large-scale scanning of the workpiece surface, quickly acquiring global 3D information about the weld area. This allows for rapid screening of potential weld areas despite strong arc light and smoke interference, allowing for precise identification and demarcation of the target range in subsequent S3 and S4 steps. It should be noted that the components of a structured light camera—the camera and the projector—are two separate pieces of hardware. The structured light camera of the present invention employs a highly integrated hardware architecture, with its core consisting of a camera module and a projector module. Specifically, the camera module refers to the CMOS / CCD image sensor assembly integrated within the structured light camera. This assembly includes an optical lens, an infrared filter, and an image acquisition circuit, and is responsible for capturing the deformed stripe image modulated by the workpiece surface. The projector module is a DLP micro-projection system integrated within the structured light camera. It consists of a DMD digital micromirror array, a collimating lens group, and an infrared laser light source, and is specifically designed to project precisely encoded sinusoidal phase-shifted stripes onto the workpiece surface.

[0032] RGB images and deformed stripe images are two completely different types of image data, each acquired through independent physical channels. An RGB image refers to a visible light image containing red, green, and blue color information. The acquisition principle is as follows: the RGB sensor integrated within the structured light camera captures the visible spectrum information in the ambient light through the Bayer color filter array. Each pixel is composed of three sub-pixels (R / G / B), and the output is a 24-bit true color image. This image directly reflects the color, texture, and other appearance characteristics of the workpiece surface. The deformed stripe image, on the other hand, is a grayscale image acquired by the infrared sensor in the structured light camera. The acquisition process is as follows: a built-in projector emits a coded stripe pattern in the near-infrared band. After being modulated by the workpiece surface topography, the camera's infrared sensor captures the stripe deformation and outputs a single-channel 16-bit grayscale image specifically used for phase resolution and 3D reconstruction.

[0033] S2, such as Figure 2 As shown in the figure, based on the workpiece image, the channel space attention mechanism and the Center Net convolutional neural network are used to perform rough recognition of the weld position and determine the weld position. This can effectively distinguish the weld from background noise, reduce the false detection rate, and ensure the reliability of coarse positioning.

[0034] Specifically, S2 includes: S21. Based on the deformed fringe image, calculate the wrapping phase of the deformed fringe image by a phase shift method. The formula is: ,in Represents the number of periodic groups (or phase shift steps) of the sinusoidal phase-shifted fringes used in structured light 3D measurement. Specifically, It indicates the total number of sinusoidal grating patterns with fixed phase difference projected sequentially by the built-in projection module of the structured light camera during a single 3D measurement. is the wrap phase, with a range of , recover the absolute phase through the phase unwrapping algorithm ,According to the calibration parameters, namely the camera and projector external parameters R,t and internal parameter matrices Kcam,Kproj, the depth map D(x,y) is calculated by triangulation, , where B is the baseline distance between the structured light camera and the projector, is the camera focal length, is the fringe wavelength, is the calibration constant.

[0035] Align and fuse the depth map with the RGB image to generate a four-channel RGB-D tensor, which is expressed as: , where H×W is the image resolution, the depth channel D is normalized to the range [0, 1], and R represents a set of real numbers, which is used here as a label for the tensor space.

[0036] S22: Input the four-channel RGB-D tensor into the Center Net convolutional neural network that integrates the channel spatial attention mechanism for processing, and output the three-dimensional information of the weld. S22 includes: S221, through the feature encoding network of the Center Net convolutional neural network, the four-channel RGB-D tensor is extracted to generate a multi-scale feature map; Center Net adopts an encoder-decoder architecture and a Res Net backbone feature encoding network. The input four-channel RGB-D tensor is encoded by Res Net-18 to generate a multi-scale feature map: Each Res Block contains two convolutions (Conv+BN+ReLU) and skip connections, and the output feature map size is gradually downsampled to H / 16×W / 16×512. Through the cascaded residual blocks, the network gradually abstracts high-level semantic information while retaining spatial details, providing multi-level feature support for subsequent heat map prediction and geometric regression.

[0037] S222, such as Figure 3 As shown, through the channel attention mechanism, channel weight processing is performed on the multi-scale feature map to obtain channel enhanced features; S222 includes: S2221. Perform global maximum pooling and average pooling on the multi-size feature maps to obtain a maximum pooling channel vector and an average pooling channel vector; By aggregating spatial information to generate channel weights, important feature channels are highlighted, and the feature map F is subjected to global maximum pooling (Max Pool) and average pooling (Avg Pool) to generate two channel description vectors: , ; S2222. Input the maximum pooled channel vector and the average pooled channel vector into the shared MLP, and generate channel attention weights through nonlinear mapping: , where σ is the Sigmoid activation function, and the MLP contains one hidden layer (number of neurons C / 16) and ReLU activation; S2223. Apply the channel attention weight to the multi-scale feature map to obtain channel enhanced features; After spatial information compression and weight generation, channel feature enhancement begins. The channel weight is multiplied by the original feature channel by channel to obtain the channel enhanced feature F′. ,in This operation strengthens the channel features related to the weld, such as edges and depth mutation areas, and suppresses irrelevant channel noise.

[0038] Channel weights essentially represent the importance of each feature channel in weld detection tasks. For example, in a four-channel RGB-D input, the depth channel (D) and red channel (R) typically receive higher weights due to their inclusion of weld geometry and oxidation color difference information, while the blue channel (B), which is affected by weld spatter, is dynamically suppressed.

[0039] S223. Perform spatial weight processing on the channel enhancement features through the spatial attention mechanism to obtain channel spatial enhancement features; S223 includes: S2231. Perform maximum pooling and average pooling along the channel on the channel enhancement features to obtain a maximum pooling spatial feature map and an average pooling spatial feature map; The spatial attention module focuses on the key spatial positions of the weld area, performs maximum pooling and average pooling on the channel-enhanced features F′ along the channel dimension, and generates two spatial feature maps: , ; S2232, concatenate the maximum pooling spatial feature map and the average pooling spatial feature map, and generate spatial attention weights through a large convolution kernel; After concatenating the two-way features, the spatial attention weight is generated through the convolution layer : ,in is the convolution operation of the 7×7 convolution kernel, [;] represents channel splicing, and σ is the Sigmoid activation function.

[0040] S2233. Apply the spatial attention weight to the channel enhancement feature to obtain the channel space enhancement feature.

[0041] Through the above-mentioned channel information aggregation and spatial weight generation, the spatial weight is multiplied pixel by pixel with the channel enhancement feature to obtain the final enhanced feature , ,in This operation strengthens the channel features related to the weld (such as edges and depth mutation areas) and suppresses irrelevant channel noise.

[0042] After the Center Net convolutional neural network analyzes the approximate weld location, S222 and S223 perform further coarse positioning. The features processed by the residual block use the channel and spatial attention mechanisms (CBAM) to readjust the feature weights in the spatial and channel dimensions. These features are then added and fused with the decoding layer features, allowing the network to fully utilize effective features and improve detection accuracy. Within the channel and spatial attention mechanisms, the CBAM module significantly improves the network's perception of key features in the weld area through a hierarchical feature screening mechanism.

[0043] The channel attention mechanism first performs global information compression on the feature map output by the residual block in the channel dimension, using maximum pooling to capture the most significant feature responses in the weld area, while simultaneously obtaining channel-level statistical features through average pooling. This dual-path pooling strategy effectively avoids the information loss caused by a single pooling operation. Maximum pooling is sensitive to local extremes, such as high-contrast areas like metal reflective spots and groove edges, while average pooling can reflect the overall activation intensity distribution of the channel.

[0044] Building on the channel-enhanced features, the spatial attention mechanism further focuses on the spatial structural characteristics of the weld. This module compresses the multi-channel feature map into two spatial response maps through a dual-channel pooling operation along the channel dimension: the maximum pooled response map highlights local extreme features such as the weld centerline and groove inflection point, while the average pooled response map reflects the overall spatial distribution trend of the weld area. The concatenation of the two response maps constructs a multi-granular feature representation in the spatial dimension. Spatial context information is then integrated using a large-scale convolution kernel (7×7). This wide receptive field design effectively correlates the weld center point with the surrounding groove structure, avoiding false activations caused by local noise. The resulting spatial attention weight map displays a distinct band-like high-response region whose morphology closely matches the actual direction of the weld, verifying the module's ability to accurately capture spatial geometric features.

[0045] Channel attention and spatial attention utilize a cascaded architecture rather than a parallel structure. This design adheres to the "channel-first, spatial-second" feature optimization logic. Feature filtering in the channel dimension provides a high-quality, redundant feature base for spatial attention, while weight adjustment in the spatial dimension further enhances the ability to express detailed information in the target area based on channel optimization. Working together, the channel attention module can be viewed as a "global filter" for feature channels, while the spatial attention module acts as a "local focuser" for spatial locations.

[0046] S224. Apply Gaussian kernel diffusion to the channel space enhancement features to obtain the weld center probability heat map; at the same time, parallel regression of weld size, center point offset, and depth correction is performed to output the weld 3D information. Specifically: Gaussian kernel diffusion: The network output layer predicts the weld center point heat map , modeled as a Gaussian distribution, the center point of the marked weld (x c ,y c ), generate Gaussian kernel diffusion heat map, the expression is: ,in and is the center coordinate after downsampling, is the adaptive standard deviation.

[0047] Then we start to define the loss function and use the improved FocalLoss to solve the category imbalance problem. The loss function is: ; Here, α = 2 and β = 4 are mainly used to suppress the gradient contribution of simple samples, and N is the number of positive samples. The above heat map branch explicitly models the weld center position in the form of probability density, providing a strong supervision signal for the target area for subsequent geometric regression.

[0048] Regression weld size: To accurately describe the weld shape, the network regresses the weld size (w, h) and center point offset (Δx, Δy) in parallel. First, the size regression is defined, and the L1 loss is used to output the weld width and height: N is the total number of weld feature points processed in the current batch. k is the feature point index variable, and the traversal range is k∈[1, N]. is the actual weld width at the kth measuring point. is the k-th weld width predicted by the network. is the actual weld height at the kth measuring point. is the k-th weld height predicted by the network.

[0049] Center point offset: defines offset compensation to correct the quantization error caused by downsampling. The L1 loss is: N is the total number of weld feature points processed in the current batch. K is the feature point index variable, traversing the range k∈[1, N]. The x-axis offset of the k-th point predicted by the network. is the true x-coordinate of the k-th measurement point. The y-axis offset of the k-th point predicted by the network. is the true y coordinate of the kth measurement point.

[0050] The above-mentioned size regression and offset compensation jointly optimize the positioning accuracy of the weld bounding box and make up for the lack of spatial resolution of single heat map prediction.

[0051] Depth correction: Since the depth measurement of the structured light camera is easily disturbed by metal reflections, the network additionally outputs the depth correction value Δz and uses Huber loss to enhance robustness: N is the total number of weld feature points processed in the current batch. K is the feature point index variable, traversing the range k∈[1, N]. is the true depth offset of the kth sampling point. The depth correction amount for the k-th sample point predicted by the network.

[0052] The deep regression branch corrects system errors through end-to-end learning, incorporates the complementarity of RGB and depth information into network training, and improves the reliability of three-dimensional positioning.

[0053] Based on the above analysis, define the weighted sum of the total loss function and optimize it for training: ; The parameters are set to typical weights: λheat=1, λsize=0.1, λoffset=1, and λdepth=0.5. By dynamically balancing multi-task losses, the network achieves collaborative optimization between heatmap confidence, geometric accuracy, and depth correction, avoiding overfitting of a single task.

[0054] S3. Scan the weld position using a line laser camera based on the identified weld position to obtain three-dimensional point cloud data of the weld; S4, such as Figure 6 As shown in the figure, based on the three-dimensional point cloud data, the three-dimensional coordinates of the weld are identified by the Marr-Hildreth edge detection algorithm.

[0055] After the aforementioned steps of using a structured light camera for rough identification, the first part of the fine identification phase of weld graded positioning and identification involves line laser scanning and data preprocessing. At this point, the robot undergoes coarse positioning and moves to the initial position for fine identification using hand-eye calibration. Guided by the weld seam's prior position provided by coarse identification, the line laser camera initiates high-precision scanning. Its operating principle is based on optical triangulation: a laser projects a linear spot onto the workpiece surface, and the camera calculates the three-dimensional coordinates based on the spot's displacement on the imaging plane. During the scanning process, the robotic arm moves at a constant speed along the weld seam to ensure sufficient point cloud density.

[0056] S4 includes: S41, such as Figure 5 As shown, the preprocessing operation of the 3D point cloud data includes sequentially performing straight-through filtering, dimension conversion, coordinate mapping, and region segmentation and repair to obtain the target depth grayscale image and dimension mapping relationship; S41 includes: S411, performing through-filter processing on the three-dimensional point cloud data according to a preset weld buffer expansion amount, and retaining point cloud data of key weld areas; 1 To eliminate environmental noise interference, first perform through filtering, the formula is: , where x min =X weld −Δ buffer , x max =X weld +Δ buffer , where Δ buffer= 20mm (buffer expansion to account for possible welding thermal deformation). The point cloud volume is reduced by approximately 60% after filtering, retaining key areas such as the groove root and blunt edge. This significantly reduces the amount of data processing.

[0057] S412: Project the retained 3D point cloud data onto the laser scanning plane and quantize it along the depth direction to generate a depth mapping grayscale image. Specifically, project the 3D point cloud onto the laser scanning plane (XY plane) and quantize it along the Z axis to generate a grayscale image, facilitating the use of traditional image processing algorithms. This process requires resolving the conflict between depth resolution (Z-axis quantization accuracy) and noise suppression.

[0058] , where zmed(u, v) is the median of the Z coordinates of all points within the grid (u, v), suppressing the influence of outliers (such as splashes). μz and σz are the Z-axis mean and standard deviation of the filtered point cloud. Depth mapping compresses 3D geometric information into 2D grayscale space, similar to converting a stereo topographic map into a contour map, providing standardized input for subsequent image processing. The median filter and 3σ cutoff in the formula essentially preserve the main groove shape while suppressing random noise. This effectively removes noise and enhances the 3D data.

[0059] S413, establishing a dimension mapping relationship and constructing a bidirectional lookup table based on the three-dimensional point cloud data and the depth mapping grayscale image; Specifically, 3D-2D coordinate index construction is performed to reversely locate the subsequent image processing results (such as edge coordinates) to the three-dimensional space. A bidirectional mapping relationship needs to be established to build a bidirectional lookup table to achieve a cross-dimensional data closed loop. The dimension mapping relationship is: ; Extended base coordinate x base ,y base To reserve a boundary buffer and prevent the boundary point cloud from indexing out of bounds due to rounding errors, the design of the round function and the extended basis coordinates ensure the numerical stability of the mapping process and avoid index misalignment due to floating point errors.

[0060] S414, separating the groove area in the depth mapping grayscale image from the parent material background to generate a binary mask image, and then performing adaptive threshold segmentation. After segmentation, performing morphological repair on the binary mask image to obtain a target depth grayscale image.

[0061] Specifically, in welding scenarios, due to interference from material reflections, oxidation color differences, and other factors, adaptive thresholding and morphological restoration are required. An improved large-scale method is used for threshold segmentation. Adaptive threshold segmentation solves the classification problem of overlapping grayscale areas. Its improved Otsu algorithm strikes a balance between maximizing inter-class variance and ensuring a reasonable target area size: , where T is the optimal segmentation threshold, w1 and w2 are optimization weight coefficients, is the between-class variance, and the parameters are defined as ω1=0.7 and ω2=0.3 to balance the between-class variance. Proportional constraint with foreground. The percentage of foreground pixels is constrained to approximately 30% (empirical value) to avoid over-segmentation. Its physical significance lies in maximizing inter-class differences while limiting the proportion of the groove area to match actual working conditions. After completing threshold segmentation, morphological processing begins. The binary image is first dilated and then eroded to fill small holes and smooth the edges to eliminate isolated spots caused by missing points or noise, ensuring the connectivity of the groove contour. The formula is: , where B3 is a 3×3 cross-shaped core that eliminates small holes, and B5 is a 5×5 rectangular core that smoothes burrs on the edge of the groove.

[0062] S42, using the Marr-Hildreth edge detection algorithm, sequentially performing Gaussian filtering, Laplace operation, and zero-crossing detection on the target depth grayscale image to obtain the two-dimensional coordinates of the weld edge; Gaussian filtering: As the core link of preprocessing, Gaussian filtering's physical significance lies in establishing a balanced relationship between spatial smoothness and edge preservation.

[0063] The Gaussian kernel function essentially reshapes the depth map grayscale image generated by laser scanning in the frequency domain through a spatially weighted averaging strategy. The exponentially decaying weight distribution characteristic causes the contribution of neighboring pixels to the central pixel to decay in a bell-shaped curve. This mechanism can suppress high-frequency noise (such as salt and pepper noise caused by spatter particles) while maintaining the gradient continuity of the groove edge. In particular, for the reflection interference of metal oxide layers commonly seen in welding scenarios, the smoothing effect of the Gaussian filter can effectively bridge the local brightness mutation area, forming a macroscopically coherent groove morphology representation. The formula is: , where σ is the standard deviation of the Gaussian kernel, which controls the smoothing strength. A too large σ will lead to blurred edges, while a too small σ will result in residual noise. The convolution operation is then performed, and the kernel is weighted summed with the local area of ​​the image. The mathematical expression is: , where I(x, y) is the original grayscale image and I′(x, y) is the smoothed image. Gaussian filtering effectively removes high-frequency noise by leveraging the exponential decay properties of the Gaussian kernel. It also ensures a centrally symmetrical distribution of weights, avoiding the edge blurring effect of mean filtering.

[0064] Laplace operation: The design of the Laplace operator reflects the sensitivity of the second-order differential to geometric features. Its discretized convolution kernel is symmetrically arranged with negative weights at the center and positive weights in the neighborhood, which is essentially a quantitative response to the curvature of the image. When acting on a Gaussian filtered image, the operator will produce strong response values ​​with opposite signs on both sides of the groove edge - the upper edge area shows a positive peak due to the sudden drop in depth value, and the root area shows a negative peak due to the sudden increase in depth value. This bidirectional response characteristic enables the algorithm to distinguish the geometric direction of the groove and provide a directional basis for subsequent slope analysis. Especially in the heat-affected zone of welding, the changing trend of the Laplace response amplitude can indirectly reflect the degree of thermal deformation of the material and provide feature-level input for adaptive path planning.

[0065] The Laplace operator is a second-order differential operator that is highly sensitive to sudden grayscale changes (such as the edge of a bevel). Its response strength is positively correlated with the edge curvature. Convolution is used to calculate the second-order derivative of an image and locate edge regions.

[0066] The formula is: , where E(x, y) is the edge intensity map. Positive values ​​indicate grayscale protrusions (such as the top edge of a groove), while negative values ​​indicate depressions (such as the root of a groove). Physical meaning: The larger the absolute value of the Laplace response, the higher the edge intensity.

[0067] Zero-crossing detection: Zero-crossing detection is performed after Gaussian filtering and Laplacian kernel operations. The engineering value of zero-crossing detection lies in converting abstract mathematical responses into actionable physical boundaries. Using a neighborhood sign comparison strategy, the algorithm can capture sub-pixel transition features of groove edges. Essentially, it locates edge crossing points in continuous space by flipping the sign of discrete differential responses. This detection mechanism is uniquely adaptable to edge blurring caused by welding deformation: when thermal deformation makes it difficult for traditional thresholding methods to accurately determine the boundary, the spatial distribution of zero-crossing points maintains a stable topological structure. The gradient threshold constraint introduced during detection achieves an adaptive balance between weak edge preservation and noise suppression by dynamically adjusting the sensitivity threshold. Zero-crossing points are defined as pixel locations where the sign of the edge intensity map E(x, y) changes, representing the extreme value of the second-order grayscale derivative (corresponding to the peak of the first-order derivative, i.e., a true edge). The detection algorithm first performs neighborhood sign comparison: for each pixel (x, y), its eight-neighborhood neighborhood is checked to see if there is a response value with the opposite sign. Next, a gradient threshold constraint is applied: only when the difference between the positive and negative response pairs exceeds the threshold Tedge is it considered a valid edge. The specific mathematical conditions are as follows: , where the parameters are set to , which dynamically adapts to image contrast, introduces non-maximum suppression (NMS), retains local maximum response points, and eliminates repeated edges. The above detection outputs a binary edge map B(x, y)∈{0, 1}, where B(x, y)=1 represents an edge pixel.

[0068] S43. Obtain the three-dimensional coordinates corresponding to the two-dimensional coordinate data of the weld edge according to the dimensional mapping relationship.

[0069] After completing zero-crossing detection to obtain the two-dimensional coordinate information of the weld edge, the system uses the pre-built 3D-2D coordinate index, that is, the dimension mapping relationship, to achieve accurate mapping from image space to three-dimensional physical space.

[0070] In S42 and S43, the Marr-Hildreth edge detection method is used to obtain the 2D coordinate information of the weld seam after the laser scanning and data preprocessing steps described above. Finally, 3D coordinate data is returned through 3D-2D coordinate indexing. The Marr-Hildreth method uses the Laplacian of Gaussian (LoG) operator combined with zero-crossing detection to achieve a balance between noise immunity and edge location accuracy. Its core principle is to capture grayscale abrupt changes (groove edges) using a second-order differential operator and map the 2D image coordinates to 3D space.

[0071] After completing the above steps, the host computer automatically generates a multi-layer, multi-pass welding path based on the characteristic parameters, plans the welding gun's motion trajectory and process parameters, and forms a fully closed-loop control chain of "detection-positioning-planning." This process seamlessly integrates the image processing results into the physical welding space through cross-dimensional data fusion and spatial mapping technology.

[0072] Corresponding to the above method steps, the present invention also provides a weld coarse and fine grading three-dimensional positioning system integrating structured light and line laser, comprising: A structured light camera is used to photograph the workpiece and obtain an image of the workpiece; The coarse recognition module integrates the channel space attention mechanism and the Center Net convolutional neural network to perform coarse recognition of the weld position and determine the weld position; Line laser camera, used to scan the weld position and obtain 3D point cloud data of the weld; The fine recognition module is used to identify the three-dimensional coordinates of the weld using the Marr-Hildreth edge detection algorithm.

[0073] Both structured light cameras and line laser cameras are visual modules, of which the structured light camera is mainly responsible for coarse weld recognition, and the line laser scanning camera is mainly responsible for fine weld recognition.

[0074] The coarse recognition structured light camera uses the Yinniu R-132 structured light camera to collect RGB-D images of the welding site, and uses the Center Net neural network to quickly and accurately find the approximate location of the weld.

[0075] For fine recognition, a DeepVision intelligent SR-7900 line laser camera is used to scan and generate a point cloud image. After coarse recognition determines the approximate weld area, the line laser camera, acting as a high-precision measurement module, performs fine-grained local scanning of the target area. Line laser technology, based on the principle of triangulation, generates submillimeter-level 3D point cloud data through laser line scanning, accurately analyzing weld geometry features such as groove angle, depth, and gap width. Its high resolution enables it to capture subtle variations in weld edges. Combined with an edge detection algorithm, it dynamically optimizes data quality. Depth information is linearly mapped into a grayscale image. Marr-Hildreth edge detection is then used to locate and select the weld edge, ultimately generating weld data. This maintains stable measurement accuracy even under complex interference such as metal spatter and oxide layer coverage. The two work together seamlessly through data fusion technology. The global coordinate information provided by the structured light camera is combined with the local high-precision data from the line laser camera to form a complete 3D model of the welding scene, providing precise input for robot path planning.

[0076] The structured light camera and line laser camera of this invention meet the technical requirements of coarse-fine data processing. The coarse-grained point cloud output by the structured light camera provides a spatial reference for initial weld positioning, while the high-precision point cloud of the line laser camera provides refined reconstruction of groove details. Their collaborative operation establishes a hierarchical data input channel for the multi-level recognition algorithm. Through spatial layout optimization and a coordinated control strategy, the structured light camera and line laser camera achieve a synergistic effect within limited resources, providing reliable hardware support for weld feature extraction under complex working conditions.

[0077] The technical solution of the present invention is particularly suitable for welding scenarios of complex workpieces such as shield machine cutter heads, whose multiple types of grooves (such as X-type, Y-type, and K-type) place higher demands on the adaptability and robustness of the recognition system. Through a phased processing strategy, it not only ensures the efficiency of large-scale scanning, but also achieves high-precision analysis of key areas, significantly improving the level of welding automation. In addition, the system's built-in dynamic anti-interference mechanism can adaptively adjust the light source intensity and data processing algorithm to effectively cope with common challenges in industrial sites such as ambient light fluctuations and smoke and dust obscuration. Ultimately, this coarse-fine collaborative recognition mode not only reduces the need for manual intervention, but also greatly improves the consistency of welding quality and process controllability, providing reliable technical support for the intelligent upgrade of high-end equipment manufacturing.

[0078] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A weld seam coarse and fine graded three-dimensional positioning method using structured light and line laser fusion, characterized by: The steps include: S1. Use a structured light camera to shoot the workpiece and obtain an image of the workpiece; S2, based on the workpiece image, the channel space attention mechanism and Center Net convolutional neural network are used to perform rough recognition of the weld position and determine the weld position; S3. Scan the weld position using a line laser camera based on the identified weld position to obtain three-dimensional point cloud data of the weld; S4. Based on the 3D point cloud data, the 3D coordinates of the weld are identified through edge detection algorithm.

2. The method for coarse and fine grading of weld seams by combining structured light and line laser according to claim 1 is characterized by: S1 includes: The workpiece is photographed by a structured light camera to obtain an RGB image. The coded stripes are projected onto the workpiece and the deformed stripe image is captured. S2 includes: S21. Calculate the depth map based on the deformed stripe image, align and fuse the depth map with the RGB image, and generate a four-channel RGB-D tensor. S22. Input the four-channel RGB-D tensor into the Center Net convolutional neural network that integrates the channel spatial attention mechanism for processing, and output the three-dimensional information of the weld.

3. The method for weld coarse and fine grading three-dimensional positioning by combining structured light and line laser according to claim 2, characterized in that: In S21, the wrapped phase of the deformed fringe image is calculated by a phase shift method, and a depth map is calculated based on the calibration parameters of the camera; S22 includes: S221. Through the feature encoding network of Center Net convolutional neural network, feature extraction is performed on the four-channel RGB-D tensor to generate a multi-scale feature map; S222. Through the channel attention mechanism, channel weight processing is performed on the multi-scale feature map to obtain channel enhanced features; S223. Perform spatial weight processing on the channel enhancement features through the spatial attention mechanism to obtain channel spatial enhancement features; S224. Gaussian kernel diffusion is applied to the channel space enhancement features to obtain the weld center probability heat map; at the same time, weld size, center point offset, and depth correction are regressed in parallel to output the three-dimensional information of the weld.

4. The method for weld coarse and fine grading three-dimensional positioning by combining structured light and line laser according to claim 3 is characterized by: S222 includes: S2221. Perform global maximum pooling and average pooling on the multi-size feature maps to obtain a maximum pooling channel vector and an average pooling channel vector; S2222: Input the maximum pooled channel vector and the average pooled channel vector into the shared MLP to generate channel attention weights; S2223. Apply the channel attention weight to the multi-scale feature map to obtain channel enhanced features; S223 includes: S2231. Perform maximum pooling and average pooling along the channel on the channel enhancement features to obtain a maximum pooling spatial feature map and an average pooling spatial feature map; S2232, concatenate the maximum pooling spatial feature map and the average pooling spatial feature map, and generate spatial attention weights through a large convolution kernel; S2233. Apply the spatial attention weight to the channel enhancement feature to obtain the channel space enhancement feature.

5. The method for three-dimensional positioning of weld seam coarse and fine grading by fusion of structured light and line laser according to claim 1, characterized in that: S4 include: S41, preprocessing the 3D point cloud data, including sequentially performing straight-through filtering, dimension conversion, coordinate mapping, and region segmentation and repair, to obtain a target depth grayscale image and a dimension mapping relationship; S42, using the Marr-Hildreth edge detection algorithm, sequentially performing Gaussian filtering, Laplace operation, and zero-crossing detection on the target depth grayscale image to obtain the two-dimensional coordinates of the weld edge; S43. Obtain the three-dimensional coordinates corresponding to the two-dimensional coordinate data of the weld edge according to the dimensional mapping relationship.

6. The method for three-dimensional positioning of weld seam coarse and fine grading by fusion of structured light and line laser according to claim 5, characterized in that: S4 1 includes: S411, performing through-filter processing on the three-dimensional point cloud data according to a preset weld buffer expansion amount, and retaining point cloud data of key weld areas; S412, projecting the retained three-dimensional point cloud data onto the laser scanning plane, and quantizing along the depth direction to generate a depth mapping grayscale image; S413, establishing a dimension mapping relationship and constructing a bidirectional lookup table based on the three-dimensional point cloud data and the depth mapping grayscale image; S414, separating the groove area in the depth mapping grayscale image from the parent material background to generate a binary mask image, and then performing adaptive threshold segmentation. After segmentation, performing morphological repair on the binary mask image to obtain a target depth grayscale image.

7. A three-dimensional positioning system for weld coarse and fine grading by integrating structured light and line laser, characterized by: include: A structured light camera is used to photograph the workpiece and obtain an image of the workpiece; The coarse recognition module is used to perform coarse recognition of the weld position and determine the weld position through the channel space attention mechanism and the Center Net convolutional neural network; Line laser camera, used to scan the weld position and obtain 3D point cloud data of the weld; The fine recognition module is used to identify the three-dimensional coordinates of the weld using the Marr-Hildreth edge detection algorithm; The time-sharing trigger control module controls the application timing of the structured light camera and line laser camera through hardware synchronization signals.

8. The weld seam coarse and fine grading three-dimensional positioning system integrating structured light and line laser according to claim 7 is characterized by: The structured light camera and the line laser camera are further equipped with filtering devices respectively, and the filtering devices eliminate the mutual crosstalk between light sources of different wavelength bands through spectral isolation technology.

Citation Information

Patent Citations

  • Multi-sensor fusion three-dimensional modeling method and system for building measurement robot

    CN110842940A

  • Welding seam track extraction method based on red line laser

    CN112381783A

  • Dual-sensing integrated type welding seam tracking sensor and deviation rectifying method

    CN113695715A

  • Weld joint quality detection method based on structured light imaging

    CN116772723A

  • Method and device for controlling welding operation of welding gun at tail end of robot

    CN117885096A

Cited By

  • Battery welding seam detection method, device and equipment and readable storage medium

    CN115829923A

  • A method, apparatus, device, and readable storage medium for inspecting battery weld seams.

    CN115829923B

  • Concrete template three-dimensional laser scanning point cloud data processing method, and concrete template detection method and device based on three-dimensional laser scanning

    CN121213665A

  • Graphite equipment weld defect intelligent detection and positioning system based on multispectral imaging

    CN121633123A

  • Graphite equipment weld defect intelligent detection and positioning system based on multispectral imaging

    CN121633123B