Underwater concrete surface three-dimensional reconstruction method based on laser stripe and stereo vision

CN122550830APending Publication Date: 2026-08-11SOUTHWEAT UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]针对现有技术中的上述不足,本发明提供的基于激光条纹与立体视觉的水下混凝土表面三维重建方法解决了水下光学测量中多介质折射、背景散射干扰和弱纹理表面匹配失效的问题

Benefits of technology

(1)本发明提供了基于激光条纹与立体视觉的水下混凝土表面三维重建方法,针对水下光学测量中受多介质非线性折射、散射引起的背景干扰以及表面纹理不足等因素影响,传统光学重建方法难以实现稳定可靠的几何测量,本发明采取以下技术改进:首先,构建了一种跨介质标定模型,将空气中标定参数转换为水下成像参数,从而避免繁琐的现场水下标定。其次,提出了结合Hessian优化的CRSA U-Net方法,用于强噪声、低对比度条件下激光条纹的稳健分割与亚像素级激光条纹中心提取,最后,构建了联合激光条纹特征与立体视觉联合约束的三维点云重建方法,并结合旋转扫描位姿变换实现多帧局部点云拼接与完整区域重建,解决了水下光学测量中多介质折射、背景散射干扰和弱纹理表面匹配失效的关键问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122550830A_ABST
    Figure CN122550830A_ABST
Patent Text Reader

Abstract

This invention discloses a method for 3D reconstruction of underwater concrete surfaces based on laser stripes and stereo vision, belonging to the field of underwater concrete surface 3D reconstruction technology. The method includes: calculating equivalent underwater camera parameters based on camera and laser plane parameters acquired in air, combined with a cross-medium calibration model; preprocessing the original underwater image using a contrast-perception adjustment algorithm after field-of-view transformation to obtain a preprocessed image; extracting laser stripes from the laser stripe image using a Hessian-optimized CRSA U-Net laser stripe extraction method to obtain laser stripe features; and jointly constraining 3D reconstruction using laser stripe features and stereo vision to generate a single-frame local point cloud, and stitching together multiple frames of local point clouds to obtain a complete point cloud. Experimental results show that the proposed method achieves optimal measurement accuracy, with the overall error drastically reduced to 3.65% and the point cloud integrity reaching a near-perfect 98.6%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of underwater concrete surface three-dimensional reconstruction technology, specifically involving a method for underwater concrete surface three-dimensional reconstruction based on laser stripes and stereo vision. Background Technology

[0002] Underwater concrete structures are widely used in major infrastructure projects such as dams, bridges, and ports. Detailed geometric information on surface defects such as cracks, scour pits, and localized erosion is crucial for structural condition assessment and quantitative defect analysis. Traditional underwater inspection relies primarily on manual underwater visual inspection and sonar imaging. Manual inspection is inefficient and highly subjective, while sonar, although capable of wide-range detection in turbid water, often lacks the spatial resolution required for close-range, detailed measurements, particularly in accurately characterizing complex geometric features like small-scale cracks and localized scour. Therefore, high-precision 3D reconstruction methods based on optical sensing offer a new technical approach for acquiring detailed geometric information of underwater concrete surfaces. However, underwater optical measurement is simultaneously affected by factors such as multi-medium nonlinear refraction, scattering-induced background interference, and insufficient surface texture, posing significant challenges to traditional optical 3D reconstruction methods in terms of geometric accuracy, point cloud integrity, and measurement stability.

[0003] With the development of underwater visual sensing technology, image acquisition and intelligent assessment using underwater robots equipped with high-definition cameras have become an important research direction for underwater concrete structure inspection. For example, existing studies have achieved two-dimensional automatic identification of dam erosion and cracks based on deep networks. However, pixel-level analysis relying solely on two-dimensional images lacks spatial depth information, making it difficult to achieve high-fidelity geometric reconstruction and accurate three-dimensional quantization of structural defects. To meet the need for fine three-dimensional information acquisition in complex underwater environments, the research team has conducted continuous exploration around underwater visual sensing and point cloud data acquisition, and has initially constructed a prototype for three-dimensional measurement of underwater defects based on stereo vision. However, engineering scenario verification shows that early methods still have significant limitations in terms of binocular matching stability, laser stripe anti-scattering interference capability, and cross-media calibration convenience. When relying solely on stereo vision, weakly textured concrete surfaces are prone to feature matching failure and point cloud voids; traditional single-line structured light is difficult to suppress strong background scattering noise in turbid water, leading to a decrease in geometric feature extraction accuracy; and the cumbersome on-site underwater calibration required to compensate for multi-media refractive distortion also restricts the efficiency of system deployment to some extent. Summary of the Invention

[0004] To address the aforementioned shortcomings in existing technologies, the underwater concrete surface three-dimensional reconstruction method based on laser stripes and stereo vision provided by this invention solves the problems of multi-medium refraction, background scattering interference, and weak texture surface matching failure in underwater optical measurement.

[0005] To achieve the aforementioned objectives, the present invention employs the following technical solution: a three-dimensional reconstruction method for underwater concrete surfaces based on laser stripes and stereoscopic vision, comprising the following steps: S1. Based on the camera and laser plane parameters obtained in the air, and combined with the cross-medium calibration model, calculate the equivalent underwater camera parameters to complete the cross-medium calibration and refraction correction. S2. The original underwater image is preprocessed by field transformation and contrast perception adjustment algorithm to obtain the preprocessed image. S3. Laser stripe extraction is performed on the laser stripe image using the Hessian-optimized CRSA U-Net laser stripe extraction method to obtain laser stripe features; S4. Based on the preprocessed image, three-dimensional reconstruction is performed by combining laser stripe features with stereo vision to generate a single-frame local point cloud. Multi-frame local point clouds are then stitched together to obtain a complete point cloud.

[0006] Furthermore: In S1, the method for calculating the equivalent underwater camera parameters using the cross-medium calibration model is as follows: First, the equivalent intrinsic parameters of the underwater image are calculated using the calibration results in air combined with the following formula; In the formula, The camera intrinsic parameter matrix calibrated in an air environment. This is the equivalent intrinsic parameter matrix for camera calibration in an underwater environment. and Focal length in air and The principal point of the image in the air. and For underwater focal length, and The principal point of the underwater image; Subsequently, the following formula is used to establish a nonlinear mapping relationship between the image in the air and the underwater refraction distortion image, thereby converting the calibration parameters in the air into underwater camera parameters and realizing underwater distortion correction; In the formula, This refers to images taken in an atmospheric environment. Images taken in an underwater environment. and The nonlinear distortion parameter is caused by the refraction of the medium. It is a second-order radial distortion. It is a fourth-order radial distortion; In the formula, For the thickness of the waterproof shell glass, , and Indicates the refractive index of the medium. This refers to the camera's focal length.

[0007] The beneficial effects of the above-mentioned further solutions are as follows: This invention proposes a physical information-driven cross-medium calibration model, which establishes an analytical mapping between multi-medium refraction and equivalent imaging parameters to achieve rapid conversion from air calibration parameters to underwater imaging parameters, thereby reducing the complexity of on-site underwater calibration.

[0008] Furthermore, S3 includes the following sub-steps: S31. Input the laser stripe image into the CRSA U-Net model and extract the ROI region of the laser stripe; S32. Based on the ROI region of the laser stripe, the laser stripe center is finely extracted using the Hessian matrix to obtain the laser stripe features.

[0009] The beneficial effects of the above-mentioned further solutions are as follows: This invention proposes a CRSA U-Net laser stripe extraction method combined with Hessian optimization, which is used for robust segmentation and sub-pixel center localization of laser stripes under strong noise and low contrast conditions, so as to improve the continuity and localization accuracy of stripe features in complex backgrounds.

[0010] Furthermore: In S31, the specific workflow of the CRSA U-Net model is as follows: S311. Input the underwater laser stripe image into the encoder, and extract multi-scale features through convolution, normalization, activation, and downsampling to obtain the feature maps of each level of the encoder. S312: The encoder output feature map of each layer is input to the channel attention module, which weights the channel importance and outputs the weighted encoder feature map. S313. The encoder outputs a high-level semantic feature map, which is further input into the decoder through the intermediate layer of the network, and the resolution is gradually restored through upsampling. S314. Each layer of the decoder is connected by skip connections, and the corresponding layer's encoder feature map after channel attention weighting is spliced ​​together. S315. The feature map output from the last layer of the decoder is input into the residual spatial attention module to enhance the spatial continuity of the laser stripes. The residual spatial attention module outputs the final feature map, which is then convolved by 1×1 to output a laser stripe segmentation mask, serving as the ROI region of the laser stripes for subsequent laser centerline extraction.

[0011] Furthermore: In S31, the CRSA U-Net model training employs cross-entropy loss. With Dice loss The objective function of weighted combination : In the formula, The output value of the CRSA U-Net model. Image data of a label with laser stripes.

[0012] Furthermore: S32 specifically refers to: A1. Based on the left and right boundaries of the ROI region, a preliminary estimate of the fringe center is made. Taking the preliminarily estimated laser center as the midpoint, several points are taken to the left and right, and all the points obtained are used as data for function fitting; among them, the position of the preliminarily estimated laser center is... The expression is: In the formula, and These are the left and right boundaries of the ROI region; A2. By analyzing the grayscale distribution of the laser stripe cross-section, a Gaussian function is applied to all obtained points. Perform fitting: In the formula, , and Let be the three unknown parameters of the Gaussian function to be solved. According to the properties of the Gaussian function, the width of the laser line is... Based on the obtained width, the variance of the Gaussian convolution is calculated as follows: ; A3. After fitting the Gaussian function, use the Hessian matrix. The coordinates of the center of the subpixel-level laser stripe are determined and used as the laser stripe feature. In the formula, Let be the distribution function of image pixel intensity, in The distribution function of image pixel intensity is expanded using a two-dimensional Taylor expansion. The expanded distribution function of image pixel intensity is... The specific expression is: In the formula, For along direction The displacement, For the image in Pixel intensity at that location Represents the gradient. This is the transpose symbol.

[0013] Furthermore: In S4, the specific method for generating a single-frame local point cloud is as follows: A joint optimization objective function based on laser stripe features and stereo vision constraints is established. The Levenberg-Marquardt algorithm is used to solve this joint optimization objective function to obtain the spatial points in the current frame. High-precision 3D coordinates in the camera coordinate system are used to further form a local point cloud of a single frame at the current scanning position; Among them, the joint optimization objective function of laser stripe features and stereo vision joint constraints. The specific expression is: In the formula, and Representing spatial points The projection functions to the left and right cameras. To balance the weighting coefficients of geometric constraints, For spatial points Projected pixel coordinates on the left camera image plane For spatial points Projected pixel coordinates on the right camera image plane This represents the spatial point corresponding to the minimum value of the function. index, Represents the norm.

[0014] Furthermore, in S4, before stitching together multi-frame local point clouds, the local point clouds are uniformly transformed to the global world coordinate system based on the scanning pose change relationship. Specifically, for a single-frame local point cloud generated at the current scanning angle, the coordinate transformation matrix corresponding to the current pose is used to map the single-frame local point cloud at that scanning position to the unified global coordinate system. The specific expression is: In the formula, This represents the rotational transformation relationship from the initial coordinate system to the current coordinate system. It is a translation vector; In the formula, This is the rotation matrix for the current scan position. During the camera's rotating scan process, around Axis rotation angle The subsequent rotation matrix; In the formula, The angle between the camera's optical axis and the rotation coordinate axis; .

[0015] The beneficial effects of the above-mentioned further solutions are as follows: This invention proposes a three-dimensional reconstruction method based on the joint constraints of laser stripe features and stereo vision, which unifies binocular epipolar geometry and laser plane constraints into the same optimization framework, and combines rotational scanning pose transformation to realize multi-frame local point cloud stitching, thereby improving the integrity and measurement accuracy of point clouds on weakly textured concrete surfaces.

[0016] The beneficial effects of this invention are as follows: (1) This invention provides a three-dimensional reconstruction method for underwater concrete surfaces based on laser stripes and stereo vision. Addressing the challenges of background interference caused by multi-medium nonlinear refraction and scattering, as well as insufficient surface texture in underwater optical measurements, traditional optical reconstruction methods struggle to achieve stable and reliable geometric measurements. This invention employs the following technical improvements: First, a cross-medium calibration model is constructed to convert air calibration parameters into underwater imaging parameters, thus avoiding cumbersome on-site underwater calibration. Second, a CRSA U-Net method optimized with Hessian is proposed for robust segmentation of laser stripes and sub-pixel-level extraction of laser stripe centers under conditions of strong noise and low contrast. Finally, a three-dimensional point cloud reconstruction method combining laser stripe features and stereo vision constraints is constructed, and multi-frame local point cloud stitching and complete region reconstruction are achieved by combining rotational scanning pose transformation. This solves the key problems of multi-medium refraction, background scattering interference, and weak texture surface matching failure in underwater optical measurements.

[0017] (2) Experimental results show that the method of the present invention exhibits high reconstruction accuracy and good stability in both standard underwater targets and real concrete defect scenarios. In the ablation experiment, the method proposed in this invention achieved the best measurement accuracy, with the overall error dropping sharply to 3.65% and the point cloud integrity reaching a near-perfect 98.6%. This leap in performance is mainly attributed to the deep synergy in two aspects: on the one hand, the CRSA U-Net model accurately strips the semantic features of laser stripes under extremely low signal-to-noise ratios, providing high-fidelity geometric input for spatial calculation; on the other hand, the deep fusion of binocular epipolar constraints and structured light plane constraints in the nonlinear optimization domain overcomes the instability of monocular ray intersection and the texture dependence of stereo matching. The data from the ablation experiment show that no single method step in the present invention can independently cope with complex underwater working conditions. The close coupling of physical correction, deep learning semantic perception, and multi-view geometric fusion is the inevitable path to achieve high-fidelity underwater three-dimensional engineering measurement. Attached Figure Description

[0018] Figure 1The flowchart of the underwater concrete surface three-dimensional reconstruction method based on laser stripes and stereo vision provided by the present invention is shown.

[0019] Figure 2 This is a schematic diagram of the CRSA U-Net model structure of the present invention.

[0020] Figure 3 This is a schematic diagram of the channel attention module structure of the present invention.

[0021] Figure 4 This is a schematic diagram of the residual space attention module structure of the present invention.

[0022] Figure 5 This is a diagram showing the relationship between the pose change of the camera coordinate system and the local point cloud stitching during the rotational scanning process of this invention. Detailed Implementation

[0023] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0024] The purpose of this invention is to provide a feasible technical path for high-precision three-dimensional information acquisition and quantitative geometric evaluation of underwater concrete surfaces by using a physical information-driven cross-medium calibration model and a CRSAU-Net laser stripe extraction method optimized by Hessian as key supporting modules, taking the three-dimensional coordinate recovery of the combined laser stripe features and stereo vision constraints as the core, and further combining rotational scanning pose transformation to realize multi-frame local point cloud stitching.

[0025] like Figure 1 As shown, in one embodiment of the present invention, the method for three-dimensional reconstruction of underwater concrete surfaces based on laser stripes and stereo vision includes the following steps: S1. Based on the camera and laser plane parameters obtained in the air, and combined with the cross-medium calibration model, calculate the equivalent underwater camera parameters to complete the cross-medium calibration and refraction correction. S2. The original underwater image is transformed by field of view (FOV) and further preprocessed using the contrast-perceptual adjustment (CPA) algorithm to obtain the preprocessed image; S3. Laser stripe extraction is performed on the laser stripe image using the Hessian-optimized CRSA U-Net (Channel and Residual-Spatial Attention U-Net) laser stripe extraction method to obtain laser stripe features; S4. Based on the preprocessed image, three-dimensional reconstruction is performed by combining laser stripe features with stereo vision to generate a single-frame local point cloud. Multi-frame local point clouds are then stitched together to obtain a complete point cloud.

[0026] In S1, the specific principle of constructing the cross-medium calibration model in this invention is as follows: In underwater imaging, light typically passes sequentially through water, a protective viewing window, and air before entering the CMOS imaging plane. Because 3D reconstruction systems require waterproof encapsulation, the actual imaging optical path often involves continuous refraction through multiple layers of parallel media. Assuming the light passes through… Parallel medium layer, first The refractive index of the layered medium Thickness is The initial angle of incidence of the light ray is The optical axis is perpendicular to each medium. According to Snell's law, light rays passing through the first medium will... When using a layered medium, for the first layer... Refraction angle in layered medium With the angle of incidence The following relationship exists between them: (1) In the formula, For the first The refractive index of the layered medium; The light in the first Lateral displacement in layered medium Total lateral displacement The specific expression is: (2) Light exhibits angular refraction propagation during refraction through multiple media. When the thickness of each medium is known, the order in which light passes through the media does not change the horizontal distance between the light's exit point and the optical axis. This embodiment verifies this conclusion using the optical simulation software PhyDemo. For any media layer, its refraction angle depends only on the refractive index and incident angle of the preceding layer. When the order of the transmission media layers changes, the final refraction angle corresponding to each layer depends only on the product of the initial incident angle and refractive index. For a target point located in water, its object distance in the underwater medium is denoted as... Its equivalent object distance It can be represented as: (3) In the formula, Indicates the equivalent parallel dielectric layer thickness. , and This represents the refractive index of the medium. Therefore, the image point coordinates under underwater imaging conditions... It can be represented as: (4) In the formula, For underwater equivalent focal length, For a point in space in the world coordinate system X Axial coordinates, For a point in space in the world coordinate system Y Axis direction coordinates; Based on the above derivation, the equivalent focal length of the camera under underwater imaging conditions is... It can be represented as: (5) In the formula, Let be the camera focal length in air. The above equation shows that multi-medium refraction can be equivalent to a change in focal length, thus providing a theoretical basis for subsequent parameter conversion based on air calibration results. According to the refraction model and camera imaging theory, underwater imaging changes the focal length of the image. Refraction also increases image distortion, affecting the lateral shift of the image principal point. Refraction distortion can be optimized by adding higher-order terms to the Brown-Conrady model, and the image principal point shift can be corrected through camera calibration. Therefore, in an air environment... underwater environment In this context, the main changes in the camera intrinsic parameter matrix are reflected in two aspects: the equivalent focal length and the principal point position. These can be expressed as follows: (6) In the formula, The camera intrinsic parameter matrix calibrated in an air environment. This is the equivalent intrinsic parameter matrix for camera calibration in an underwater environment. and Focal length in air and The principal point of the image in the air. and For underwater focal length, and This is the main point for underwater imaging. Considering that the thickness of the camera's waterproof housing is typically only 3–5 mm, its additional impact on the overall refractive light path is relatively small.

[0027] If the camera's field of view is 60° and the distance from the camera is 400mm, the field of view (FOV) in air is 461.6mm. When the glass thickness is 3mm, the FOV is 324.2mm. Ignoring the glass thickness, the FOV is 324.4mm. Therefore, the glass thickness can be ignored in the calculation.

[0028] During underwater imaging, the severe refraction of light by the waterproof shell and water not only alters the equivalent focal length but also amplifies image distortion. Previous studies have shown that this underwater distortion is not linearly amplified but rather gradually intensifies from the image center towards the edges. To further quantify the nonlinear characteristics of underwater image distortion, this embodiment statistically analyzed the corner pixel distances of the checkerboard calibration board (PG100-6-12*9) acquired underwater, and the results are shown in Table 1.

[0029] Table 1. Statistical analysis of the distances between the corners of the underwater chessboard (unit: pixels) As shown in Table 1, the corner spacing exhibits a nonlinear trend of gradually increasing from the image center to the edge. In practical engineering applications of underwater 3D reconstruction, due to the harsh operating environment, direct underwater on-site calibration is extremely difficult and inefficient. Therefore, this embodiment introduces a multi-medium refraction distortion model, derives and establishes a nonlinear mapping relationship between air images and underwater refraction distortion images, and proposes a parameter conversion method that relies solely on a single air calibration. The nonlinear mapping relationship from air images to underwater images is expressed as: (7) In the formula, This refers to images taken in an atmospheric environment. Images taken in an underwater environment. and The nonlinear distortion parameter is caused by the refraction of the medium. It is a second-order radial distortion. For fourth-order radial distortion, using Snell's law and the Taylor expansion of the refraction angle, the distortion parameters are... and The specific expression is: (8) In the formula, For the thickness of the waterproof shell glass, Let be the camera's focal length. Based on the equivalent focal length relationship and the simulation results of medium refraction in PhyDemo, the influence of the waterproof layer on the change in the optical path is negligible. Therefore, based on the refractive index of water and the calibration parameters in the air environment, the theoretical calculation results of the intrinsic parameters of the camera when taking images underwater can be obtained. According to the expression of the distortion parameter, accurately obtaining the refractive index of the water body is crucial. This is crucial for the theoretical calculation and calibration of parameters. Therefore, this embodiment uses a turbidity sensor (WGZ-1B) and a refractive index sensor (PAL-SALT) to measure the refractive index of water bodies with different turbidities. Table 2 shows the refractive index changes from clear water to highly turbid water bodies (including natural lakes). The measurement results clearly show that the influence of water turbidity on the refractive index is minimal and can be reasonably ignored.

[0030] Table 2 Relationship between turbidity and refractive index Based on the above derivation, the system can calculate the equivalent intrinsic parameters of the underwater image using the calibration results in the air combined with equation (6), and then use equation (7) to achieve underwater distortion correction. Through this method, the present invention achieves the derivation of underwater equivalent shooting parameters solely based on the air calibration results, thus eliminating the need for cumbersome on-site underwater calibration. To verify the accuracy of this theoretical calculation model, the present invention compares and evaluates the theoretically derived underwater intrinsic parameters with those obtained through traditional on-site underwater calibration. Taking one of the lenses mounted on the system as an example, the statistical results of the test data are shown in Table 3.

[0031] Table 3. Comparison and evaluation of camera calibration parameters derived theoretically and measured underwater. As shown in Table 3, the equivalent intrinsic parameter matrix obtained by the theoretical calculation method proposed in this invention has a relative deviation of only 4.5% in numerical terms compared to the result obtained by direct underwater calibration. It should be noted that this 4.5% represents only the relative deviation in numerical terms between the theoretically derived intrinsic parameters and the underwater measured intrinsic parameters, not the final system 3D dimension measurement error. To further quantify the impact of this parameter deviation on actual imaging and 3D measurement, this invention uses reprojection error for verification. After converting the underwater checkerboard image to a field of view (FOV), the reprojection error using the underwater calibration parameters is 0.49 pixels, while the reprojection error using the theoretically derived parameters from the air calibration parameters is 0.71 pixels. These results indicate that although there is a 4.5% numerical deviation between the theoretically derived intrinsic parameters and the underwater measured intrinsic parameters, the additional reprojection error transmitted to the image plane only increases by 0.22 pixels. In practical engineering applications, such minute, sub-pixel-level errors have a relatively small impact on the final 3D dimension measurement accuracy, falling within the acceptable error range for engineering measurements. Therefore, this theoretical calculation model avoids the cumbersome on-site underwater calibration process at a relatively low cost in terms of accuracy, and significantly improves the system's deployment efficiency and engineering practicality.

[0032] In S2, the Contrast-Aware Adjustment (CPA) algorithm is used to suppress background lighting noise and enhance image features.

[0033] In this embodiment, for the image acquisition stage, the present invention designs a stereo vision-assisted line structured light scanning system for 3D reconstruction of underwater concrete surfaces. This system consists of a vision module and a rotating scanning section. The vision module comprises a color camera, a monochrome camera, and a set of line lasers, with the line lasers positioned between the two cameras. The original underwater images are acquired by the color and monochrome cameras, with the color camera acting as the left camera and acquiring the image as the left view. The monochrome camera acts as the right camera and acquires the image as the right view. The rotating scanning section consists of a stepper motor and a controller.

[0034] S3 includes the following steps: S31. Input the laser stripe image into the CRSA U-Net model and extract the ROI region of the laser stripe; S32. Based on the ROI region of the laser stripe, the laser stripe center is finely extracted using the Hessian matrix to obtain the laser stripe features.

[0035] like Figure 2 As shown in Figure S31, the specific workflow of the CRSA U-Net model is as follows: S311. Input the underwater laser stripe image into the encoder, and extract multi-scale features through convolution, normalization, activation, and downsampling to obtain the feature maps of each level of the encoder. S312: The encoder output feature map of each layer is input to the channel attention module, which weights the channel importance and outputs the weighted encoder feature map. S313. The encoder outputs a high-level semantic feature map, which is further input into the decoder through the intermediate layer of the network, and the resolution is gradually restored through upsampling. S314. Each layer of the decoder is connected by skip connections, and the corresponding layer's encoder feature map after channel attention weighting is spliced ​​together. S315. The feature map output from the last layer of the decoder is input into the residual spatial attention module to enhance the spatial continuity of the laser stripes. The residual spatial attention module outputs the final feature map, which is then convolved by 1×1 to output a laser stripe segmentation mask, serving as the ROI region of the laser stripes for subsequent laser centerline extraction.

[0036] The CRSA U-Net model is used for automatic segmentation of laser stripe regions in complex underwater backgrounds. Its overall workflow is as follows: First, the acquired underwater laser stripe image is input into the encoder, where image features at different scales are extracted through continuous convolution, normalization, activation functions, and downsampling. Second, a channel attention module is introduced in the encoding stage to adaptively weight the importance of each feature channel, making the network focus more on effective features related to the laser stripes and suppressing irrelevant responses such as background scattering and ambient light noise. Then, high-level semantic features are obtained through intermediate layers of the network. Next, in the decoder, the feature map resolution is gradually restored through upsampling, and skip connections are used to fuse shallow edge information from the encoder and deep semantic information from the decoder. Finally, a residual spatial attention module enhances the continuous response of the laser stripes in spatial location, outputting a segmentation mask image of the laser stripes. This mask image is used to define the region of interest for subsequent Hessian sub-pixel center extraction, thereby improving the stability and accuracy of laser centerline extraction. The paper also explains that CRSA U-Net is used to extract the ROI of the laser stripes and combines it with the Hessian matrix to complete sub-pixel-level center localization.

[0037] like Figure 3 As shown, the channel attention module workflow is as follows: The channel attention module is used to determine the importance of different feature channels to the laser stripe recognition task. Its workflow is as follows: First, the input feature map is subjected to global average pooling and global max pooling respectively. The size after average pooling is... , To determine the number of channels, a channel description vector representing the global response intensity of each channel is obtained. This vector is then input into a shared multilayer perceptron or convolutional mapping layer to learn the channel weights for each channel. Subsequently, these channel weights are normalized to the range of 0 to 1 using a sigmoid activation function. Finally, each channel weight is multiplied channel-by-channel with the original input feature map to enhance effective channels and suppress irrelevant channels. Through this process, the model can highlight laser stripe-related features and reduce the impact of underwater scattering noise, background illumination, and local reflections on the segmentation results.

[0038] like Figure 4 As shown, the workflow of the residual space attention module is as follows: The residual spatial attention module is used to enhance the continuity and localization response of laser stripes in the image space. Its workflow is as follows: First, the input feature map is compressed along the channel dimension. The input feature map size is... , For width, To determine the batch size, a two-dimensional feature response map reflecting spatial saliency is obtained. Next, the importance of different spatial locations is learned through convolution operations, generating a spatial attention weight map. Then, the spatial weights are normalized using the sigmoid function, limiting the weight values ​​to between 0 and 1. Subsequently, the spatial attention weight map is multiplied pixel-by-pixel with the input feature map, thereby enhancing the region containing the laser stripes and suppressing non-striped background regions. Finally, the weighted feature map is residually added to the original input feature map to obtain the output feature map. This residual connection avoids the loss of effective detail information during attention weighting, while improving the recognition continuity of narrow, weak-contrast, and locally broken laser stripes.

[0039] like Figures 2 to 4 As shown, the CRSA U-Net model is based on the U-Net encoder-decoder structure. It extracts contextual semantic information of the stripe region through multi-scale feature extraction and uses skip connections to fuse shallow edge details with deep discriminative features. Specifically, the channel attention module adaptively assigns weights to different feature channels, highlighting responses related to laser stripes and suppressing redundant information caused by scattering noise and background illumination. The residual spatial attention module further enhances the continuous response of the stripes in local space, improving the model's ability to perceive weak contrast, narrow stripes, and locally missing stripes. Through this structural design, the model no longer relies solely on local grayscale differences but combines semantic and spatial structural information to robustly segment the laser stripe region.

[0040] To alleviate the severe class imbalance problem between fine laser stripes and a large background, the CRSA U-Net model training employs cross-entropy loss. With Dice loss The objective function of weighted combination : (9) In the formula, The output value of the CRSA U-Net model. For label image data with laser stripes, by assigning greater weights to the Dice loss, the model is able to focus more on the overlapping areas of the stripes.

[0041] Regarding hyperparameter settings and hardware platform for CRSA U-Net model training, this invention adopts a unified and reproducible configuration strategy. All model training and testing were performed on a workstation equipped with an NVIDIA GeForce RTX3090 (24GB VRAM) GPU and an Intel Core i9-10900K CPU, using the PyTorch deep learning framework. During the training phase, to balance convergence speed and model accuracy, the Adam optimizer was used, and the weight decay coefficient was set to [value missing]. To suppress overfitting, the network's initial learning rate is set to... Simultaneously, a cosine annealing learning rate decay strategy was introduced, allowing the learning rate to decrease smoothly in the later stages of training to explore better local minima. Considering memory limitations and data characteristics, the batch size was set to 8, and the model was trained for a total of 150 epochs. Early stopping was applied when the validation set performance tended to stabilize. Finally, a dataset containing 2000 training images and 200 test images was constructed, and the model evaluation results are shown in Table 4.

[0042] Table 4 Evaluation results of CRSA U-Net model In Table 4 This represents the proportion of pixels that actually belong to laser stripes out of all pixels predicted as laser stripes, and its expression is: (10) In the formula, For a real example, This is a false positive example. This represents the proportion of pixels correctly predicted by the model out of all real laser stripe pixels, and its expression is: (11) In the formula, This is a false negative example; The test results in Table 4 show that the CRSA U-Net model performs well on the test set. , , and overall The results show that the proposed model can preserve the main laser stripe region relatively completely under complex background noise conditions, while effectively suppressing false detections in non-stripe regions, providing a reliable segmentation prior for subsequent stripe center extraction.

[0043] To further improve the stability of the extraction results, this invention performs morphological opening operations on the mask image output by the CRSA U-Net model to remove isolated local noise points in the ROI region and avoid small-scale false detections interfering with the localization of the stripe range. After the morphological opening operation, the stripe region boundaries are more stable, and the influence of background noise is further suppressed. The results show that the combination of the CRSA U-Net model and the post-processing steps can effectively reduce the influence of ambient light and scattering noise on laser stripe extraction.

[0044] To further achieve sub-pixel center localization of laser stripes, this invention introduces the Hessian matrix to finely solve for the stripe center. The specific method is as follows: First, the fringe center is initially estimated based on the left and right boundaries of the ROI region, and its expression is as follows: (12) In the formula, The preliminary calculated location of the laser center. and These represent the left and right boundaries of the ROI region. Taking the initially estimated laser center as the midpoint, 15 points are taken to the left and right respectively, resulting in 31 points as data for function fitting.

[0045] Secondly, by analyzing the gray-scale distribution of the laser stripe cross-section, it was found that its linewidth approximately follows a Gaussian distribution; therefore, a Gaussian function was used. Perform fitting: (13) In the formula, , and Let be the three unknown parameters of the Gaussian function to be solved. According to the properties of the Gaussian function, the width of the laser line is... Based on the obtained width, the variance of the Gaussian convolution is calculated as follows: After fitting with a Gaussian function, the Hessian matrix is ​​used. Further, the sub-pixel level coordinates of the laser stripe center were determined and used as the laser stripe characteristics: (14) In the formula, Let be the distribution function of image pixel intensity. The distribution function of image pixel intensity is expanded using a two-dimensional Taylor expansion. The expanded distribution function of image pixel intensity is... The specific expression is: (15) In the formula, For along direction The displacement, For the image in Pixel intensity at that location Represents the gradient. The transpose symbol is used for... By taking the derivative and setting it to zero, we can solve for... The value when When the value is between -0.5 and 0.5, the current point is considered the laser center point. Considering the existence of multiple extreme points, a threshold is used as the standard for maximum suppression, and the point with the largest pixel intensity among all calculated center points in each row is retained as the final solution.

[0046] In summary, the CRSA U-Net model robustly identifies laser stripe regions in complex and degraded backgrounds, morphological processing enhances the continuity and stability of the ROI, and Hessian optimization further achieves sub-pixel-level geometric refinement at the stripe center. This method not only enhances the separability of laser stripes in complex underwater environments but also provides a reliable feature base for subsequent joint-constraint 3D reconstruction.

[0047] In 3D reconstruction of underwater concrete surfaces, single-modal methods struggle to simultaneously ensure point cloud integrity and measurement accuracy. For weakly textured concrete surfaces, traditional binocular matching often fails to find corresponding points or suffers from unstable disparity estimation. Furthermore, relying solely on laser stripes for spatial calculations results in stripe breaks, local reflections, and center extraction errors directly transferring to the 3D coordinates, leading to local point cloud instability. Therefore, this paper proposes a high-precision 3D reconstruction method based on joint constraints of laser stripe features and stereo vision. This method uses the sub-pixel laser stripe center extracted from the left camera as the active geometric feature, binocular epipolar geometry as a cross-viewpoint consistency constraint, and combines laser plane parameters to unify different geometric constraints into a single optimization framework for 3D coordinate recovery of spatial points, thereby improving reconstruction accuracy and point cloud integrity in degraded underwater environments.

[0048] In this embodiment, during the image preprocessing stage S2, the CPA algorithm is applied to eliminate background light interference caused by turbid water. Subsequently, in the candidate matching generation stage, the system calculates the matching rate based on the consistency of candidate features in the left and right views, and uses an empirical threshold to initially screen low-confidence responses. This invention considers regions with a matching rate greater than 15% as valid candidate regions. When the matching rate is greater than 85%, the correspondence is considered to have high reliability, and the region directly proceeds to the subsequent 3D solution process. It should be noted that the matching rate determination is only an auxiliary screening step before joint constraint reconstruction, its purpose being to eliminate spurious responses and improve the reliability of candidate matching. The core of this invention's method still lies in the 3D coordinate recovery process defined by the subsequent joint optimization objective function.

[0049] However, in complex underwater environments, relying solely on the aforementioned candidate screening mechanism is insufficient to guarantee stable reconstruction. Because concrete surfaces lack texture and are susceptible to scattering interference from suspended particles, simple stereo matching can easily produce parallax holes, while single-line structured light may exhibit fringe breaks in areas of strong reflection. Therefore, this invention introduces a joint constraint mechanism of laser fringe features and stereo vision during single-frame spatial point recovery. Assuming a spatial point... The projected pixel coordinates on the image planes of the left and right cameras are respectively and After correcting underwater nonlinear refractive distortion using the S1 method's cross-medium calibration model, the epipolar geometric constraint relationship between the left and right cameras can be established through the fundamental matrix. Represented as: (16) Equation (16) serves to limit the search range of the corresponding point in the right camera to the epipolar line determined by the left camera point, thereby utilizing binocular geometric consistency to suppress mismatches caused by weakly textured surfaces and scattering noise. In other words, stereo vision in this invention does not solely undertake depth recovery, but rather provides cross-view geometric verification for active laser features. Meanwhile, the equation of the light plane formed by the line laser in the left camera coordinate system can be expressed as: (17) In the formula, Let be the normal vector of the laser plane. This represents the distance from the origin of the coordinate system to the plane. For spatial points The high-precision three-dimensional coordinates in the camera coordinate system are obtained precisely during the calibration phase. Equation (17) gives the geometric prior of active structured light, namely, the spatial point corresponding to the center of the laser stripe should satisfy the laser plane constraint. Compared with passive binocular matching, the laser plane constraint can provide more direct local high-precision geometric information on textureless surfaces.

[0050] In conventional line structured light reconstruction, the three-dimensional coordinates can be calculated simply by finding the intersection between the left camera ray and the laser plane. However, in complex underwater environments, this method is highly sensitive to stripe quality. Therefore, this invention introduces observation information from the right camera as a joint verification mechanism. For each sub-pixel laser center point extracted by the CRSA U-Net model in the left camera image... First, the epipolar line corresponding to it in the right camera is calculated using the fundamental matrix. And search for candidate feature points that satisfy response consistency in the neighborhood of the epipolar line. Based on this, the following joint optimization objective function is constructed. : (18) In the formula, and Representing spatial points The projection functions to the left and right cameras. To balance the weighting coefficients of geometric constraints, This represents the spatial point corresponding to the minimum value of the function. index, The norm is represented. Equation (18) is the core of the method of this invention. Its essence lies in unifying the binocular reprojection consistency and the laser plane geometric prior into the same nonlinear least squares framework. The first two terms ensure that the projection results of the spatial point in the left and right viewpoints are consistent with the observed features, and the third term ensures that the spatial point satisfies the geometric constraints provided by the active structured light. Through this joint constraint design, the method of this invention can simultaneously utilize the local high-precision observation capability of laser stripe features and the cross-viewpoint consistency constraint capability of stereo vision, thereby alleviating the matching failure under weak texture conditions and the local instability of single-line structured light.

[0051] To solve equation (18), this invention employs the Levenberg-Marquardt (LM) algorithm for iterative optimization. The LM algorithm, by dynamically introducing a damping factor, can adaptively switch between the Gauss-Newton method and the gradient descent method, thereby improving the robustness of the joint constraint solution in degraded underwater environments. By solving this nonlinear least squares problem, the spatial points in the current frame can be obtained. High-precision 3D coordinates in the camera coordinate system This further forms a single-frame local point cloud at the current scanning position.

[0052] It is important to emphasize that equation (18) performs local point cloud reconstruction for a single frame, while the complete point cloud of the entire scanned area still depends on the stitching of multiple frames during the rotational scanning process. Since the acquisition system is in a continuous rotational state during scanning, the local point clouds reconstructed at different scanning angles are located in their respective camera coordinate systems. Therefore, they must be uniformly transformed to the global world coordinate system according to the scanning pose change relationship, such as... Figure 5 As shown.

[0053] Assume that the camera is initially located in Cartesian coordinates in the world coordinate system. The angle between the camera's optical axis and the rotation coordinate axis is First, tilt the camera coordinate system from its initial orientation at the tilt angle. It needs to be rotated around the Cartesian coordinate system. Axis rotation Corresponding rotation matrix Represented as: (19) During the camera's rotational scanning process, the rotation angle around the axis is [value missing]. The new Cartesian coordinates after rotation are , around Axis rotation The subsequent rotation matrix It can be represented as: (20) The rotation transformation relationship from the initial coordinate system to the current coordinate system. Represented as: (twenty one) Translation vector The expression in the initial coordinate system: (twenty two) Therefore, during the rotational scan reconstruction process, the coordinate transformation matrix corresponding to the current pose... Represented as: (twenty three) Equations (19)–(23) describe the pose change relationship of the camera coordinate system relative to the initial coordinate system during the rotational scanning process. Their physical meaning is that: for each joint constraint 3D reconstruction completed by the system at a scanning angle, the coordinate transformation matrix corresponding to the current pose is used. The local point cloud at the scanned location is mapped to a unified global coordinate system. As the scan angle continuously changes, the local point clouds recovered from different locations are accumulated and stitched together frame by frame, ultimately forming a complete 3D point cloud of the entire scanned area. In other words, Figure 5 It describes not only the changing relationship of the camera coordinate system, but also the key process of gradually expanding the observation range through rotational scanning and stitching together the local reconstruction results of multiple frames into a complete point cloud of the target area. Figure 5 middle, The x-coordinate of the position after rotation. The ordinate of the position after rotation is... The vertical coordinate of the position after rotation; In summary, the proposed joint constraint reconstruction method utilizes laser stripe features to provide high-precision active geometric observations, employs stereo vision to provide cross-viewpoint consistency constraints, and completes coordinate recovery of single-frame spatial points through a unified optimization framework. Furthermore, by incorporating pose transformation relationships during rotational scanning, it achieves global stitching and complete region reconstruction of multi-frame local point clouds. Compared to methods relying solely on binocular vision or single-line structured light, this method more effectively improves reconstruction accuracy, point cloud integrity, and measurement stability in degraded underwater environments.

[0054] In this embodiment, in order to comprehensively evaluate the effectiveness of the underwater concrete surface three-dimensional reconstruction method proposed in this invention in underwater close-range high-precision three-dimensional reconstruction, an ablation experiment was conducted.

[0055] To further verify the independent contributions and joint necessity of each key step (stereo vision epipolar constraint, CRSA U-Net deep learning feature extraction, and line structured light spatial constraint) in the proposed 3D reconstruction method, this study conducted a series of ablation experiments on an underwater standard component (Target area_1, reference dimensions: 150 mm × 80 mm × 30 mm). Three different reconstruction computation configurations were set up: Configuration A (Pure Stereo Vision): 3D reconstruction relies solely on traditional stereo vision (binocular matching), without introducing laser stripe features. Configuration B (Pure Structured Light): Reconstruction relies solely on single-line structured light, and laser stripes are extracted using the traditional gray-scale centroid method, without binocular epipolar geometric assistance. Configuration C (Our Proposed Method): The complete method flow of this paper is adopted, namely, robust stripe extraction based on the CRSA U-Net model, and deep fusion solution combining stereo vision epipolar constraint and laser plane equation.

[0056] To objectively quantify the impact of complex underwater environments on point cloud topology, this experiment, in addition to the relative errors in length, width, and depth measurements, introduced the Point Cloud Completeness index, which is the ratio of the number of effectively reconstructed 3D points to the theoretical number of surface points. The statistical results of the ablation experiments are shown in Table 5.

[0057] Table 5 Comparison of ablation experimental results under different configurations (Unit: mm) As shown in Table 5, single-sensor modalities have significant limitations in dealing with extreme underwater degradation conditions. Relying solely on traditional stereo vision (Configuration A) results in an overall relative error as high as 14.16%, with a point cloud integrity of only 68.5%. This is mainly due to the lack of high-frequency physical textures on the surface of underwater standard components, coupled with water scattering leading to decreased image contrast. This makes simple binocular epipolar line search prone to getting trapped in local extrema, causing large-scale matching failures and disparity holes. While introducing traditional active structured light (Configuration B) reduces the overall error to 8.73% and improves the point cloud integrity to 84.2%, the strong backscattering and background noise underwater make it difficult for traditional grayscale-based feature extraction algorithms to guarantee the continuity of sub-pixel stripes. Local stripe breaks still cause significant spatial calculation errors. This demonstrates that relying solely on natural passive features or traditional low-level algorithms for 3D calculation in complex underwater environments is unreliable.

[0058] In contrast, the proposed method (configuration C) achieves optimal measurement accuracy, with the overall error plummeting to 3.65% and point cloud integrity reaching a near-perfect 98.6%. This leap in performance is primarily attributed to two key synergies: firstly, CRSA U-Net accurately extracts the semantic features of laser stripes at extremely low signal-to-noise ratios, providing high-fidelity geometric input for spatial computation; secondly, the deep fusion of binocular epipolar constraints and structured light plane constraints in the nonlinear optimization domain overcomes the instability of monocular ray intersection and the texture dependence of stereo matching. Ablation experiments demonstrate that no single module in the proposed method can independently handle complex underwater conditions; the tight coupling of physical correction, deep learning semantic perception, and multi-view geometric fusion is the necessary path to achieving high-fidelity underwater 3D engineering measurement.

[0059] Based on the experimental results above, the proposed underwater concrete surface 3D reconstruction method exhibits significant advantages over both the commercial underwater laser scanner ULS-100 and typical benchmark reconstruction methods. In standardized underwater target reconstruction experiments, the average relative error of the proposed method is 3.65%, and the root mean square error (RMSE) is 1.6 mm; in contrast, the average relative error of ULS-100 is 8.64%, and the RMSE is 3.77 mm. In the quantitative measurement of typical underwater concrete defects, the average relative error of the proposed method is 4.67%, significantly lower than the 8.97% of ULS-100. Furthermore, comparative experiments with existing underwater 3D reconstruction methods further demonstrate that the proposed method has superior reconstruction accuracy and robustness in degraded underwater environments.

[0060] The aforementioned performance improvements are primarily attributed to the synergistic effect of three key modules in our proposed method. First, the physical mechanism-based cross-medium calibration model maintains geometric accuracy sufficient for engineering measurements while avoiding cumbersome on-site underwater calibration. Second, the Hessian-optimized CRSA U-Net laser fringe extraction method effectively improves the continuity and positioning accuracy of fringe extraction under conditions of strong background noise and weak contrast. More importantly, this paper constructs a 3D reconstruction model based on joint constraints of laser fringe features and stereo vision, unifying binocular reprojection consistency and laser plane geometric priors into a single optimization framework. This overcomes the limitations of single-modal stereo matching and single-line structured light methods, and further enables multi-frame local point cloud stitching through rotational scanning pose transformation. Ablation experiments further validate this, demonstrating that the complete method significantly outperforms both pure stereo vision and pure structured light configurations in terms of overall error and point cloud integrity.

[0061] In addition to its superior geometric accuracy, the proposed method can acquire high-fidelity point cloud data with accurate color information, which is significant for subsequent defect identification, semantic analysis, and digital representation. Compared to traditional monochrome underwater scanning equipment, this method can provide richer geometric and appearance information in close-range, fine-grained measurement scenarios. For the detection of local defects in underwater concrete structures such as dams and bridge foundations, this data acquisition method helps improve the reliability of quantitative defect characterization and provides a data foundation for subsequent digital assessment.

[0062] In the description of this invention, the above are merely preferred embodiments and are not intended to limit the scope of protection of this invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for three-dimensional reconstruction of underwater concrete surface based on laser stripe and stereo vision, characterized in that, Includes the following steps: S1. Based on the camera and laser plane parameters obtained in the air, and combined with the cross-medium calibration model, calculate the equivalent underwater camera parameters to complete the cross-medium calibration and refraction correction. S2. The original underwater image is preprocessed by field transformation and contrast perception adjustment algorithm to obtain the preprocessed image. S3. Laser stripe extraction is performed on the laser stripe image using the Hessian-optimized CRSA U-Net laser stripe extraction method to obtain laser stripe features; S4. Based on the preprocessed image, three-dimensional reconstruction is performed by combining laser stripe features with stereo vision to generate a single-frame local point cloud. Multi-frame local point clouds are then stitched together to obtain a complete point cloud.

2. The method according to claim 1, characterized in that, In S1, the method for calculating the equivalent underwater camera parameters using the cross-medium calibration model is as follows: First, the equivalent intrinsic parameters of the underwater image are calculated using the calibration results in air combined with the following formula; In the formula, The camera intrinsic parameter matrix calibrated in an air environment. This is the equivalent intrinsic parameter matrix for camera calibration in an underwater environment. and Focal length in air and The principal point of the image in the air. and For underwater focal length, and The principal point of the underwater image; Subsequently, the following formula is used to establish a nonlinear mapping relationship between the image in the air and the underwater refraction distortion image, thereby converting the calibration parameters in the air into underwater camera parameters and realizing underwater distortion correction. In the formula, This refers to images taken in an atmospheric environment. Images taken in an underwater environment. and The nonlinear distortion parameter is caused by the refraction of the medium. It is a second-order radial distortion. It is a fourth-order radial distortion; wherein is the thickness of the waterproof case glass, , and denotes the refractive index of the medium, is the focal length of the camera.

3. The method according to claim 1, wherein, S3 includes the following steps: S31. Input the laser stripe image into the CRSA U-Net model and extract the ROI region of the laser stripe; S32. Based on the ROI region of the laser stripe, the laser stripe center is finely extracted using the Hessian matrix to obtain the laser stripe features.

4. The method for three-dimensional reconstruction of underwater concrete surfaces based on laser stripes and stereo vision according to claim 3, characterized in that, In S31, the workflow of the CRSA U-Net model is as follows: S311. Input the underwater laser stripe image into the encoder, and extract multi-scale features through convolution, normalization, activation, and downsampling to obtain the feature maps of each level of the encoder. S312: The encoder output feature map of each layer is input to the channel attention module, which weights the channel importance and outputs the weighted encoder feature map. S313. The encoder outputs a high-level semantic feature map, which is further input into the decoder through the intermediate layer of the network, and the resolution is gradually restored through upsampling. S314. Each layer of the decoder is connected by skip connections, and the corresponding layer's encoder feature map after channel attention weighting is spliced ​​together. S315. The feature map output from the last layer of the decoder is input into the residual spatial attention module to enhance the spatial continuity of the laser stripes. The residual spatial attention module outputs the final feature map, which is then convolved by 1×1 to output a laser stripe segmentation mask, serving as the ROI region of the laser stripes for subsequent laser centerline extraction.

5. The method according to claim 3, wherein, In S31, the CRSA U-Net model training adopts cross-entropy loss with Dice loss a weighted combination of target functions : In the formula, is an output value of the CRSA U-Net model, is a label image data with laser stripes.

6. The method according to claim 3, wherein, S32 specifically refers to: A1、According to the left and right boundaries of the ROI region, the center of the stripe is preliminarily estimated, and a plurality of points are taken to the left and right respectively with the preliminarily estimated laser center as the midpoint, and all the points are taken as the data for function fitting; wherein the position of the preliminarily estimated laser center is The expression is: In the formula, and are the left and right boundaries of the ROI region; A2, by analyzing the gray scale distribution of the laser stripe cross section, according to the obtained all points using Gaussian function fitting: In the formula, , and Let be the three unknown parameters of the Gaussian function to be solved. According to the properties of the Gaussian function, the width of the laser line is... Based on the obtained width, the variance of the Gaussian convolution is calculated as follows: ; A3, after Gaussian function fitting, using Hessian matrix The sub-pixel level laser stripe center coordinates are solved as the laser stripe features; In the formula, Let be the distribution function of image pixel intensity, in The distribution function of image pixel intensity is expanded using a two-dimensional Taylor expansion. The expanded distribution function of image pixel intensity is... The specific expression is: In the formula, For along direction The displacement, For the image in Pixel intensity at that location Represents the gradient. This is the transpose symbol.

7. The method according to claim 6, characterized in that, In S4, the specific method for generating a single-frame local point cloud is as follows: A joint optimization objective function of the laser stripe feature and the stereo vision joint constraint is established, and a Levenberg-Marquardt algorithm is used to solve the joint optimization objective function, so as to obtain a space point under a current frame A high-precision three-dimensional coordinate in a camera coordinate system is further formed into a single-frame local point cloud at a current scanning position. In the expression, the joint optimization objective function of the laser stripe feature and the stereo vision joint constraint The expression is specifically: In the formula, and Representing spatial points The projection functions to the left and right cameras. To balance the weighting coefficients of geometric constraints, For spatial points Projected pixel coordinates on the left camera image plane For spatial points Projected pixel coordinates on the right camera image plane This represents the spatial point corresponding to the minimum value of the function. index, Represents the norm.

8. The method according to claim 7, characterized in that, In S4, before stitching together multi-frame local point clouds, the local point clouds are uniformly transformed to the global world coordinate system based on the scanning pose change relationship. Specifically, for a single-frame local point cloud generated at the current scanning angle, the coordinate transformation matrix corresponding to the current pose is used to map the single-frame local point cloud at that scanning position to the unified global coordinate system. The specific expression is: In the formula, is a rotation transformation relationship of the initial coordinate system to the current coordinate system, is a translation vector; wherein is the rotation matrix for the current scan position, is the rotation matrix after the rotation around the axis by the rotation angle axis by the rotation angle In the formula, is the angle between the camera optical axis and the rotation coordinate axis; 。