Road segmentation and gradient estimation method based on multi-sensor fusion
By registering and data fusion of cameras and lidars, road segmentation results are generated and slope angles are estimated, the problem of inaccurate slope angle estimation in the prior art is solved, and high-precision road segmentation and slope detection are achieved, which is suitable for complex environments.
Patent Information
- Application Number
- CN202510626471.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-15
AI Technical Summary
The prior art cannot accurately estimate the slope angle, making it difficult to meet the needs of high-precision pavement segmentation and slope detection, especially in complex environments.
By registering the camera and lidar, road segmentation results are generated and slope estimation of the road surface is realized based on three-dimensional point cloud images. Specific steps include time synchronization, spatial registration, back projection of road segmentation results and calculation of slope angle.
It effectively reduces the error rate of road surface segmentation and slope estimation errors, ensures that stable output can be maintained in complex environments such as vibration, rain, fog, and low light, and supports access to more sensors to adapt to special scenarios.
Smart Images

Figure CN120125592A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a road segmentation and slope estimation method based on multi-sensor fusion. Background Art
[0002] The environmental perception of a single sensor has great limitations. The RGB camera is limited by lighting conditions and cannot perceive depth information, while lidar performs poorly in complex weather such as rain and snow. In the fields of intelligent transportation and autonomous driving, as a core technical link, the accuracy and reliability of environmental perception directly affect the quality of system decision-making. At present, the limitations of the application of a single sensor are as follows: The RGB camera realizes scene recognition by virtue of high-resolution texture information, but is easily interfered by lighting conditions. In a low-illumination environment at night, the signal-to-noise ratio of the image drops sharply, resulting in the failure of feature extraction; in a scene with direct strong light or backlight, the overexposure / underexposure problem of the image causes the target edge to be blurred. At the same time, the RGB camera is essentially a two-dimensional imaging device and lacks the ability to directly obtain the depth information of the scene, making it difficult to accurately model the distance of obstacles and the undulation of the road surface. Although lidar can obtain high-precision three-dimensional point cloud data through the time-of-flight principle (TOF), its performance decays significantly under complex meteorological conditions. In rainy and snowy weather, Mie scattering occurs between the laser beam and suspended particles, resulting in disordered echo signals, and phenomena such as point cloud data jumps and false alarms; in a sandy environment, the sensor window is polluted, leading to a decrease in ranging accuracy, and even a complete loss of perception ability in extreme cases. In the prior art, the perception blind area of a single sensor has become a key bottleneck restricting the robustness of the autonomous driving system. Taking the ramp working condition as an example, the traditional vision scheme cannot accurately estimate the slope angle, and lidar is prone to sparse point cloud problems on a wet and slippery road surface. Both are difficult to meet the requirements of high-precision road surface segmentation and slope detection. Therefore, it is urgent to construct a complementary perception system through multi-sensor information fusion technology to adapt to complex road environments. Summary of the Invention
[0003] In view of this, the present invention aims to provide a road segmentation and slope estimation method based on multi-sensor fusion to solve the problems that the prior art cannot accurately estimate the slope angle and is difficult to meet the requirements of high-precision road surface segmentation and slope detection. The present invention effectively reduces the error rate of road surface segmentation and the error of slope estimation through multi-sensor complementarity, and still maintains stable output under complex backgrounds such as vibration, rain, fog, and low light.
[0004] To achieve the above object, the technical solution of the present invention is realized as follows: A road segmentation and slope estimation method based on multi-sensor fusion specifically includes the following steps: S1: Register the camera and the lidar; S2: Generate a road segmentation result based on the output data of the registered camera and lidar; S3: Obtain a 3D point cloud image of the road surface based on the road segmentation result, and implement slope estimation of the road surface based on the 3D point cloud image.
[0005] Furthermore, step S1 specifically includes the following steps: S11: Use ApproximateTimeSynchronizer to set an approximate time, and match the timestamps of the output data of the camera and the lidar based on the approximate time to achieve time synchronization of the camera and the lidar; S12: Use a checkerboard as a planar target, and simultaneously collect data on the planar target using the camera and the lidar to obtain a camera image and a lidar point cloud correspondingly; S13: Extract the feature points of the planar target from the camera image, and detect the black and white corner points of the planar target based on the Harris corner detection algorithm; In the lidar point cloud, use the RANSIC plane fitting algorithm to fit the point cloud plane of the planar target, and calculate the transformation relationship between the point cloud plane and the plane where the black and white corner points are located through the following formula to achieve spatial synchronization of the camera and the lidar, and complete the registration of the camera and the lidar: ;
[0006]
[0007] where are respectively the angles of rotation of the point cloud around the x-axis, y-axis, and z-axis of the lidar coordinate system, R is the rotation matrix, T is the translation matrix, ( , , ) is the coordinate position of the point cloud in the camera coordinate system before rotation, ( , , ) is the coordinate position of the point cloud in the lidar coordinate system after rotation, ( , , ) respectively represent the translation vectors of the point cloud in the x-direction, y-direction, and z-direction, and ( , , ) respectively represent the matrices of rotation of the point cloud around the x-axis, y-axis, and z-axis.
[0008] Furthermore, step S2 specifically includes the following steps: S21; Preprocess the lidar point cloud and camera image respectively, and obtain a point cloud sparse map and a corrected image correspondingly; S22: Obtain a trained lightweight segmentation model, input the point cloud sparse map and the corrected image into the lightweight segmentation model for processing, obtain a point cloud confidence map and an image confidence map, and perform result-level fusion on the point cloud confidence map and the image confidence map to obtain a binary result; S23: Construct a LiDAR confidence evaluation formula and a visual confidence evaluation formula, and perform weight assignment on the point cloud confidence map and the image confidence map according to the LiDAR confidence evaluation formula and the visual confidence evaluation formula to obtain a road segmentation result.
[0009] Further, in step S21, the steps of preprocessing the lidar point cloud to obtain a point cloud sparse map include: S21A1: Take the x-axis position, y-axis position, and z-axis position of each point in the lidar point cloud; S21A2: Horizontally splice the rotation matrix and the translation vector to obtain a 4×3 transformation matrix, and expand the bottom of the transformation matrix to obtain a 4x4 homogeneous transformation matrix; S21A3: After transforming each point in the lidar point cloud processed in step S21A1 through the transposed homogeneous transformation matrix, take the first 3 column elements of the transformed matrix as the point cloud transformation matrix, and use the elements of the point cloud transformation matrix as the new 3D coordinates corresponding to each point in the lidar point cloud; The first column element of the point cloud transformation matrix corresponds to the x value of each point in the lidar point cloud, the second column element of the point cloud transformation matrix corresponds to the y value of each point in the lidar point cloud, and the third column element of the point cloud transformation matrix corresponds to the z value of each point in the lidar point cloud; S21A4: Remove all points with negative x values in the point cloud transformation matrix, take the first two column elements of the point cloud transformation matrix after removing the corresponding points as two-dimensional image coordinate values, use the camera internal parameters and distortion parameters to convert the two-dimensional image coordinate values into a 2D image coordinate matrix, and round the values of the 2D image coordinate matrix; S21A5: Construct a depth map based on the z values of each point in the point cloud transformation matrix and the 2D image coordinate matrix, and perform normalization processing on the depth map so that the values of the pixel points included in the depth map are in the range of 0 to 255 to obtain a point cloud sparse map; The steps of preprocessing the camera image to obtain a corrected image include: S21B1: Based on the checkerboard images taken from multiple angles, use the Zhang Zhengyou calibration algorithm to calculate the camera internal parameter matrix and the distortion coefficient, and establish the mapping relationship between the camera coordinate system and the image coordinate system constructed by the camera image taken by the camera; S21B2: Based on the calibration result of step S21B1, perform real-time correction on the camera image using bilinear interpolation to obtain a corrected image: ; ; ; where r is the Euclidean distance from the pixel point with two-dimensional coordinates in the camera image to the center of the camera image, is the two-dimensional coordinate of the corrected image, and are the radial distortion coefficients of the camera, and are the tangential distortion coefficients of the camera.
[0010] Furthermore, the lightweight segmentation model includes an encoder, a decoder, and a multi-scale evidence collection module. Input the corrected image and the point cloud sparse map into the encoder-decoder structure for feature extraction to obtain a first feature map and a second feature map respectively. Input the first feature map and the second feature map into the multi-scale evidence collection module for processing, and obtain a binary result after result-level fusion; The multi-scale evidence collection module includes a first branch and a second branch with the same network structure. The first branch includes a first sub-branch, a second sub-branch, and a third sub-branch. The first sub-branch, the second sub-branch, and the third sub-branch all contain a convolution, an upsampling layer, and a softplus activation layer connected in sequence. Among them, the convolution of the first sub-branch is a 1×1 convolution, the convolution of the second sub-branch is a 3×3 convolution, and the convolution of the third sub-branch is a 5×5 convolution.
[0011] Furthermore, step S23 includes: S231: Construct a LiDAR confidence evaluation formula and a visual confidence evaluation formula : ; ; where the fitting residual is the deviation between the predicted value and the true value output by the lightweight segmentation model, E is the probability map entropy value output by the encoder-decoder structure, L is the light intensity, is the maximum light value obtained according to prior knowledge; S232: Calculate the final road surface through the following formula and use the final road surface as the road segmentation result: .
[0012] Furthermore, step S3 specifically includes the following steps: S31: Reproject the road segmentation result back to the point cloud to obtain the 3D point cloud image of the road surface; S32: Segment the 3D point cloud image, and use the RANSAC algorithm to fit a plane to each section of the road surface, and calculate the relative slope angle of each section of the road surface according to the plane normal vector obtained after fitting each section of the road surface : ; where, is the plane normal vector obtained after fitting the i-th section of the road surface, is the arccosine function.
[0013] Furthermore, if the vehicle-mounted IMU provides the plane normal vector of the vehicle body itself , calculate the absolute slope angle of each section of the road surface through the following formula : .
[0014] Furthermore, in step S31, reproject the road segmentation result back to the point cloud through the following formula: ; ; where, x and y are the calculated x value and y value in the camera coordinate system corresponding to the z value of the current point, u and v are the pixel coordinates in the image coordinate system corresponding to the z value of the current point, P is the camera intrinsic matrix, z is the depth value of the current point, P(0, 0) and P(1, 1) are both the focal lengths of the camera, and P(0, 2) and P(1, 2) are both the principal point positions of the camera.
[0015] Compared with the prior art, the present invention can achieve the following beneficial effects: (1) For the road segmentation and slope estimation method based on multi-sensor fusion of the present invention, multi-sensor complementarity effectively reduces the error rate of road surface segmentation and the error of slope estimation; it still maintains stable output in complex environments such as vibration, rain and fog, and low light; it achieves low latency through low-error time synchronization and simplified Lidar input; it supports access to more sensors (such as millimeter-wave radar) to adapt to special scenarios.
[0016] (2) The method for road segmentation and slope estimation based on multi-sensor fusion according to the present invention for the first time uses the method of fusing lidar and camera to estimate the slope; innovatively converts the point cloud into a sparse point cloud map, and uses the sparse point cloud map instead of the point cloud as the input of the road segmentation network (lightweight segmentation model), greatly reducing the computational complexity brought by the multi-channel data input into the network; by obtaining the confidence of the RGB image and the point cloud data and performing weighted fusion, it can effectively achieve complementary advantages and increase the accuracy of road segmentation. The present invention is of great significance for autonomous driving, robot navigation, intelligent transportation systems, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings: Figure 1 is a schematic flow chart of the method for road segmentation and slope estimation based on multi-sensor fusion according to the embodiment of the present invention; Figure 2 is a schematic structural diagram of the method for road segmentation and slope estimation based on multi-sensor fusion according to the embodiment of the present invention; Figure 3 is a schematic diagram of the transformation from the lidar coordinate system to the camera coordinate system according to the embodiment of the present invention; Figure 4 is a schematic diagram of the network structure of the lightweight segmentation model according to the embodiment of the present invention; Figure 5 is a schematic diagram of the fusion of the point cloud road surface and the image road surface according to the embodiment of the present invention; Figure 6 is a schematic diagram of the point cloud road surface after segmentation according to the embodiment of the present invention; Figure 7 is a schematic diagram of the calculation of the absolute slope angle according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention.
[0019] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0020] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.
[0021] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "mounted", "connected", "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.
[0022] The present invention will be described in detail below with reference to the drawings and in combination with embodiments.
[0023] As Figure 1 - Figure 2 shown, the present invention provides a road segmentation and slope estimation method based on multi-sensor fusion, which specifically includes the following steps: S1: Register the camera and the lidar; S2: Generate a road segmentation result based on the output data of the registered camera and lidar; S3: Obtain a three-dimensional point cloud image of the road surface based on the road segmentation result, and implement slope estimation of the road surface based on the three-dimensional point cloud image.
[0024] Through the synchronous registration of lidar and camera data, the present invention realizes the efficient fusion of these two sensors in time and space. The lidar can provide rich depth information, while the camera captures detailed image texture information. By combining these two types of information, the road surface can be more accurately identified and segmented. During the process of road surface segmentation, first, the multi-source data is preprocessed and optimized to ensure the accurate alignment of lidar and camera data in the same coordinate system. Then, using the texture features in the image, the road and other objects can be effectively distinguished, thus extracting a clear road area. Next, the segmented road surface is back-projected and converted back into three-dimensional space. This process can not only retain the geometric features of the road surface but also provide accurate basic data for subsequent three-dimensional point cloud processing. On this basis, plane fitting analysis is performed on the three-dimensional point cloud to identify the planar structure of the road. Finally, through the analysis of the plane fitting results, the slope angle of the road ahead can be calculated.
[0025] That is, the present invention first realizes the collaborative work of multiple sensors, including camera-lidar time synchronization, camera-lidar spatial registration, and camera-lidar joint calibration; secondly, road segmentation is performed based on the data after spatio-temporal alignment, including RGB images (camera images) and point cloud sparse maps, and the fusion of multi-source heterogeneous data adopts a result-level fusion strategy. Then, the two-dimensional "RGB image + point cloud sparse map" fusion segmentation result is back-projected into a three-dimensional point cloud to obtain a three-dimensional point cloud image of the road; finally, the road point cloud plane is fitted, the normal vector of the road plane is calculated, and the slope angle of the road ahead is obtained.
[0026] In some embodiments, step S1 specifically includes the following steps: S11: Use ApproximateTimeSynchronizer to set an approximate time, and match the timestamps of the respective output data of the camera and the lidar based on the approximate time to achieve the time synchronization of the camera and the lidar; S12: Use a checkerboard as a planar target, and simultaneously collect data from the planar target using the camera and the lidar to obtain the camera image and the lidar point cloud respectively; S13: Extract the feature points of the planar target from the camera image, and detect the black and white corner points of the planar target based on the Harris corner detection algorithm; In the lidar point cloud, use the RANSIC plane fitting algorithm to fit the point cloud plane of the planar target, and calculate the transformation relationship between the point cloud plane and the plane where the black and white corner points are located through the following formula to achieve the spatial synchronization of the camera and the lidar and complete the registration of the camera and the lidar: ;
[0027]
[0028] wherein, are respectively the angles of the point cloud rotating around the x-axis, y-axis and z-axis of the lidar coordinate system, R is the rotation matrix, T is the translation matrix, ( , , ) is the coordinate position of the point cloud before rotation in the camera coordinate system, ( , , ) is the coordinate position of the point cloud after rotation in the lidar coordinate system, ( , , ) respectively represent the translation vectors of the point cloud in the x-direction, y-direction and z-direction, ( , , ) respectively represent the matrices of the point cloud rotating around the x-axis, y-axis and z-axis.
[0029] It should be noted that in order to prevent the scene information obtained between the lidar and the camera from being inconsistent, which may lead to errors in the subsequent data fusion and segmentation processes, it is necessary to perform time and space synchronization processing on these two sensors (the camera and the lidar, i.e., multi-sensors).
[0030] In terms of time synchronization, since there may be different delays in data acquisition of these two sensors, it may lead to inconsistent images captured by the two sensors in a high-speed motion scene. Through time synchronization, data inconsistency caused by time differences can be avoided, thereby improving the accuracy of subsequent fusion.
[0031] A soft synchronization method is adopted for time alignment. The tool used is the message filter message_filters of the ROS system. The message filter message_filters is similar to a message buffer. When a message arrives at the message filter, it will not be output immediately, but at a later time point and under certain conditions. In the present invention, the nearest neighbor frame is found by looking for adjacent timestamps. The method adopted is to use the ApproximateTimeSynchronizer (approximate time synchronizer) for time matching. Since the probability that the timestamps of data collected by different sensors are the same is close to zero, it is necessary to set a time threshold and use the approximate time synchronizer to match the messages of the two sensors with a time error less than the given threshold. The proximity time can be customized, rather than the absolute time being exactly the same. In this way, time alignment can be achieved through the approximate time synchronizer to achieve the purpose of synchronous callback output.
[0032] In terms of spatial synchronization, first, camera feature extraction is carried out: using a checkerboard as a planar target, the camera collects images of the planar target, and feature points of the planar target are extracted from the camera images. The black and white corner points in the checkerboard are detected through the Harris corner detection algorithm, and the formula used is: ; where R is the response of the Harris corner detection algorithm, used as a corner judgment flag bit, , is the determinant of the matrix, is the trace of the matrix, and are the change components in two orthogonal directions, and k is an empirical constant within the range (0.04, 0.06).
[0033] In terms of spatial synchronization, lidar feature extraction: three-dimensional points corresponding to the camera feature points are extracted from the lidar point cloud, and the plane of the calibration target is fitted through the RANSIC plane fitting algorithm.
[0034] The coordinate transformation relationship between the lidar and the camera is constructed, and the least squares method is used to estimate the external parameters (rotation matrix and translation vector) between the camera and the lidar, which is achieved by minimizing the distance error between the reprojection error of the pixel points in the camera image and the lidar point cloud.
[0035] As Figure 3 shown, the coordinate transformation relationship between the lidar and the camera is as follows: ; ; ; where, are the angles of rotation of the point cloud around the x-axis, y-axis, and z-axis of the lidar coordinate system respectively, R is the rotation matrix, T is the translation matrix, ( , , ) is the coordinate position of the point cloud before rotation corresponding to the camera coordinate system, ( , , ) is the coordinate position of the point cloud after rotation corresponding to the lidar coordinate system, ( , , ) respectively represent the translation vectors of the point cloud in the x-direction, y-direction, and z-direction, and ( , , ) respectively represent the matrices of the point cloud rotating around the x-axis, y-axis, and z-axis.
[0036] In some embodiments, step S2 specifically includes the following steps: S21: Preprocess the lidar point cloud and camera image respectively to obtain a point cloud sparse map and a corrected image correspondingly; S22: Obtain a trained lightweight segmentation model, input the point cloud sparse map and the corrected image into the lightweight segmentation model for processing to obtain a point cloud confidence map and an image confidence map, and perform result-level fusion on the point cloud confidence map and the image confidence map to obtain a binary result; S23: Construct a LiDAR confidence evaluation formula and a visual confidence evaluation formula, and perform weight assignment on the point cloud confidence map and the image confidence map according to the LiDAR confidence evaluation formula and the visual confidence evaluation formula to obtain a road segmentation result.
[0037] It should be noted that in order to effectively estimate the slope of the road surface, other interferences are excluded by segmenting the road area. Specifically: First, preprocess the data obtained by the camera and lidar respectively to reduce the computational complexity and improve the data quality; then design a dual-branch structure network (lightweight segmentation model), and obtain a point cloud confidence map and an image confidence map by performing feature extraction and learning on the RGB image (camera image) and the point cloud sparse map; finally, perform result-level fusion on the two confidences to finally generate a fused road surface segmentation result.
[0038] Point cloud data preprocessing: Since the amount of point cloud data obtained by the lidar is huge, for tasks with high real-time requirements such as autonomous driving, directly processing the point cloud data will result in slow processing speed. Therefore, by preprocessing the three-dimensional point cloud, it is converted from a three-dimensional form to a two-dimensional point cloud sparse map, reducing the computational complexity while retaining useful information.
[0039] The specific steps for preprocessing the lidar point cloud to obtain a point cloud sparse map include: a. Preprocess the obtained lidar point cloud data. The common point cloud data has 4 bits, namely: the x-axis position, the y-axis position, the z-axis position, and the reflectivity. The reflectivity does not need to be reflected in the point cloud sparse map, so only the first three bits are taken during the data preprocessing process.
[0040] b. Obtain calibration information and construct a transformation matrix: Horizontally concatenate the rotation matrix and the translation vector to form a 4x3 transformation matrix. Expand it at the bottom to extend it to a 4x4 homogeneous transformation matrix to be applicable to 3D transformation.
[0041] The purpose of expanding the matrix is to enable it to perform calculations with other matrices. To avoid the impact of the expanded values on matrix operations, 1 should be filled in the main diagonal position and 0 in the remaining positions. At this point, the fourth row elements to be expanded are [0, 0, 0, 1].
[0042] c. Convert the point cloud coordinate system to the camera coordinate system: Apply the rotation and translation transformation to the point cloud using matrix multiplication, and transform each point of the lidar point cloud through the transposed homogeneous transformation matrix. Extract the first 3 columns of the transformed result to represent the new 3D coordinates.
[0043] d. Since the camera's perspective is forward, the point cloud at the rear naturally cannot be projected onto the camera. Therefore, remove the points in the point cloud that are behind the camera, that is, remove the points with negative X in the new 3D coordinates.
[0044] e. Use the camera's internal parameters and distortion parameters to convert the 3D points to 2D image coordinates. The rotation and translation here are both zero because they have been combined in the previous steps. Then, convert the rotated point cloud to integer values by rounding.
[0045] f. According to the coordinate values of the third dimension in the point cloud matrix after the above transformation, obtain the pixel values corresponding to the depth values. Generate a depth map of the same size as the image, and normalize this depth map so that its values are in the range of 0 to 255, that is, generate the sparse point cloud map required by the network.
[0046] Preprocess the camera image: Collect checkerboard images taken from multiple angles, and calculate the camera's internal parameter matrix and distortion coefficients (radial distortion and tangential distortion and ) through Zhang Zhengyou calibration algorithm, and establish the mapping relationship between the camera coordinate system and the image coordinate system.
[0047] Based on the calibration results (part of which is the camera's internal parameter matrix and distortion coefficients, and the other part is the mapping relationship between the camera coordinate system and the image coordinate system (established based on the images captured by the camera)), use bilinear interpolation to perform real-time correction on the camera image: ; ; ; where r is the Euclidean distance from the pixel point with two-dimensional coordinates in the camera image to the center of the camera image, is the two-dimensional coordinate of the corrected image, and is the radial distortion coefficient of the camera, and is the tangential distortion coefficient of the camera.
[0048] Next, the noise of the corrected image is reduced by smoothing the pixel values or using local neighborhood information, thereby preserving the main features and details of the corrected image. This denoising step helps subsequent image processing tasks and enables higher-precision image segmentation.
[0049] In some embodiments, the lightweight segmentation model includes an encoder, a decoder, and a multi-scale evidence collection module. The corrected image and the sparse point cloud map are input into the encoder-decoder structure for feature extraction, and the first feature map and the second feature map are obtained correspondingly. The first feature map and the second feature map are input into the multi-scale evidence collection module for processing, and a binary result is obtained after result-level fusion; The multi-scale evidence collection module includes a first branch and a second branch with the same network structure. The first branch includes a first sub-branch, a second sub-branch, and a third sub-branch. The first sub-branch, the second sub-branch, and the third sub-branch all include a convolution, an upsampling layer, and a softplus activation layer connected in sequence. Among them, the convolution of the first sub-branch is a 1×1 convolution, the convolution of the second sub-branch is a 3×3 convolution, and the convolution of the third sub-branch is a 5×5 convolution.
[0050] Furthermore, as Figure 4 shown, the lightweight segmentation model is designed based on the encoder-decoder structure. In the encoder part, a feature extraction network based on depthwise separable convolution optimized V3 (MobileNet-V3) is adopted. In addition, an attention mechanism CBAM module (Convolutional Block Attention Module) is introduced in the decoder to improve the recovery ability of road edge details. The encoder-decoder structure is the same as the U-net structure. The sparse point cloud map and the corrected RGB (corrected image) are input in parallel to the feature extraction network (parameters are not shared), and the corresponding feature maps are obtained. The two feature maps are respectively fed into the multi-scale evidence collection module to generate the road segmentation results of the point cloud and the image. The multi-scale evidence collection module consists of three parallel branches. In each branch, first, convolution is used to obtain feature maps of multiple scales with two channels. The convolutions in these branches are 1×1, 3×3, and 5×5 respectively. Then, an upsampling layer is used to magnify the two-channel feature map to the size of the original image. Finally, a softplus activation layer is used at the end of the branch to ensure that the output result is non-negative. The result-level fusion is represented by a binary image, that is, the road is 1 and the non-road is 0. The road area and the non-road area of the result-level fusion are determined based on the binary result, and the road area is divided into a point cloud confidence map and an image confidence map based on the characteristics of the point cloud and the RGB image.
[0051] Calculate the confidence levels of the road segmentation results of the point cloud and the image respectively (the calculation method is as follows) to obtain the corresponding fusion weights, and then calculate the final road segmentation result according to the respective fusion weights of the two.
[0052] 1) LiDAR confidence evaluation formula: Calculate the LiDAR confidence according to the point cloud density (such as number of points / m²) and the fitting residual : ; Among them, the fitting residual is the deviation between the predicted value and the true value output by the lightweight segmentation model, and the fitting residual threshold is set according to user requirements.
[0053] 2) Visual confidence evaluation formula: Based on the entropy value of the probability map output by the lightweight segmentation model and the lighting condition Calculate the visual confidence : ; Among them, E is the entropy value of the probability map output by the encoder-decoder structure (the sum of the entropy values calculated from the first feature image and the second feature image respectively), L is the lighting intensity, is the maximum lighting value obtained according to prior knowledge.
[0054] Such as Figure 5 shown, determine the result-level fusion weight assignment according to the confidence calculation result, and the calculation method of the final road surface segmentation result is as follows: .
[0055] In some embodiments, step S3 specifically includes the following steps: S31: Reproject the road segmentation result back to the point cloud to obtain the three-dimensional point cloud image of the road surface; S32: Segment the three-dimensional point cloud image, and use the RANSIC algorithm to perform plane fitting on each section of the road surface, and calculate the relative slope angle of each section of the road surface according to the plane normal vector obtained after fitting each section of the road surface : ; Among them, is the plane normal vector obtained after plane fitting of the i-th section of the road surface, is the arccosine function.
[0056] The slope estimation consists of three steps: the first step is point cloud back-projection, the second step is plane fitting, and the third step is slope estimation. The specific implementation process is as follows: 1) Point cloud back-projection As Figure 6 shown, the segmentation result of the final road surface is a two-dimensional mask, lacking depth information. Therefore, slope estimation cannot be performed yet. So, it is necessary to re-project it back into the point cloud and use the three-dimensional measurement values of the point cloud to estimate the slope of the road ahead. The specific process is as follows: First, according to the projection matrix P of the given camera, the depth value z, and the pixel coordinates u, v (i.e., the pixel coordinates of the road surface after road segmentation), calculate the three-dimensional coordinates x and y in the camera coordinate system corresponding to the depth value z. The formula is: ; ; where x and y are the calculated x value and y value in the camera coordinate system corresponding to the z value of the current point, u and v are the pixel coordinates in the image coordinate system corresponding to the z value of the current point, P is the camera internal parameter matrix, z is the depth value of the current point, P(0, 0) and P(1, 1) are both the focal lengths of the camera, and P(0, 2) and P(1, 2) are both the principal point positions of the camera.
[0057] Convert the 2D coordinate system of the road surface to the 3D coordinate system of the camera through the above formula. Then, stack all the pixel points of the calculated three-dimensional coordinates into a two-dimensional array to form points in the camera coordinate system. Invert the rotation matrix and translation matrix obtained through spatial synchronization, and transform the points in the obtained camera coordinate system with this matrix to the lidar coordinate system. In this way, generate the final three-dimensional point coordinates, and the returned result only contains the three coordinate components of x, y, and z.
[0058] In this way, the point cloud information containing only the road surface can be obtained. At the same time, the point cloud points of non-road surfaces can be effectively removed, reducing the complexity of fitting the road surface in subsequent slope estimation.
[0059] 2) Plane fitting + slope estimation The three-dimensional point cloud image obtained through back-projection is segmented according to the distance d. For example , generally, the higher the segmentation accuracy, the more accurate the slope estimation. Use the RANSIC algorithm to perform plane fitting on each section of the road surface. The normal vector of the plane after fitting is , and at this time, the relative slope angle of the front slope can be calculated : .
[0060] As Figure 7 shown, if the vehicle-mounted IMU provides the normal vector of the vehicle body itself , calculate the absolute slope angle of each section of the road surface through the following formula : 。
[0061] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in the disclosure of the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and no limitation is imposed herein.
[0062] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A road segmentation and slope estimation method based on multi-sensor fusion, characterized in that: The specific steps include: S1: align the camera and lidar; S2: Generate road segmentation results based on the output data of the registered camera and lidar; S3: Obtain a three-dimensional point cloud image of the road surface based on the road segmentation result, and estimate the slope of the road surface based on the three-dimensional point cloud image.
2. The road segmentation and slope estimation method based on multi-sensor fusion according to claim 1 is characterized by: Step S1 specifically includes the following steps: S11: Use ApproximateTimeSynchronizer to set the approach time, match the timestamps of the camera and lidar output data based on the approach time, and achieve time synchronization between the camera and lidar; S12: Using the chessboard as a plane target, the camera and the laser radar are used to collect data on the plane target at the same time, and the camera image and the laser radar point cloud are obtained accordingly; S13: extracting feature points of the plane target from the camera image, and detecting black and white corner points of the plane target based on the Haris corner point detection algorithm; In the lidar point cloud, the RANSIC plane fitting algorithm is used to fit the point cloud plane of the planar target. The transformation relationship between the point cloud plane and the plane where the black and white corner points are located is calculated by the following formula to achieve spatial synchronization between the camera and the lidar, and complete the registration of the camera and the lidar: ; in, are the angles of rotation of the point cloud around the x-axis, y-axis, and z-axis of the laser radar coordinate system, R is the rotation matrix, and T is the translation matrix. ( , , ) is the coordinate position of the point cloud before rotation in the camera coordinate system, ( , , ) is the coordinate position of the rotated point cloud corresponding to the laser radar coordinate system, ( , , ) represent the translation vectors of the point cloud in the x, y and z directions respectively, ( , , ) represent the matrices for rotating the point cloud around the x-axis, y-axis, and z-axis, respectively.
3. The road segmentation and slope estimation method based on multi-sensor fusion according to claim 1, characterized in that: Step S2 specifically includes the following steps: S21; preprocessing the laser radar point cloud and the camera image respectively, and obtaining a point cloud sparse map and a corrected image correspondingly; S22: obtaining a trained lightweight segmentation model, inputting the point cloud sparse map and the corrected image into the lightweight segmentation model for processing, obtaining a point cloud confidence map and an image confidence map, and performing result-level fusion on the point cloud confidence map and the image confidence map to obtain a binarization result; S23: constructing a LiDAR confidence evaluation formula and a visual confidence evaluation formula, and performing weight allocation on the point cloud confidence map and the image confidence map according to the LiDAR confidence evaluation formula and the visual confidence evaluation formula to obtain a road segmentation result.
4. The road segmentation and slope estimation method based on multi-sensor fusion according to claim 3 is characterized by: In step S21, the laser radar point cloud is preprocessed to obtain a point cloud sparse map, including: S21A1: Get the x-axis position, y-axis position, and z-axis position of each point in the lidar point cloud; S21A2: horizontally concatenate the rotation matrix and the translation vector to obtain a 4×3 transformation matrix, and expand the bottom of the transformation matrix to obtain a 4x4 homogeneous transformation matrix; S21A3: After transforming each point in the laser radar point cloud processed by step S21A1 through the transposed homogeneous transformation matrix, take the first three columns of the transformed matrix as the point cloud transformation matrix, and use the elements of the point cloud transformation matrix as the new 3D coordinates corresponding to each point in the laser radar point cloud; The first column of the point cloud transformation matrix corresponds to the x value of each point in the laser radar point cloud, the second column of the point cloud transformation matrix corresponds to the y value of each point in the laser radar point cloud, and the third column of the point cloud transformation matrix corresponds to the z value of each point in the laser radar point cloud; S21A4: Remove all points with negative x values in the point cloud transformation matrix, take the first two columns of the point cloud transformation matrix after removing the corresponding points as the two-dimensional image coordinate values, use the camera intrinsic parameters and distortion parameters to convert the two-dimensional image coordinate values into a 2D image coordinate matrix, and round the values of the 2D image coordinate matrix; S21A5: construct a depth map based on the z value of each point in the point cloud transformation matrix and the 2D image coordinate matrix, and normalize the depth map so that the values of the pixels contained in the depth map are in the range of 0 to 255, thereby obtaining a point cloud sparse map; The steps of preprocessing the camera image to obtain the corrected image include: S21B1: Based on the checkerboard images taken from multiple angles, the camera intrinsic parameter matrix and distortion coefficients are calculated using the Zhang Zhengyou calibration algorithm, and the mapping relationship between the camera coordinate system and the image coordinate system is established; S21B2: Based on the calibration result of step S21B1, the camera image is corrected in real time using bilinear interpolation to obtain a corrected image: ; ; ; Where r is the two-dimensional coordinate in the camera image. The Euclidean distance from the pixel point to the center of the camera image, is the two-dimensional coordinate of the rectified image, and is the radial distortion coefficient of the camera, and is the tangential distortion coefficient of the camera.
5. The road segmentation and slope estimation method based on multi-sensor fusion according to claim 3 is characterized by: The lightweight segmentation model includes an encoder-decoder structure and a multi-scale evidence collection module. The corrected image and the point cloud sparse map are input into the encoder-decoder structure for feature extraction, and the first feature map and the second feature map are obtained accordingly. The first feature map and the second feature map are input into the multi-scale evidence collection module for processing, and the binarization result is obtained after result-level fusion. The multi-scale evidence collection module includes a first branch and a second branch with the same network structure. The first branch includes a first sub-branch, a second sub-branch and a third sub-branch. The first sub-branch, the second sub-branch and the third sub-branch all contain sequentially connected convolution, upsampling layers and softplus activation layers. Among them, the convolution of the first sub-branch is 1×1 convolution, the convolution of the second sub-branch is 3×3 convolution, and the convolution of the third sub-branch is 5×5 convolution.
6. The road segmentation and slope estimation method based on multi-sensor fusion according to claim 3 is characterized by: Step S23 includes: S231: Constructing LiDAR confidence evaluation formula And the visual confidence evaluation formula : ; ; Among them, the fitting residual is the deviation between the predicted value output by the lightweight segmentation model and the true value, E is the probability map entropy value output by the encoder-decoder structure, L is the light intensity, is the maximum illumination value obtained based on prior knowledge; S232: Calculate the final road surface by the following formula, and use the final road surface as the road segmentation result: 。 7. The road segmentation and slope estimation method based on multi-sensor fusion according to claim 1, characterized in that: Step S3 specifically includes the following steps: S31: reprojecting the road segmentation result back to the point cloud to obtain a three-dimensional point cloud image of the road surface; S32: Segment the 3D point cloud image, and use the RANSIC algorithm to perform plane fitting on each section of the road surface, and calculate the relative slope angle of each section of the road surface based on the plane normal vector corresponding to each section of the road surface after fitting. : ; in, is the plane normal vector obtained after plane fitting of the i-th section of road surface, is the inverse cosine function.
8. The road segmentation and slope estimation method based on multi-sensor fusion according to claim 7, characterized in that: If the vehicle-mounted IMU provides the vehicle's own plane normal vector , the absolute slope angle of each road section is calculated by the following formula : 。 9. The road segmentation and slope estimation method based on multi-sensor fusion according to claim 7, characterized in that: In step S31, the road segmentation result is reprojected back to the point cloud by the following formula: ; ; Among them, x and y are the calculated x and y values in the camera coordinate system corresponding to the z value of the current point, u and v are the pixel coordinates in the image coordinate system corresponding to the z value of the current point, P is the camera intrinsic parameter matrix, z is the depth value of the current point, P(0, 0) and P(1, 1) are the focal lengths of the camera, and P(0, 2) and P(1, 2) are the principal point positions of the camera.
Citation Information
Patent Citations
Unstructured road state parameter estimation method and system
CN114565616A
Vehicle target detection method and system based on Leiyu semantic segmentation adaptive fusion
CN114724120A
Scene object classification method and system based on multi-view depth image
CN119169288A
Calibration device, calibration method, optical device, imaging device, projection device, measurement system, and measurement method
WO2016076400A1
Cited By
Laser radar and camera fused road segmentation method, system and device and medium
CN120279277A
Road segmentation method, system, device and medium for laser radar and camera fusion
CN120279277B
Front road gradient estimation method based on multivariate information fusion
CN120279522A
Cross-country road gradient prediction method and device based on multi-modal data fusion
CN122023530A
Road image processing method based on inverse perspective transformation and electronic equipment
CN122492431A