Road segmentation method based on LC-HRNet pixel level fusion
By adopting a pixel-level fusion method based on LC-HRNet in road segmentation, combining point cloud and image data, multi-channel RGBXYZ data is generated, and HRNet neural network processing is carried out, the problems of light and weather changes and multi-modal data fusion in the existing technology are solved, and the road segmentation effect with high accuracy and robustness is achieved.
Patent Information
- Application Number
- CN202510061483.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively deal with light and weather changes in road segmentation, and the method of a single data source has the problem of complementary advantages in multimodal data fusion.
Using a pixel-level fusion method based on LC-HRNet, dense LiDAR-Image maps are obtained through point cloud-image coordinate system conversion, used in combination with RGB images, multi-channel RGBXYZ data is generated, and inputted to the HRNet neural network for road segmentation.
The accuracy and robustness of road segmentation are improved, with a maximum F value of 93.26%, which significantly improves the accuracy and computing efficiency of multimodal data fusion.
Smart Images

Figure CN119992494A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of road segmentation, and in particular relates to a road segmentation method based on LC-HRNet pixel-level fusion. Background Art
[0002] Road segmentation refers to segmenting the road area from the data obtained by the on-board camera or LiDAR. Accurate road segmentation is a necessary condition for smart cars to generate safe paths. The road segmentation problem can be regarded as a pixel-level binary classification problem, that is, the collected environmental data (such as RGB images or 3D LiDAR point clouds) are classified, and each pixel or LiDAR point cloud data is assigned a label of road area or non-road area.
[0003] Currently, most of them have developed multiple algorithm frameworks through a single data source and deep learning technology. For methods that only use RGB images, Multi-Net provides a unified platform for multi-task processing, while RBNet uses an Encoder-Decoder structure to collect features at different scales. Fan et al. also created a new driving scene dataset. However, the performance of these image-based methods is often negatively affected when facing changes in lighting and weather.
[0004] For methods that only use sparse point cloud data, Fernandes et al. used sliding window technology and morphological post-processing to improve the results. Projection-based methods transform point clouds to BEV views or spherical front views, which are suitable for real-time systems. Thrun et al. observed the distribution of point cloud data in the vertical direction and proposed a minimum and maximum elevation map representation to simplify the road segmentation problem. Ly et al. rearranged the point cloud data into specific views for input. Gu et al. identified road areas by extracting vertical and horizontal histograms.
[0005] The forms and sizes of RGB images and 3D point cloud data are completely different. How to effectively design multimodal features to achieve complementary advantages of data is the key to improving the perception level of smart cars. Summary of the invention
[0006] Purpose of the invention: In order to overcome the above shortcomings, the purpose of the present invention is to provide a road segmentation method based on LC-HRNet pixel-level fusion, combined with high-resolution HRNet neural network and upsampling theory, through point cloud-image coordinate system conversion, to obtain the image corresponding to the dense LiDAR-Image map from the lidar point cloud, and then combine the RGB channel data to generate multi-channel RGBXYZ data; then, these data are sent to the HRNet neural network for road segmentation, which can effectively improve the accuracy and robustness of the segmentation.
[0007] Technical solution: In order to achieve the above object, the present invention provides a road segmentation method based on LC-HRNet pixel-level fusion, comprising:
[0008] S1): obtaining an RGB color image, that is, taking an RGB color image of the road by a camera;
[0009] S2): Obtain 2D data from the LiDAR point cloud, that is, obtain a large number of point clouds by scanning the surrounding environment with the LiDAR sensor, accurately reflect the 3D coordinates of the real world, and extract the point cloud corresponding to the image pixel according to the range constraints within the camera field of view. After the point cloud-image coordinate system conversion, obtain the LiDAR 2D data with the same structure as the RGB image from the LiDAR point cloud;
[0010] S3): The 2D LiDAR data obtained in S2) is upsampled to obtain a dense LiDAR-Image; then the LiDAR-Image is combined with the three channels R, G, and B in the RGB color image to construct a new data type, namely RGBXYZ multi-channel tensor data, where X, Y, and Z are the three-dimensional coordinates of the LiDAR point cloud, corresponding to the LiDAR 2D depth map, range map, and height map, respectively;
[0011] S4): Send the fused data into the HRNet neural network model for road segmentation and output the road visualization result;
[0012] S5): Result verification: The effect achieved by the above segmentation method is verified through algorithm evaluation, HRNet neural network model training experiment, and comparative experiment of multimodal data fusion and single modal data input.
[0013] The road segmentation method based on LC-HRNet pixel-level fusion described in the present invention adopts an upsampling method and a height difference method respectively. The specific process of obtaining a dense LiDAR-Image map by upsampling the LiDAR 2D data obtained in S2) in S3) is as follows:
[0014] Densify the sparse information of the lidar 2D depth map, range map and height map in S2) into pseudo camera data;
[0015] For the LiDAR 2D depth map and range map, an upsampling operation is used to improve the image with lower resolution through interpolation method and enlarge the image, as follows:
[0016] P=F(Q1,Q2,Q3,Q4) (1)
[0017] Among them, Q1, Q2, Q3 and Q4 are low-resolution image pixels, and P is the pixel point to be constructed on the high-resolution image;
[0018] The dense depth of each pixel is calculated using bilateral filtering, and the output image D p ,as follows:
[0019]
[0020] Where I is the sparse radar depth map, I q is the depth information corresponding to the lidar point cloud at the q-point position of the sparse radar depth map, N represents the spatial domain (neighborhood); W p is the normalization factor, ensuring that the sum of the weights is 1, so that the converted grayscale value range is between 0 and 255; is the distance penalty term, and its size is inversely proportional to the Euclidean distance (‖pq‖) between pixel positions p and q; is the weight of the depth of point q to point p, The value of is inversely proportional to the distance value;
[0021] Similarly, a dense range map can be obtained;
[0022] Radar height map, introduce height difference transformation operation, calculate the pixel value H at (x, y) x,y Traverse each pixel of the height image, calculate the absolute value of the offset between two positions, and obtain the height difference, so as to densify the sparse radar height map. The specific formula is as follows:
[0023]
[0024] Among them, Z x,y is the height value of the corresponding laser radar point, N x and N y is the location of the point in the neighborhood, and M is the total number of neighbors.
[0025] The road segmentation method based on LC-HRNet pixel-level fusion described in the present invention, the new data type is pixel-level data, which is composed of RGB channel data and LiDAR-Image, all channel data are located in the Cartesian coordinate system, in the X (depth)-Y (range)-Z (height) channel image, the pixel values are respectively normalized by x, y, and z coordinates;
[0026] The pixel-level data is then input into the CNN; the data size is 6×m×n, where m×n is the image size.
[0027] The road segmentation method based on LC-HRNet pixel-level fusion described in the present invention, the specific method of obtaining the point cloud label is as follows: Since the training data set of the official road data set KITTI (Karlsruhe Institute of Technology and Toyota Technological Institute, KITTI) only provides the true value of the RGB image, and does not provide the label data of the point cloud, it is necessary to additionally prepare the point cloud label. The specific method of obtaining the point cloud label is as follows:
[0028] Since the two modal data of RGB image and 3D point cloud have been aligned, the true value of the image data can be directly transferred to the corresponding radar point cloud, thereby generating the true value label of the lidar point cloud data, as follows:
[0029]
[0030] in, is the label of the i-th point cloud,
[0031] This means that the semantic label of the i-th point cloud on the projected image is the road area;
[0032] Label Image is the semantic label of the image, T LiDARtoImage ×LiDAR i is the point on the image corresponding to the i-th point cloud data.
[0033] The road segmentation method based on LC-HRNet pixel-level fusion of the present invention adopts a parallel HRNet network structure and is divided into multiple stages; each stage performs a downsampling operation before entering the next stage to reduce the scale of the feature map; then, at the end of each stage, feature maps of different resolutions are fused to retain multi-scale information.
[0034] The road segmentation method based on LC-HRNet pixel-level fusion of the present invention, the multiple stages of the HRNet network are specifically:
[0035] In stage I, the fused data is used as the network input. After passing through the residual module, two branches are connected in parallel, one of which maintains the original image resolution and the other has half the feature size.
[0036] Phase II takes the output of Phase I as input, connects multiple branches in parallel, obtains multiple feature maps with different resolutions, and then fuses the features of different resolutions through a multi-scale feature fusion module;
[0037] Similarly, in each branch of stage III and stage IV, one branch maintains the original image resolution, and the feature sizes of the remaining branches are halved in turn, and the reduced feature maps are obtained, so that feature maps of different resolutions can repeatedly obtain information from other feature maps;
[0038] Finally, each segmentation result of different resolutions is sampled to the original image resolution, and the segmentation result is predicted on the first branch.
[0039] In the road segmentation method based on LC-HRNet pixel-level fusion of the present invention, when the multi-scale feature fusion module is fused, branches with the same resolution are directly copied;
[0040] When the high-resolution branch is fused to the features of the low-resolution branch, it first goes through one or several convolutions with a step size of 2 to sample to the corresponding low-resolution feature size, and then all features are superimposed according to the corresponding channels;
[0041] When the low-resolution branch is fused to the high-resolution branch, an upsampling operation is required to make it consistent with the high-resolution feature size;
[0042] Finally, a 1×1 convolution is used to maintain a uniform number of channels, and the feature fusion method is channel superposition.
[0043] It can be seen from the above technical solution that the present invention has the following beneficial effects:
[0044] 1. The road segmentation method based on LC-HRNet pixel-level data fusion described in the present invention obtains a dense LiDAR-Image corresponding to the image from the lidar point cloud through joint calibration (one-to-one correspondence between point cloud and image modal data) and upsampling operations, and splices it with the RGB channel data of the RGB image to construct a new type of multi-channel RGBXYZ data; the RGBXYZ data contains color information and depth information, providing richer input features for the deep learning model; and a high-resolution HRNet neural network is used to extract features from the multi-channel RGBXYZ data, segment the image, and obtain the road segmentation result. The experimental results of the KITTI road dataset show that compared with the method using only a single modal data, the Max F value of this method reaches 93.26%; compared with the related multi-modal fusion road segmentation algorithm, its accuracy and robustness are significantly improved.
[0045] 2. The present invention combines the high-resolution HRNet neural network and upsampling theory, obtains a dense LiDAR-Image corresponding to the image from the LiDAR point cloud through point cloud-image coordinate system conversion, and then combines the RGB channel data to generate multi-channel RGBXYZ data; then, these data are sent to the HRNet neural network for road segmentation, which can effectively improve the accuracy and robustness of segmentation.
[0046] 3. The HRNet neural network with multi-scale feature fusion in the present invention can accurately retain the effective information in the field of view while maintaining high computational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 A block diagram of a road segmentation method based on LC-HRNet pixel-level data fusion in the present invention;
[0048] Figure 2 It is the HRNet network structure in the present invention;
[0049] Figure 3 The residual network configuration of each stage of HRNet in the present invention;
[0050] Figure 4 is the residual module of HRNet in the present invention, wherein (a) is the Bottleneck residual module, and (b) is the Basic·Block residual module;
[0051] Figure 5 It is the HRNet multi-scale feature fusion module in the present invention, wherein (a) the feature fusion module of medium and low resolution branches to high resolution branches; (b) the feature fusion module of high and low resolution branches to medium resolution branches; (c) the feature fusion module of high and medium resolution branches to low resolution branches;
[0052] Figure 6 The generation process of LiDAR-Image in the present invention;
[0053] Figure 7 Schematic diagram of pixel interpolation in the present invention;
[0054] Figure 8 This is an example of generating dense LiDAR-Image data in the present invention; wherein (a) is an RGB image; (b) is an RGB image with superimposed sparse point cloud; (c) is a sparse depth map of the point cloud within the image field of view; (d) is a dense depth map of the point cloud within the image field of view; (e) is a sparse range map of the point cloud within the image field of view; (f) is a dense range map of the point cloud within the image field of view; (g) is a sparse height map of the point cloud within the image field of view; (h) is a dense height map of the point cloud within the image field of view;
[0055] Fig. 9 An example of channel data in a Cartesian coordinate system for pixel-level data constructed in the present invention;
[0056] Fig.10 The upper middle picture shows the true value of the image, and the lower picture shows the true value label of the point cloud;
[0057] Fig.11 is the loss during the training of the neural network in the present invention;
[0058] Fig.12 is the accuracy during the training of the neural network in the present invention;
[0059] Fig.13 This is a table of evaluation (%) of road segmentation results on the KITTI-Road validation dataset in the present invention;
[0060] Fig.14 The first row in the figure is the original RGB image, and the second to last rows are the road segmentation results using the RGB image, LiDAR-Image data, and multi-modal data as input respectively;
[0061] Fig.15 is the road area segmentation result on the KITTI road test set in the present invention (%);
[0062] Fig.16 It is the Precision-Recall curve (PR curve) of the present invention, wherein (a) is the PR curve of UM environment data; (b) is the PR curve of UMM environment data; (c) is the PR curve of UU environment data; (d) is the PR curve of URBAN environment data;
[0063] Fig.17 The visualization result of the road segmentation on the KITTI test set in the present invention under the camera's forward viewing angle is shown in red, where the area that is not correctly divided into a road region, the area that is incorrectly divided into a road region, and the area that is correctly segmented into a road region are shown in blue, and the area that is correctly segmented into a road region are shown in green.
[0064] Fig.18 The visualization result of road segmentation from a bird's eye view in the present invention;
[0065] Fig.19 is the online evaluation result of UM (bird's eye view) in the present invention (%);
[0066] Fig. 20 is the UMM (bird's eye view) online evaluation result (%) in the present invention;
[0067] Fig.21It is the UU (bird's eye view) online evaluation result (%) in the present invention. DETAILED DESCRIPTION
[0068] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments.
[0069] Example
[0070] like Figures 1 to 21 A road segmentation method based on LC-HRNet pixel-level fusion is shown, comprising:
[0071] S1): obtaining an RGB color image, that is, taking an RGB color image of the road by a camera;
[0072] S2): Obtain 2D data from the LiDAR point cloud, that is, obtain a large number of point clouds by scanning the surrounding environment with the LiDAR sensor, accurately reflect the 3D coordinates of the real world, and extract the point cloud corresponding to the image pixel according to the range constraints within the camera field of view. After the point cloud-image coordinate system conversion, obtain the LiDAR 2D data with the same structure as the RGB image from the LiDAR point cloud;
[0073] S3): The 2D LiDAR data obtained in S2) is subjected to upsampling operation to obtain a dense LiDAR-Image image; the LiDAR-Image image is then combined with the three channels R, G, and B in the RGB color image and fused through a multi-scale feature fusion module to construct a new data type, namely RGBXYZ multi-channel tensor data, where X, Y, and Z are the three-dimensional coordinates of the LiDAR point cloud, corresponding to the LiDAR 2D depth map, range map, and height map, respectively;
[0074] S4): Send the fused data into the HRNet neural network model for road segmentation and output the road visualization result;
[0075] S5): Result verification: The effect achieved by the above segmentation method is verified through algorithm evaluation, HRNet neural network model training experiment, and comparative experiment of multimodal data fusion and single modal data input.
[0076] The road segmentation method based on LC-HRNet pixel-level fusion described in the present invention adopts an upsampling method and a height difference method respectively. The specific process of obtaining a dense LiDAR-Image map by upsampling the LiDAR 2D data obtained in S2) in S3) is as follows:
[0077] Densify the sparse information of the lidar 2D depth map, range map and height map in S2) into pseudo camera data;
[0078] For the LiDAR 2D depth map and range map, an upsampling operation is used to improve the image with lower resolution through interpolation method and enlarge the image, as follows:
[0079] P=F(Q1,Q2,Q3,Q4) (1)
[0080] Among them, Q1, Q2, Q3 and Q4 are low-resolution image pixels, and P is the pixel point to be constructed on the high-resolution image;
[0081] The dense depth of each pixel is calculated using bilateral filtering, and the output image D p ,as follows:
[0082]
[0083] Where I is the sparse radar depth map, I q is the depth information corresponding to the lidar point cloud at the q-point position of the sparse radar depth map, N represents the spatial domain (neighborhood); W p is the normalization factor, ensuring that the sum of the weights is 1, so that the converted grayscale value range is between 0 and 255; is the distance penalty term, and its size is inversely proportional to the Euclidean distance (||pq||) between pixel positions p and q; is the weight of the depth of point q to point p, The value of is inversely proportional to the distance value;
[0084] Similarly, a dense range map can be obtained;
[0085] Radar height map, introduce height difference transformation operation, calculate the pixel value H at (x, y) x,y Traverse each pixel of the height image, calculate the absolute value of the offset between two positions, and obtain the height difference, so as to densify the sparse radar height map. The specific formula is as follows:
[0086]
[0087] Among them, Z x,y is the height value of the corresponding laser radar point, N x and N y is the location of the point in the neighborhood, and M is the total number of neighbors.
[0088] The road segmentation method based on LC-HRNet pixel-level fusion described in the present invention, the new data type is pixel-level data, which is composed of RGB channel data and LiDAR-Image, all channel data are located in the Cartesian coordinate system, in the X (depth)-Y (range)-Z (height) channel image, the pixel values are respectively normalized by x, y, and z coordinates;
[0089] The pixel-level data is then input into the CNN; the data size is 6×m×n, where m×n is the image size.
[0090] The road segmentation method based on LC-HRNet pixel-level fusion described in the present invention, the specific method of obtaining the point cloud label is as follows: Since the training data set of the official road data set KITTI (Karlsruhe Institute of Technology and Toyota Technological Institute, KITTI) only provides the true value of the RGB image, and does not provide the label data of the point cloud, it is necessary to additionally prepare the point cloud label. The specific method of obtaining the point cloud label is as follows:
[0091] Since the two modal data of RGB image and 3D point cloud have been aligned, the true value of the image data can be directly transferred to the corresponding radar point cloud, thereby generating the true value label of the lidar point cloud data, as follows:
[0092]
[0093] in, is the label of the i-th point cloud,
[0094] This means that the semantic label of the i-th point cloud on the projected image is the road area;
[0095] Label Image is the semantic label of the image, T LiDARtoImage ×LiDAR i is the point on the image corresponding to the i-th point cloud data.
[0096] Figure 8 (a) is an RGB image, Figure 8 (b) is the effect of converting the sparse LiDAR point cloud to an image (i.e., the RGB image with the sparse point cloud superimposed). Figure 8 (c), (e) and (g) are sparse LiDAR-Images obtained by converting to the image plane. Figure 8 (d), (f) and (h) are dense LiDAR-Images obtained by upsampling. Figure 8 (c)-(h) show that the pixel value changes with the detection value in the LiDAR coordinate system. When the detection value of the point cloud in the LiDAR coordinate system is larger, the pixel value in the corresponding image will become larger or brighter; and the density of the point cloud decreases as the distance between the sensor and the detection target increases.
[0097] In addition, from the sparse radar data Figure 8(c), (e) and (g) can be directly observed that when the difference between roads and non-roads is not high, directly using this data will reduce the ability of the road detection model to identify the road area. The upsampling conversion method is used to obtain dense depth maps and range maps through interpolation methods. For dense height images, the height difference conversion method is used to calculate the height change to obtain the height difference image, such as Figure 8 As shown in (h), the road features in the radar data are retained; when there are buildings or other objects, there is an obvious height difference between the road and other objects, so the height map is very helpful for distinguishing the road area.
[0098] In the road segmentation method based on LC-HRNet pixel-level fusion described in this embodiment, the HRNet network structure adopts a parallel method and is divided into multiple stages; each stage performs a downsampling operation before entering the next stage to reduce the scale of the feature map; then, at the end of each stage, feature maps of different resolutions are fused to retain multi-scale information.
[0099] The road segmentation method based on LC-HRNet pixel-level fusion described in this embodiment, the multiple stages of the HRNet network are specifically:
[0100] In stage I, the fused data is used as the network input. After passing through the residual module, two branches are connected in parallel, one of which maintains the original image resolution and the other has half the feature size.
[0101] Phase II takes the output of Phase I as input, connects multiple branches in parallel, obtains multiple feature maps with different resolutions, and then fuses the features of different resolutions through a multi-scale feature fusion module;
[0102] Similarly, in each branch of stage III and stage IV, one branch maintains the original image resolution, and the feature sizes of the remaining branches are halved in turn, and the reduced feature maps are obtained, so that feature maps of different resolutions can repeatedly obtain information from other feature maps;
[0103] Finally, each segmentation result of different resolutions is sampled to the original image resolution, and the segmentation result is predicted on the first branch.
[0104] In the road segmentation method based on LC-HRNet pixel-level fusion described in this embodiment, when the multi-scale feature fusion module performs fusion, branches with the same resolution are directly copied;
[0105] When the high-resolution branch is fused to the features of the low-resolution branch, it first goes through one or several convolutions with a step size of 2 to sample to the corresponding low-resolution feature size, and then all features are superimposed according to the corresponding channels;
[0106] When the low-resolution branch is fused to the high-resolution branch, an upsampling operation is required to make it consistent with the high-resolution feature size;
[0107] Finally, a 1×1 convolution is used to maintain a uniform number of channels, and the feature fusion method is channel superposition.
[0108] In the road segmentation method based on LC-HRNet pixel-level fusion described in this embodiment, the new data type is pixel-level data, which is composed of RGB channel data and LiDAR-Image. All channel data are located in the Cartesian coordinate system. In the X (depth)-Y (range)-Z (height) channel image, the pixel values are respectively normalized by x, y, and z coordinates;
[0109] The pixel-level data is then input into the CNN; the data size is 6×m×n, where m×n is the image size.
[0110] The road segmentation method based on LC-HRNet pixel-level fusion described in this embodiment, the S4) result verification, through algorithm evaluation, model training experiments and multi-modal data fusion and single-modal data input comparison experiments to verify the effect achieved by the above segmentation method;
[0111] The algorithm evaluation includes two methods: qualitative evaluation and quantitative evaluation.
[0112] Qualitative evaluation relies on the experience of observers and is mainly based on observation and analysis, which contains subjective elements. It lacks comprehensiveness and objectivity, cannot accurately measure the pros and cons of algorithms, and is usually regarded as an auxiliary means.
[0113] Quantitative evaluation uses mathematical statistics to be scientific and precise, and uses digital standards to measure the differences between different algorithms.
[0114] KITTI provides an evaluation script in the bird's-eye view space for road segmentation. The evaluation indicators mainly include the maximum F1-measure (Max F), precision (PRE), recall (REC), average precision (AP), false positive rate (FPR) and false negative rate (FNR). These indicators involve:
[0115] TP (True Positive): The number of positive samples detected as correct, indicating the road pixels that are correctly classified;
[0116] TN (True Negative): The number of negative samples detected as correct, indicating non-road pixels that are correctly classified;
[0117] FP (False Positive): The number of positive samples detected as errors, indicating non-road pixels that were misclassified;
[0118] FN (False Negative): The number of negative samples detected as errors, indicating road pixels that are misclassified.
[0119] Precision: Evaluates the ratio of “the number of road pixels correctly detected by the algorithm” to “the number of road pixels detected in the true value”:
[0120]
[0121] Recall: Evaluates the ratio of "the number of correctly detected road pixels in the algorithm's detection results" to "the total number of real road pixels in the algorithm's results":
[0122]
[0123] The maximum F1 measurement value is a comprehensive indicator, which is a weighted harmonic average of precision and recall to balance the impact of precision and recall. It is often used for pixel-level evaluation. The F1 measurement value represents the measurement result of β = 1.
[0124]
[0125] Average precision is derived from the PR curve.
[0126] Precision and recall provide different insights into the performance of the method: lower precision means that many background pixels are classified as road, while lower recall means that the road surface cannot be detected.
[0127] False positive rate FPR: gives the ratio of negative samples predicted as positive to all negative samples.
[0128] False Negative Rate FNR: gives the ratio of positive samples predicted as negative to all positive samples.
[0129]
[0130] In addition, Accuracy indicates the proportion of detected road pixels in all pixels:
[0131]
[0132] The prediction results on the road test set are submitted to the KITTI official evaluation website for evaluation, and ranked according to Max F from a bird's-eye view to compare different road segmentation algorithms.
[0133] The specific process of the model training experiment is as follows:
[0134] Experimental comparison based on fused RGBXYZ data and traditional RGB image data. The loss of the network during model training is as follows Fig.11 shown.
[0135] In the early stage of model training, the training losses based on the fusion of RGBXYZ data and RGB images were both around 2.4. When the number of training times reached 800, the training loss based on RGBXYZ data dropped rapidly to around 0.2, while the training loss based on RGB images only dropped to around 0.5.
[0136] When the training iteration reached 1250 times, the network training based on RGBXYZ data had basically converged; while the network based on RGB images basically converged after 1750 iterations.
[0137] It can be seen that in model training, the loss curve when using RGBXYZ data drops faster than when using only RGB data. This is because RGBXYZ data provides richer features, and the XYZ channel data is converted from the lidar point cloud. The neural network using RGBXYZ data converges very quickly in 600 iterations, with lower losses when stable. Throughout the training process, it is always better than the model using only RGB image data.
[0138] Fig.12 The figure shows the average accuracy of the HRNet model when training the model using RGBXYZ data and RGB image data. After 1800 iterations, the accuracy based on RGBXYZ data can reach up to 92%, compared to only about 84% using RGB images.
[0139] In summary, using RGBXYZ data makes the model converge faster than traditional RGB images. In other words, using the XYZ channel data converted by the lidar sensor can improve the accuracy by 8%.
[0140] Comparative experiment of multimodal data fusion and single modal data input:
[0141] The KITTI road dataset is used to verify the role of multimodal data fusion. The road segmentation results using multimodal data and single-modal data as network input are compared on the KITTI validation set, and the role of multimodal data fusion strategy in improving road segmentation results is analyzed.
[0142] Fig.13It represents the road segmentation results based on multimodal data and single-modal data as network input. Among them, "RGB-based" means that the network input only has three-channel data of RGB image; "LiDAR-Image-based" means that the network input only has dense XYZ channels (depth map, range map and height map); "RGBXYZ-based" means that multimodal fusion data is used as network input.
[0143] The quantitative segmentation results show that the road segmentation results based on the point cloud method and the multimodal fusion method are similar. In terms of MaxF values, they are 3.23% and 4.29% higher than the road segmentation method based on RGB images, respectively. The reason is that the image data is affected by different lighting conditions. This also verifies that multimodal data fusion has a significant improvement effect on road segmentation results.
[0144] Fig.14 It shows the visualization results of using different modal data as network input. The first row is the RGB original image, the second row is the road segmentation result based on the RGB image, the third row is the road segmentation result based on dense LiDAR-Image data, and the last row is the road segmentation result based on multi-modal data fusion. Fig.14 It is concluded that after using multimodal fusion data, the number of incorrect segmentations is significantly reduced.
[0145] Fig.15 The segmentation results of various environment types on the KITTI road test dataset. From Figure 15, it can be concluded that the performance of this method in the UM and UMM environment datasets is better than that in the UU environment dataset. This is because the road areas in the UM and UMM environments are flatter and have clear road dividing lines.
[0146] Fig.16 The figure is the Precision-Recall (PR) curve. The Precision-Recall curve represents the recall rate (horizontal axis) and precision (vertical axis) corresponding to different IoU and confidence thresholds, reflecting the algorithm performance from two perspectives. Ideally, the curve shows a trend in the upper right corner. In summary, the Precision-Recall curve based on the pixel-level fusion method in the UMM road environment dataset is the best.
[0147] In the road segmentation task, the visualization results on the KITTI test set are as follows: Fig.17 and Fig.18As shown in the figure, the camera's front view and bird's-eye view are respectively used, and different colors are used to represent the difference between the road prediction result and the road true value. Among them, the red area represents the false negative area, the blue area represents the false positive area, and the green area represents the true positive area. It can be seen from the visualization results that the multimodal data fusion method can better complete the road segmentation task in a complex environment. This is because the model input data contains high-precision information of the point cloud, so it will not cause obvious wrong segmentation results in the shadow part of the road surface. The wrong segmentation basically occurs at the edge of the road and the boundary between the road and obstacles such as pedestrians and vehicles.
[0148] Comparison with other methods
[0149] Considering the existing multimodal data fusion methods, Figure 19 to Figure 20 The quantitative comparison results of the method in this chapter and other data fusion methods are respectively for road test sets in different environments in KITTI, including RES3D+VELO, FusedCRF, Hybrid CRF and Han algorithm. The method in this chapter performs best in UM and UMM datasets, and performs slightly worse than Hybrid CRF in UU environment dataset, but is better than the methods in the list overall. Bold indicates the best performance.
[0150] from Figures 19 to 21 It is concluded that the pixel-level fusion method proposed in this chapter is superior to Fig.19 To table Fig.21 The performance of the four methods listed in . In addition, there are some problems in this model that need to be improved. The feature points collected by the LiDAR point cloud in the distance are lost, which leads to poor segmentation ability in the distant area; and when the sidewalk and road area are very close, there will be segmentation errors, that is, the sidewalk is regarded as the road area, which makes it difficult to accurately segment the road.
[0151] The above description is only a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements without departing from the principle of the present invention, and these improvements should also be regarded as within the protection scope of the present invention.
Claims
1. A road segmentation method based on LC-HRNet pixel-level fusion, characterized by: include: S1): obtaining an RGB color image, that is, taking an RGB color image of the road by a camera; S2): Obtain 2D data from the LiDAR point cloud, that is, obtain a large number of point clouds by scanning the surrounding environment with the LiDAR sensor, accurately reflect the 3D coordinates of the real world, and extract the point cloud corresponding to the image pixel according to the range constraints within the camera field of view. After the point cloud-image coordinate system conversion, obtain the LiDAR 2D data with the same structure as the RGB image from the LiDAR point cloud; S3): The 2D LiDAR data obtained in S2) is upsampled to obtain a dense LiDAR-Image; then the LiDAR-Image is combined with the three channels R, G, and B in the RGB color image to construct a new data type, namely RGBXYZ multi-channel tensor data, where X, Y, and Z are the three-dimensional coordinates of the LiDAR point cloud, corresponding to the LiDAR 2D depth map, range map, and height map, respectively; S4): Send the fused data into the HRNet neural network model for road segmentation and output the road visualization result; S5): Result verification: The effect achieved by the above segmentation method is verified through algorithm evaluation, HRNet neural network model training experiment, and comparative experiment of multimodal data fusion and single modal data input.
2. The road segmentation method based on LC-HRNet pixel-level fusion according to claim 1, characterized in that: The specific process of obtaining a dense LiDAR-Image image by upsampling the LiDAR 2D data obtained in S2) in S3) is as follows: Densify the sparse information of the lidar 2D depth map, range map and height map in S2) into pseudo camera data; For the radar 2D depth map and range map, an upsampling operation is used to improve the image with smaller resolution through interpolation method and enlarge the image, as follows: P=F(Q1,Q2,Q3,Q4) (1) Among them, Q1, Q2, Q3 and Q4 are low-resolution image pixels, and P is the pixel point to be constructed on the high-resolution image; The dense depth of each pixel is calculated using bilateral filtering, and the output image D p ,as follows: Where I is the sparse radar depth map, I q is the depth information corresponding to the lidar point cloud at the q-point position of the sparse radar depth map, N represents the spatial domain (neighborhood); W p is the normalization factor, ensuring that the sum of the weights is 1, so that the converted grayscale value range is between 0 and 255; is the distance penalty term, and its size is inversely proportional to the Euclidean distance (||pq||) between pixel positions p and q; is the weight of the depth of point q to point p, The value of is inversely proportional to the distance value; Similarly, a dense range map can be obtained; Radar height map, introduce height difference transformation operation, calculate the pixel value H at (x, y) x,y Traverse each pixel of the height image, calculate the absolute value of the offset between two positions, and obtain the height difference, so as to densify the sparse radar height map. The specific formula is as follows: Among them, Z x,y is the height value of the corresponding laser radar point, N x and N y is the location of the point in the neighborhood, and M is the total number of neighbors.
3. The road segmentation method based on LC-HRNet pixel-level fusion according to claim 1, characterized in that: The new data type is pixel-level data, which consists of RGB channel data and LiDAR-Image. All channel data are located in the Cartesian coordinate system. In the X (depth)-Y (range)-Z (height) channel image, the pixel values are normalized by x, y, and z coordinates respectively. The pixel-level data is then input into the CNN; the data size is 6×m×n, where m×n is the image size.
4. The road segmentation method based on LC-HRNet pixel-level fusion according to claim 2 is characterized in that: Since the training dataset of the official road dataset KITTI (Karlsruhe Institute of Technology and Toyota Technological Institute, KITTI) only provides the true value of the RGB image and does not provide the label data of the point cloud, it is necessary to make additional point cloud labels. The specific method of obtaining point cloud labels is as follows: Since the two modal data of RGB image and 3D point cloud have been aligned, the true value of the image data can be directly transferred to the corresponding radar point cloud, thereby generating the true value label of the lidar point cloud data, as follows: in, is the label of the i-th point cloud, This means that the semantic label of the i-th point cloud on the projected image is the road area; Label Image is the semantic label of the image, T LiDARtoImage ×LiDAR i is the point on the image corresponding to the i-th point cloud data.
5. The road segmentation method based on LC-HRNet pixel-level fusion according to claim 1, characterized in that: The HRNet neural network adopts a parallel approach and is divided into multiple stages; each stage performs a downsampling operation before entering the next stage to reduce the scale of the feature map; Then, feature maps of different resolutions are fused at the end of each stage to preserve multi-scale information.
6. The road segmentation method based on LC-HRNet pixel-level fusion according to claim 3 is characterized in that: The multiple stages of the HRNet neural network are specifically: In stage I, the fused data is used as the network input. After passing through the residual module, two branches are connected in parallel, one of which maintains the original image resolution and the other has half the feature size. Phase II takes the output of Phase I as input, connects multiple branches in parallel, obtains multiple feature maps with different resolutions, and then fuses the features of different resolutions through a multi-scale feature fusion module; Similarly, in each branch of stage III and stage IV, one branch maintains the original image resolution, and the feature sizes of the remaining branches are halved in turn, and the reduced feature maps are obtained, so that feature maps of different resolutions can repeatedly obtain information from other feature maps; Finally, each segmentation result of different resolutions is sampled to the original image resolution, and the segmentation result is predicted on the first branch.
7. The road segmentation method based on LC-HRNet pixel-level fusion according to claim 1, characterized in that: When the multi-scale feature fusion module performs fusion, branches with the same resolution are directly copied; When the high-resolution branch is fused to the features of the low-resolution branch, it first goes through one or several convolutions with a step size of 2 to sample to the corresponding low-resolution feature size, and then all features are superimposed according to the corresponding channels; When the low-resolution branch is fused to the high-resolution branch, an upsampling operation is required to make it consistent with the high-resolution feature size; Finally, a 1×1 convolution is used to maintain a uniform number of channels, and the feature fusion method is channel superposition.
Citation Information
Cited By
Laser radar and camera fused road segmentation method, system and device and medium
CN120279277A