Laser radar and camera fused road segmentation method, system and device and medium
Through the road segmentation method of integrating lidar and camera, the segmentation problem of a single sensor in complex scenarios is solved, and the stable segmentation performance and high-precision segmentation effect under different lighting conditions are achieved, which improves the robustness and computing efficiency of the system.
Patent Information
- Application Number
- CN202510779393.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-12
AI Technical Summary
Existing single sensor methods are difficult to achieve high-precision road segmentation in complex scenarios, especially in the case of light changes or occlusion, and cannot fully perceive the environment.
The road segmentation method of lidar and camera is adopted to construct the road image data set, preprocess and synchronous processing are performed, road depth maps are generated and adjusted and fused, and the optimized DeepLabv3+ is used for segmentation, combining CBAM to introduce an attention mechanism in the encoder and decoder.
Maintain stable segmentation performance in complex scenarios, reduce occlusion impact, improve segmentation accuracy and robustness, enhance the model's understanding of complex scenarios, and improve edge accuracy and computing efficiency.
Smart Images

Figure CN120279277A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and specifically relates to a road segmentation method, system, device and medium for the fusion of lidar and camera. Background Art
[0002] Road segmentation is one of the key technologies in the autonomous driving system, aiming to identify the drivable area of the vehicle. The existing road segmentation methods are mainly divided into two categories: camera-based and lidar-based. Camera-based methods rely on image texture information, but perform poorly under light changes or occlusion; lidar-based methods have high accuracy, but are costly and data-sparse. A single sensor can only provide partial information. For example, lidar provides depth information and the camera provides color information. It cannot comprehensively perceive the environment, and its performance will decline to varying degrees in bad weather and complex scenarios, making it difficult to achieve high-precision object detection and segmentation. To overcome the deficiencies of a single sensor, multi-sensor fusion technology has become the mainstream solution. By combining the advantages of different sensors, the defects of a single sensor can be made up for, and the robustness and accuracy of the system can be improved. For example, the fusion of lidar and camera can effectively make up for these defects: (1) The camera performs excellently under good lighting conditions, while lidar is not affected by lighting. The fusion of the two can maintain stable segmentation performance under different lighting conditions; (2) The fusion technology combines the point cloud data of lidar and the image information of the camera, reduces the occlusion effect between objects, and improves the integrity of segmentation; (3) The data provided by the two sensors verify each other, reduce the influence of single-sensor data errors on the segmentation result, and improve the accuracy rate.
[0003] In summary, the existing single-sensor methods are difficult to handle complex scenarios. Therefore, a road segmentation method that can combine the advantages of multiple sensors is needed. Summary of the Invention
[0004] The present invention solves the problem that the existing single-sensor methods are difficult to handle complex scenarios, and a road segmentation method that can combine the advantages of multiple sensors is needed.
[0005] The road segmentation method for the fusion of lidar and camera according to the present invention includes the following steps: Step S1, constructing a road image data set, where the road image data set includes a road RGB image obtained by camera photographing and a road point cloud image obtained by lidar; Step S2, respectively preprocessing the road RGB image and the road point cloud image to obtain a road grayscale image and a denoised road point cloud image respectively, and synchronously processing the road grayscale image and the denoised road point cloud image; Step S3: Obtain a road depth map based on the synchronized road grayscale image and the synchronized and denoised road point cloud image, adjust the road depth map, and fuse the adjusted road depth map with the road RGB image to obtain a road feature image; Step S4: Input the road feature image into the optimized DeepLabv3+ for segmentation, and then complete the segmentation of the road image.
[0006] Further, in an embodiment of the present invention, in step S2, the synchronization processing includes spatial synchronization and temporal synchronization; The spatial synchronization is specifically as follows: Synchronize the coordinate system of the camera and the coordinate system of the lidar, and then complete the spatial synchronization; The temporal synchronization is specifically as follows: The camera and the lidar respectively record a timestamp. Based on the downward compatibility principle, synchronize the road grayscale image and the denoised road point cloud image through the timestamp, and then complete the temporal synchronization.
[0007] Further, in an embodiment of the present invention, in step S3, the adjustment of the road depth map and the fusion of the adjusted road depth map with the road RGB image to obtain a road feature image are specifically as follows: Complete the complementation of the road depth map. After the complemented road depth map is successively subjected to operations such as cropping, normalization, bilinear interpolation, convolution, max pooling, and average pooling, and then adjusted through normalization, fuse it with the road RGB image through dot multiplication to obtain a road feature image.
[0008] Further, in an embodiment of the present invention, the complementation of the road depth map is specifically as follows: Retain the high-reflection intensity information of the road depth map, and complete the complementation of the low-reflection intensity information of the road depth map through interpolation.
[0009] Further, in an embodiment of the present invention, in step S4, the optimized DeepLabv3+ is specifically as follows: Introduce CBAM into the encoder and decoder of DeepLabv3+ respectively. In the decoder of DeepLabv3+, after the road feature image is successively processed by a convolutional neural network algorithm and convolution, perform dot multiplication with the road depth map adjusted by attention weights.
[0010] The road segmentation system for fusing lidar and camera according to the present invention includes the following modules: Building module, which constructs a road image dataset. The road image dataset includes road RGB images obtained by camera photography and road point cloud images obtained by lidar. Synchronization module, which preprocesses the road RGB image and the road point cloud image respectively to obtain a road grayscale image and a denoised road point cloud image, and performs synchronization processing on the road grayscale image and the denoised road point cloud image. Fusion module, which obtains a road depth map based on the synchronized road grayscale image and the synchronized denoised road point cloud image, adjusts the road depth map, and fuses the adjusted road depth map and the road RGB image to obtain a road feature image. Segmentation module, which inputs the road feature image into the optimized DeepLabv3+ for segmentation, thus completing the segmentation of the road image.
[0011] An electronic device according to the present invention includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The memory is used to store a computer program. The processor is used to implement the road segmentation method for lidar and camera fusion according to any one of the above methods when executing the program stored on the memory.
[0012] A computer-readable storage medium according to the present invention stores a computer program therein. When the computer program is executed by a processor, it implements the road segmentation method for lidar and camera fusion according to any one of the above methods.
[0013] The present invention solves the problem that the existing single-sensor method is difficult to handle complex scenarios and a road segmentation method that can combine the advantages of multiple sensors is needed. The specific beneficial effects include: 1. For the road segmentation method for lidar and camera fusion according to the present invention, in the prior art, the single-sensor method is difficult to handle complex scenarios and a road segmentation method that can combine the advantages of multiple sensors is needed. By effectively fusing the road RGB image obtained by camera photography and the road point cloud image obtained by lidar, the present invention can maintain stable performance in complex scenarios, reduce the occlusion between targets, and avoid the influence of single-sensor errors on the segmentation result. 2. The road segmentation method for the fusion of lidar and camera according to the present invention can avoid the waiting time for unarrived data by synchronously processing the data collected by the lidar and the denoised data collected by the camera, process the data collected by the lidar and the camera in parallel, reduce the calculation delay, make full use of the computing resources, and improve the overall efficiency of the system. This processing method can also continue to use the data of another sensor for processing when a certain sensor fails or the data is interfered, or adapt to these changes by dynamically adjusting the weights, improving the fault tolerance and reliability of the system; 3. The road segmentation method for the fusion of lidar and camera according to the present invention can help the model better handle the occlusion problem and distinguish occluded objects by using the road feature image as the input of the segmentation model. The addition of the lidar depth information can reduce the influence of light changes on the segmentation result, improve the segmentation performance of the segmentation model in low-light or backlight environments, enhance the model's understanding ability of complex scenes, and improve the robustness and accuracy at the same time; 4. The road segmentation method for the fusion of lidar and camera according to the present invention can help the model better capture the boundary information of the object, improve the edge accuracy of the segmentation result, reduce the calculation of irrelevant regions, and improve the calculation efficiency by adding CBAM to both the encoder and decoder in Deeplabv3+. Brief Description of the Drawings
[0014] The above and / or additional aspects and advantages of the present invention will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where: Figure 1 It is the time synchronization diagram described in Embodiment 2; Figure 2 It is the optimized DeepLabv3+ diagram with the input of the road feature image described in Embodiment 3. Detailed Embodiments
[0015] The following will clearly and completely describe various embodiments of the present invention in conjunction with the drawings. The embodiments described by referring to the drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.
[0016] Embodiment 1. The road segmentation method for the fusion of lidar and camera described in this embodiment includes the following steps: Step S1, construct a road image dataset, where the road image dataset includes the road RGB image obtained by camera photographing and the road point cloud image obtained by lidar; Step S2, preprocess the road RGB image and the road point cloud image respectively to obtain the road grayscale image and the denoised road point cloud image, and synchronously process the road grayscale image and the denoised road point cloud image; Step S3: Obtain a road depth map based on the synchronized processed road grayscale image and the synchronized processed denoised road point cloud image, adjust the road depth map, and fuse the adjusted road depth map with the road RGB image to obtain a road feature image; Step S4: Input the road feature image into the optimized DeepLabv3+ for segmentation, and then complete the segmentation of the road image.
[0017] In the prior art, a single sensor method is difficult to handle complex scenarios, so a road segmentation method that can combine the advantages of multiple sensors is needed.
[0018] To solve the above technical problems, this embodiment proposes a road segmentation method based on the fusion of lidar and camera, including the following steps: Step S1: Construct a road image dataset, where the road image dataset includes a road RGB (red, green, blue) image obtained by camera photographing and a road point cloud image obtained by lidar, and the road image dataset is the KITTI dataset (a computer vision algorithm evaluation dataset in the autonomous driving scenario); Step S2: Preprocess the road RGB image and the road point cloud image respectively. Specifically: The preprocessing refers to denoising the road RGB image and the road point cloud image respectively. For the denoising of the road point cloud image, Gaussian filtering and downsampling are used respectively. Gaussian filtering smooths the road point cloud image and reduces the influence of noise. Downsampling uses voxel grid filtering: the road point cloud image is divided into regular voxel grids, and the points in each grid are averaged or randomly sampled, so as to reduce the number of the road point cloud image and can quickly process the road point cloud image. For the road RGB image, denoising and grayscale processing are performed. Median filtering technology is used to remove the noise in the road RGB image, and the size of the road RGB image is adjusted to adapt to the input of the subsequent network. The road RGB image is converted into a noise-free road grayscale image by the weighted average method. The common formula of the weighted average method is: ; In the formula, is the road grayscale image, , and are the red pixel, green pixel, and blue pixel in the road RGB image respectively.
[0019] Synchronize the road grayscale image and the denoised road point cloud image.
[0020] Step S3: Filter out the coordinates in the denoised road point cloud image For the points, only the points located in front of the camera are retained to prevent incorrect projection. Then, the internal parameters of the camera are used to project the filtered and denoised road point cloud image onto the image plane to obtain the pixel positions of each point in the image. .
[0021] The mathematical expression is: ; ; In the formula, and are the different focal lengths of the camera respectively, is the coordinate of the principal point of the image, is the coordinate point after transformation (in the camera coordinate system).
[0022] A two-dimensional matrix with the same size as the image is created, and the initial values are all 0. The depth value corresponding to each projected pixel position is written into the matrix with an initial value of 0, and the road depth map can be obtained. Then, the road depth map is adjusted, and the adjusted road depth map and the road RGB image are fused to obtain the road feature image.
[0023] Step S4: Input the road feature image into the optimized DeepLabv3+ (deep learning model) for segmentation. The obtained segmentation result is optimized for the edge and the isolated regions or holes are removed through morphological operations and noise removal operations, etc., to obtain the final segmentation result, and then the segmentation of the road image is completed.
[0024] Therefore, in this embodiment, by fusing the road RGB image obtained by camera photographing and the road point cloud image obtained by lidar, stable performance can be maintained under different lighting conditions, the occlusion between targets can be reduced, and the influence of single sensor errors on the segmentation result can be avoided.
[0025] Embodiment 2: This embodiment further limits the road segmentation method of lidar and camera fusion described in Embodiment 1. In step S2, the synchronization processing includes spatial synchronization and time synchronization; The spatial synchronization is specifically: Synchronize the coordinate system of the camera and the coordinate system of the lidar, and then the spatial synchronization is completed; The time synchronization is specifically: The camera and the lidar respectively record a timestamp. Based on the downward compatibility principle, the road grayscale image and the denoised road point cloud image are synchronized through the timestamp, and then the time synchronization is completed.
[0026] In this embodiment, the denoised road point cloud image and the road grayscale image are generated in the lidar coordinate system and the camera coordinate system respectively. To achieve the fusion of lidar and camera data, it is necessary to unify the coordinate systems and timestamps of the two. Therefore, in this embodiment, the road grayscale image and the denoised road point cloud image are respectively subjected to spatial synchronization and time synchronization.
[0027] The specific method for separately performing spatial synchronization on the road grayscale image and the denoised road point cloud image is as follows: The lidar and the camera are calibrated using the checkerboard method, and the internal parameters of the camera and the external parameters between the lidar can be obtained: namely, the rotation matrix and the translation vector. The process of converting lidar data from the lidar coordinate system to the camera coordinate system is as follows: converting the three-dimensional lidar points from the three-dimensional lidar coordinate system to the three-dimensional camera coordinate system to obtain the corresponding three-dimensional camera points , and the specific conversion process is as follows: ; In the formula, is the conversion matrix from the three-dimensional lidar coordinate system to the three-dimensional camera coordinate system, is the rotation optimization matrix.
[0028] According to the internal parameter information of the camera, the three-dimensional camera points are converted from the three-dimensional camera coordinate system to the two-dimensional image coordinate system to obtain the corresponding two-dimensional image points , and the specific conversion process is as follows: ; In the formula, is the conversion matrix from the three-dimensional camera coordinate system to the two-dimensional image coordinate system. The horizontal viewing angle of the camera is about 40 to 50 degrees. Only the two-dimensional points falling within the camera viewing angle range will be retained, that is, the two-dimensional image points need to satisfy 0 < <= , and 0 < <= , and are the width and height of the camera image respectively, is the horizontal coordinate conforming to the width of the camera image, is the vertical coordinate conforming to the height of the camera image, is the vertical coordinate in the three-dimensional camera coordinate system.
[0029] According to the above conversion, any three-dimensional lidar point can be converted to the corresponding two-dimensional image point . The final conversion formula is: ; In this way, the spatial synchronization between the lidar and the camera is completed. The time synchronization of the road grayscale image and the denoised road point cloud image is specifically as follows: The data collected is synchronized by means of software synchronization. After the lidar and the camera collect data, the system records a timestamp, and the data collected by the two lidars and the camera is aligned through the timestamp. The implementation means of software synchronization is simple, does not require a hardware circuit, and is easy to build on platforms such as MATLAB (commercial mathematical software) and Python (computer programming language).
[0030] For example, if the sampling frame frequency of the lidar used is 10 frames per second and the frame rate of the video collected by the camera is 20 frames per second. Then, according to the downward compatibility of the lower sampling frequency, and based on the characteristics of the sampling frequencies of the lidar and the camera, one frame of data of the lidar at intervals and two frames of data of the camera at intervals are selected as valid data and recorded, that is, the data at time nodes such as 100 ms, 200 ms, 300 ms, etc. after the starting moment. Therefore, the time alignment method is as Figure 1 shown.
[0031] In this way, the time synchronization between the lidar and the camera data is completed.
[0032] Therefore, in this embodiment, by separately performing spatial synchronization and time synchronization on the road grayscale image and the denoised road point cloud image, it helps to achieve the fusion between the lidar data and the camera data.
[0033] Embodiment 3: This embodiment further limits the road segmentation method for lidar and camera fusion described in Embodiment 1. In step S3, when adjusting the road depth map and fusing the adjusted road depth map with the road RGB image to obtain a road feature image, specifically: The road depth map is completed, and after the completed road depth map is successively processed by cropping, normalization, bilinear interpolation, convolution, max pooling, and average pooling, and then adjusted by normalization, it is fused with the road RGB image through dot multiplication to obtain a road feature image.
[0034] In this embodiment, the completion of the road depth map is specifically: The high-reflection intensity information of the road depth map is retained, and the low-reflection intensity information of the road depth map is completed by interpolation.
[0035] In this embodiment, in the visual perception task, the road RGB image is good at capturing apparent information such as texture and color, while the road depth map contains rich spatial structure and geometric shape information, which can effectively supplement the lack of spatial information in the road RGB image. Effectively fusing these two complementary pieces of information will greatly improve the model's perception ability of the scene and the recognition accuracy of the target.
[0036] However, if the road depth map is not sufficiently preprocessed and guided for modeling, and is only simply stitched with the road RGB image or directly fed into the subsequent model for segmentation, not only its advantages cannot be fully utilized, but it may even bring additional information redundancy and misleading interference.
[0037] In addition, the road depth map often contains a large number of invalid or low-quality pixel values, such as sensor acquisition errors, data loss in areas with too far distances, depth jumps in edge-blurred areas, and depth holes caused by occlusion or reflection. These invalid pieces of information directly participate in the subsequent feature extraction process without screening, which will cause the subsequent model to wrongly focus on non-target areas during segmentation. Thus, it reduces the semantic clarity and discriminability of the overall expression.
[0038] Moreover, there are differences between the road RGB image and the road depth map in terms of spatial resolution, scale, receptive field, etc. If no unified alignment processing is performed, it will be difficult for the subsequent model to achieve pixel-level precise alignment in the fusion stage, resulting in information mismatch and feature perturbation. For example, the numerical range of the road depth map usually has a large span and uneven distribution. If directly used as the model input without processing, it is extremely likely to cause instability in subsequent model training, deviation in gradient update, and even lead to the failure of the attention mechanism.
[0039] To solve the above technical problems, in this embodiment, while generating the road depth map, the unique reflection intensity of lidar data is added, and the depth information with high reflection intensity is preferentially retained, and the low reflection intensity area is filled in by interpolation to obtain a higher-quality road depth map. The following processing is performed on the higher-quality road depth map: As Figure 2 shown, a depth threshold cropping strategy is set to eliminate unreasonable pixels, and the higher-quality road depth map is normalized so that its numerical range is stable between [0, 1]. Subsequently, the higher-quality road depth map is adjusted to the same resolution as the road RGB image through interpolation (for example, bilinear interpolation) to ensure spatial alignment and lay the foundation for subsequent fusion.
[0040] In addition, in the existing fusion methods, most subsequent models ignore the non-uniformity of information contribution between different modalities during segmentation and lack the ability to selectively model depth information. In other words, when processing road RGB images, subsequent models do not actively "pay attention" to the important geometric structure areas contained in the depth, but treat all depth information on an average or uniform basis. This "indiscriminate fusion" strategy often causes subsequent models to perform poorly in complex backgrounds, occluded areas, or under extreme lighting conditions.
[0041] In order to solve the above technical problems, this implementation introduces a 3×3 convolution operation on the interpolated higher quality road depth map to extract its shallow feature representation. This operation can capture local contour information and has a positive effect on enhancing the ability to perceive the target shape. In order to further improve the subsequent model's ability to select the area of interest, the maximum pooling and average pooling are used to perform spatial statistical modeling on the deep features extracted by convolution to generate a saliency attention map.
[0042] In addition, if the attention weight map is not generated based on the road depth map, the semantic guidance ability in the shallow feature map will be limited, resulting in performance degradation in multiple aspects. First, due to the lack of attention to deep structural information, it is difficult for the subsequent model to accurately perceive the edges and geometric details of objects in the image, which can easily cause problems such as blurred contours and broken edges, affecting the final segmentation accuracy. Secondly, shallow features often contain a lot of background interference information. Without superimposed attention weights for suppression, these redundant information may be mistaken for foreground targets, thereby reducing the subsequent model's ability to distinguish foreground areas. In addition, for targets with small sizes or blurred boundaries, the lack of depth guidance will further weaken the recognition ability of subsequent models, making small targets easily ignored or misclassified. In summary, the lack of a fusion strategy based on deep attention limits the decoder's ability to focus on important areas, and also leads to obvious deficiencies in the accuracy, edge perception, and environmental adaptability of the entire semantic segmentation model.
[0043] In order to solve the above technical problems, this implementation method obtains an attention weight map between [0,1] after normalization by the Sigmoid function (activation function), which can effectively depict the key areas in the image with drastic depth changes and prominent structures. Finally, the attention weight map is element-by-element weighted fusion with the road RGB image, so that the subsequent model pays more attention to the semantically significant areas in the process of trunk feature extraction, suppresses background interference, and improves the overall perception effect.
[0044] Therefore, in this embodiment, a 3×3 convolution is used to separately perform convolution on the interpolated higher-quality road depth map to extract its shallow features and structural information, and a 1-channel feature map is output. Based on the extracted depth map features, an attention map is generated through max pooling and average pooling processing, and finally, after normalization by the Sigmoid function, a weight map between [0, 1] is obtained. This process can highlight the areas with obvious depth changes. Finally, the generated attention map is applied to the road RGB image, and each position of the road RGB image is re-weighted by the weights of the attention weight map through dot multiplication, suppressing background noise, strengthening the feature expression of important regions, and improving the performance of the subsequent model under complex backgrounds, occluded regions, or extreme lighting conditions.
[0045] In summary, in this embodiment, after a series of feature extractions are performed on a higher-quality road depth map and then fused with the road RGB image, it not only solves the problems brought to the subsequent model segmentation process due to the existence of a large number of invalid or low-quality pixel values in the road depth map, but also gives full play to the advantages of the road RGB image and the road depth map.
[0046] Embodiment 4: This embodiment further limits the lidar and camera fusion-based road segmentation method described in Embodiment 1. In step S4, the optimized DeepLabv3+ is specifically: CBAM is introduced into both the encoder and decoder of DeepLabv3+. In the decoder of DeepLabv3+, after the road feature image passes through the convolutional neural network algorithm and convolution processing in sequence, it is dot-multiplied with the road depth map adjusted by the attention weights.
[0047] In this embodiment, CBAM (Convolutional Block Attention Module, attention mechanism) is introduced into both the encoder and decoder of DeepLabv3+. Adding CBAM to the encoder enables the model to more effectively integrate features at different levels, and adding CBAM to the decoder enables the model to pay more attention to the target area.
[0048] Introducing CBAM after each branch result, the main purpose is to refine the input feature representation through the attention mechanism. CBAM combines two mechanisms, the channel attention module and the spatial attention module, which can automatically learn the weight distribution relationship between different channels and the response degree of important regions in the spatial dimension, thereby enhancing the expression ability of key features. Specifically, the channel attention module focuses on which channels are important, while the spatial attention module pays more attention to which regions of the image are more important. This special mechanism helps the model to more effectively mine the potential structural and semantic information in the input feature map, improve the model's perception ability of foreground objects, weaken the interference of irrelevant backgrounds, and thus achieve more accurate feature extraction in subsequent tasks.
[0049] To further improve the model's perception ability of key target regions and suppress irrelevant or redundant information noise in the background, a spatial attention mechanism module is introduced. As an important part of CBAM, it performs targeted feature enhancement on the fused feature map.
[0050] The core idea of the spatial attention module is to adaptively adjust the feature weights of each spatial position according to the response differences in the spatial dimension, so as to achieve the model's focus on key regions in the image. In the specific implementation, the spatial attention module first performs compression processing on the fused three-dimensional feature map along the channel dimension. This process generates two two-dimensional maps that only retain spatial information (with a size of ) by performing global average pooling and global max pooling on the feature map respectively, representing different spatial statistical attributes: (1) The average pooling feature map can capture the average response of each spatial position in the image in the overall channel dimension, reflecting the importance of regions, and is suitable for retaining overall structural information; (2) The max pooling feature map focuses on the strongest activation value in the channel, highlighting target saliency more, and is suitable for detecting local high-response regions.
[0051] These two two-dimensional pooling maps are then concatenated along the channel dimension to form a fused map with a size of . This fused map then undergoes a convolution operation (usually a convolution kernel of ) to learn a single-channel spatial attention weight map. After being activated by the Sigmoid function, this weight map is element-wise multiplied with the original fused feature map to achieve re-weighting of the importance of features in different spatial regions.
[0052] The Sigmoid function mentioned above is specifically: ; In the formula, is the output variable, is the input variable, to the power of the base of the natural logarithm , which is the exponential decay of the input.
[0053] In this way, CBAM can automatically mine and enhance the discriminative regions in the image, improving the spatial perception ability in the road segmentation task. Compared with traditional feature fusion methods, CBAM not only achieves more fine-grained region selection but also significantly improves the robustness of the model in difficult scenarios such as occlusion and complex backgrounds.
[0054] In addition, in the decoder of DeepLabv3+, after the road feature image passes through the convolutional neural network algorithm and convolution in sequence, it is dot-multiplied with the road depth map adjusted by the attention weight.
[0055] Input the road feature image into DeepLabv3+ for semantic segmentation. Specifically: Normalize the pixel values of the road feature image to the range, and enhance the road feature image through operations such as flipping. Finally, adjust the road feature image to a fixed size (such as ).
[0056] Such as Figure 2 shown, the adjusted road feature image will enter the backbone network of DeepLabv3+. The backbone network includes two parts: an encoder and a decoder. In order to obtain a feature map with higher resolution in the encoder part, Xception (a convolutional neural network algorithm) with Atrous Convolution (dilated convolution) is selected as the feature extraction network, and low-level feature maps and high-level feature maps are obtained by using multiple convolutions with different dilation rates respectively. Among them, multiple convolutions with different dilation rates are a certain expansion based on the original convolution module, and a larger visual receptive field can be obtained under the premise of the same computational cost and number of parameters.
[0057] The low-level feature maps obtained by Xception directly enter the decoder, and the high-level feature maps are processed by ASPP (Atrous Spatial Pyramid Pooling). ASPP consists of four convolutions with different dilation rates and an Image Pooling used in parallel to capture multi-scale context information and improve the segmentation accuracy by fusing multi-scale information. The results of each branch are respectively passed through CBAM, and the results of each branch output by CBAM are concatenated and then fused through a 1×1 convolution to obtain a feature map containing multi-scale information. The output of ASPP is upsampled by 4 times and fused with the low-level features obtained through 1×1 convolution operation. The attention weight map generated from the road depth map is multiplied pointwise to the low-level feature map, and the fused result after multiplication is passed through CBAM. The output result of CBAM is passed through a 3×3 convolution and upsampled by 4 times to obtain a segmentation map with the same resolution as the input image. The segmentation result map is converted into a binary image, and noise is removed and the main target area is retained through opening and closing operations to obtain a segmentation result with a complete boundary.
[0058] To better process the segmentation results output by the model, the output results are further processed through morphological operations. The main purpose is to remove some isolated misclassified pixels that may exist in the segmentation results output by the model processing, and fill small holes that may exist inside the target area. Specifically: the segmentation result map is converted into a binary image through the maximum probability category, and then small noises are removed while retaining the main target area through the opening operation in morphological operations. Then, the holes inside the target area are filled through the closing operation while retaining the external shape of the target and making the boundary of the target more complete, resulting in a more accurate segmentation result.
[0059] Regarding the evaluation of road segmentation results, the evaluation metrics mainly include AP (average precision), PRE (precision), FPR (false positive rate), and FNR (false negative rate).
[0060] ; ; In the formula, FN (False Negative) is a missed detection, misjudging a positive sample as a negative sample, FP (False Positive) is a false alarm, misjudging a negative sample as a positive sample. TN (True Negative) and TP (True Positive) are both correctly judged.
[0061] Embodiment 5. The road segmentation system for fusing lidar and camera described in this embodiment includes the following modules: A construction module for constructing a road image dataset, where the road image dataset includes a road RGB image obtained by camera photographing and a road point cloud image obtained by lidar; A synchronization module for preprocessing the road RGB image and the road point cloud image respectively to obtain a road grayscale image and a denoised road point cloud image, and performing synchronization processing on the road grayscale image and the denoised road point cloud image; A fusion module for obtaining a road depth map based on the synchronized road grayscale image and the synchronized denoised road point cloud image, adjusting the road depth map, and fusing the adjusted road depth map and the road RGB image to obtain a road feature image; A segmentation module for inputting the road feature image into the optimized DeepLabv3+ for segmentation, thereby completing the segmentation of the road image.
[0062] Embodiment 6. An electronic device described in this embodiment includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store a computer program; The processor is used to implement the road segmentation method for fusing lidar and camera described in any one of Embodiments 1-4 when executing the program stored on the memory.
[0063] Embodiment 7. A computer-readable storage medium described in this embodiment stores a computer program, and the computer program implements the road segmentation method for fusing lidar and camera described in any one of Embodiments 1-4 when executed by a processor.
[0064] The above has introduced in detail the road segmentation method, system, device, and medium for fusing lidar and camera proposed by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A road segmentation method for the fusion of lidar and camera, characterized in that, It includes the following steps: Step S1, construct a road image dataset, where the road image dataset includes a road RGB image obtained by camera photographing and a road point cloud image obtained by lidar; Step S2, preprocess the road RGB image and the road point cloud image respectively to obtain a road grayscale image and a denoised road point cloud image, and synchronously process the road grayscale image and the denoised road point cloud image; Step S3, obtain a road depth map based on the synchronously processed road grayscale image and the synchronously processed denoised road point cloud image, adjust the road depth map, and fuse the adjusted road depth map and the road RGB image to obtain a road feature image; Step S4, input the road feature image into the optimized DeepLabv3+ for segmentation, and then complete the segmentation of the road image.
2. The method for road segmentation by fusing lidar and camera according to claim 1, characterized in that In the step S2, the synchronous processing includes spatial synchronization and temporal synchronization; The spatial synchronization is specifically: Synchronize the coordinate systems of the camera and the lidar, and then complete the spatial synchronization; The temporal synchronization is specifically: The camera and the lidar respectively record a timestamp. Based on the downward compatibility principle, synchronize the road grayscale image and the denoised road point cloud image through the timestamp, and then complete the temporal synchronization.
3. The method for road segmentation by fusing lidar and camera according to claim 1, characterized in that, In the step S3, the adjustment of the road depth map and the fusion of the adjusted road depth map and the road RGB image to obtain a road feature image are specifically: Complete the complementation of the road depth map. After the complemented road depth map is sequentially processed by cropping, normalization, bilinear interpolation, convolution, max pooling, and average pooling, and then adjusted by normalization, fuse it with the road RGB image through dot multiplication to obtain a road feature image.
4. The method for segmenting a road by fusing a lidar and a camera according to claim 3, wherein, The complementation of the road depth map is specifically: Retain the high-reflection intensity information of the road depth map, and complete the complementation of the low-reflection intensity information of the road depth map through interpolation.
5. The method for road segmentation by fusing lidar and camera according to claim 1, characterized in that, In the step S4, the optimized DeepLabv3+ is specifically: Introduce CBAM into the encoder and decoder of DeepLabv3+ respectively. In the decoder of DeepLabv3+, the road feature image is sequentially processed by a convolutional neural network algorithm and convolution, and then dot-multiplied with the road depth map after attention weight adjustment.
6. A road segmentation system integrating lidar and camera, characterized in that, It includes the following modules: A construction module that constructs a road image dataset, where the road image dataset includes a road RGB image obtained by camera photographing and a road point cloud image obtained by lidar; A synchronization module that preprocesses the road RGB image and the road point cloud image respectively to obtain a road grayscale image and a denoised road point cloud image, and synchronously processes the road grayscale image and the denoised road point cloud image; A fusion module that obtains a road depth map based on the synchronously processed road grayscale image and the synchronously processed denoised road point cloud image, adjusts the road depth map, and fuses the adjusted road depth map and the road RGB image to obtain a road feature image; The segmentation module inputs the road feature image into the optimized DeepLabv3+ for segmentation, and thus completes the segmentation of the road image.
7. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store computer programs; The processor is used to implement the method for segmenting a road by fusing a lidar and a camera according to any one of claims 1-5 when executing the program stored on the memory.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the method for segmenting a road by fusing a lidar and a camera according to any one of claims 1-5.
Citation Information
Patent Citations
Small target detection method and system based on laser radar and camera fusion, and medium
CN117671635A
Bridge component three-dimensional point cloud segmentation method based on multi-view data fusion
CN117876397A
Multi-sensor fusion perception method based on attention mechanism and ensemble learning
CN119295874A
Semantic information fused laser radar and vision fusion depth estimation system and method
CN119418337A
Multi-modal target detection method and device and multi-modal identification system
CN119625279A