A 3D Road Scene Generation Method Based on Binocular Depth Estimation
Through a three-dimensional road scene generation method based on binocular depth estimation, combined with edge constraint algorithm and Open3D point cloud processing capability, the problems of low reconstruction accuracy and efficiency in complex road scenes are solved, real-time three-dimensional reconstruction on low-cost hardware is realized.
Patent Information
- Application Number
- CN202510361155.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-26
AI Technical Summary
The existing binocular depth estimation algorithm has problems of insufficient accuracy and high computational complexity in complex road scenarios, making it difficult to realize real-time perception and efficient reconstruction in low-cost hardware environments.
A three-dimensional road scene generation method based on binocular depth estimation is used to obtain left and right viewing images of the road environment through a binocular camera, and the depth information of each pixel is generated by combining the depth estimation algorithm, and the key features of the road scene are highlighted through the edge constraint algorithm, and finally, Open3D is used to achieve efficient three-dimensional road scene generation.
It significantly improves the reconstruction accuracy and efficiency of complex road scenarios, reduces hardware costs, and realizes real-time perception and three-dimensional reconstruction in a low-cost hardware environment.
Smart Images

Figure CN119888093B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the display technology of automotive autonomous driving systems and the field of road scene reconstruction, and particularly to a method for generating a three-dimensional road scene based on binocular depth estimation. Background Art
[0002] In the fields of autonomous driving and intelligent transportation, the three-dimensional reconstruction of road environments is a crucial research direction. Currently, many technical solutions have been applied in this field, but they all have certain limitations. For example, Patent CN110796728B uses lidar to obtain three-dimensional point cloud data and reconstructs the shape size, structure, position, and posture of the target through a greedy projection algorithm. Such traditional three-dimensional reconstruction methods rely on high-precision lidar. Although their accuracy is relatively high, the equipment is expensive and it is difficult to achieve economy in large-scale application scenarios.
[0003] Relatively speaking, cameras, as an economical and efficient sensor, have gradually replaced lidar in some scenarios. For example, Patent CN116091695A uses hierarchical reinforcement learning technology for three-dimensional reconstruction. Although this method can achieve relatively high precision, it mainly targets single objects rather than complex road scenes, so its applicability is limited. Another patent, CN116091703A, adopts a real-time three-dimensional reconstruction technology based on multi-view stereo matching and constructs a three-dimensional model by taking pictures one by one with a monocular camera. Although this solution can improve the reconstruction accuracy to a certain extent, the shooting efficiency is low and the accuracy performance in a dynamic environment is not ideal.
[0004] In contrast, binocular cameras use the parallax information of the left and right views to achieve depth estimation, which is a sensor solution with relatively low cost and low hardware requirements. However, current binocular depth estimation algorithms still face challenges such as insufficient accuracy and high computational complexity when dealing with complex road scenes. Therefore, there is an urgent need for a new technical solution that can effectively reduce the hardware cost and improve the reconstruction accuracy and efficiency in complex scenarios to overcome the limitations of the existing technology. Summary of the Invention
[0005] The present invention proposes a method for generating a three-dimensional road scene based on binocular depth estimation. By using binocular cameras to obtain left and right view images of the road environment, combining depth estimation algorithms to generate depth information for each pixel point, highlighting the key features of the road scene through edge constraint algorithms, and finally using Open3D to achieve efficient generation of three-dimensional road scenes. The present invention can significantly improve the reconstruction accuracy and efficiency of complex road scenes, contribute to the application of autonomous driving and intelligent transportation systems, and particularly achieve real-time perception in a low-cost hardware environment.
[0006] The present invention discloses a method for generating a three-dimensional road scene based on binocular depth estimation, which is based on left and right monocular depth camera sensors, a binocular depth estimation neural network model, an edge-constrained road segmentation module, and a point cloud data optimization module. The method for generating a three-dimensional road scene based on binocular depth estimation includes the following steps:
[0007] Step 1: Dynamically capture road scene images through left and right monocular depth camera sensors;
[0008] Step 2: Calculate the left and right monocular depth images through a binocular depth estimation neural network model;
[0009] Step 3: Generate point cloud data based on the predicted depth map and the camera intrinsic matrix.
[0010] Further, in Step 1, the following steps are also included:
[0011] Step 1.1: Configure the left and right monocular depth cameras according to the intrinsic matrix of the sensor to make them work synchronously, and set the capture frame rate and resolution to meet the real-time processing requirements;
[0012] Step 1.2: Collect left and right images of the road scene through the left and right monocular vision sensors installed on the vehicle and perform preprocessing; including:
[0013] Step 1.2.1: Geometrically correct the images by calibrating the intrinsic matrix and distortion coefficients of the camera:
[0014] , ;
[0015] , ;
[0016] where (x, y, z) are pixel coordinates, (x′, y′) are standard coordinates, ( , ) are corrected coordinates, , , are distortion coefficients, ;
[0017] Step 1.2.2: Extract the road region of interest from the corrected images, and use an edge detection algorithm combined with a set region mask to retain the road surface and important edge features.
[0018] Further, in Step 2, the following steps are also included:
[0019] Step 2.1: Input the preprocessed left and right monocular images into the neural network as the basic data of the model;
[0020] Step 2.2. Extract and fuse multi-level features of the left and right eye images through an encoder;
[0021] Step 2.3. Calculate the initial disparity map through a disparity estimation module;
[0022] ;
[0023] Among them, is the disparity value at the pixel coordinate , which is used to represent the displacement difference between the left and right eye images at this pixel position. is the pixel coordinate, and are the abscissas of the corresponding points in the left and right images respectively;
[0024] Step 2.4. Use a spatial attention mechanism to improve the saliency of key regions in the scene;
[0025] Step 2.5. Enhance the boundary information in the depth map through a depth edge constraint module;
[0026] ;
[0027] Among them, represents the depth value after being optimized by the depth edge constraint module, is the original depth value, is the edge detection value, and λ is the weight parameter;
[0028] Step 2.6. Perform weighted fusion on the depth maps output by the disparity estimation module, the spatial attention module, and the depth edge constraint module;
[0029] Perform weighted fusion on the depth maps output by the disparity estimation module, the spatial attention module, and the depth edge constraint module:
[0030] ;
[0031] Among them, is the final depth map obtained after being weighted by the three modules, is the disparity estimation depth map, is the depth map enhanced by spatial attention, is the depth map enhanced by depth edges, , , are the fusion weights.
[0032] Furthermore, in Step 2.1, the following steps are also included:
[0033] Step 2.2.1. Multi-level feature extraction, including:
[0034] Low-level feature extraction: Edge and texture features of the left and right eye images are extracted through the initial convolutional layer:
[0035] ;
[0036] Among them, is the pixel value of the input image, is the width of the convolutional kernel, is the height of the convolutional kernel, is the low-level convolutional kernel, is the low-level feature output;
[0037] Middle-level feature extraction: A convolutional module with a pooling layer is used to extract geometric and texture information in the scene and reduce redundant data:
[0038] ;
[0039] Among them, is the pixel coordinate range of the pooling window in the horizontal direction, is the pixel coordinate range of the pooling window in the vertical direction, is the pixel coordinate index within the pooling window, is the middle-level feature output;
[0040] High-level feature extraction: Stacked convolutional layers are used to extract scene semantic features:
[0041] ;
[0042] Among them:
[0043] is the high-level feature output;
[0044] is the feature map output of the previous layer, and the feature value on channel ; is the feature position within the local window;
[0045] is the convolutional kernel, with a size of , an input channel of , and an output channel of ; is the bias term for output channel c; is the activation function, used to introduce non-linearity;
[0046] By stacking multiple convolutional layers, the receptive field is gradually expanded and semantic features are captured. The formula is iteratively carried out until finally ;
[0047] Combine low, medium, and high-level features through skip connections to form hierarchical features:
[0048] ;
[0049] , , respectively represent the weights of low, medium, and high-level features in the fusion process.
[0050] Furthermore, in step 2.4, the following steps are also included:
[0051] Step 2.4.1, multi-scale feature generation;
[0052] Input feature map Extract features of different scales through different convolutional kernels:
[0053] ;
[0054] Among them, represents the convolutional operation with a kernel size of , represents the -th scale feature map generated at the pixel coordinate after the convolutional operation with a kernel size of ;
[0055] Step 2.4.2, superimpose the feature maps of different scales according to the weights:
[0056] ;
[0057] Among them, is the number of scales, ensures weight normalization, represents the final feature value after fusing all scale features at the pixel coordinate ;
[0058] Step 2.4.3, superimpose the fused feature map on the input feature map to obtain the final enhanced feature:
[0059] ;
[0060] represents the enhanced feature map value that combines the original hierarchical feature and the multi-scale fusion feature at the pixel coordinate .
[0061] Furthermore, in step 3, the following steps are also included:
[0062] Step 3.1: Combine the depth map with the camera intrinsic matrix to generate initial point cloud data:
[0063] ;
[0064] where, is the pixel coordinate; ( ) is the optical center position of the camera; , are the focal lengths of the camera in the horizontal and vertical directions respectively; is the depth value; X and Y represent the positions of the 3D point in the horizontal and vertical directions of the camera coordinate system respectively, and Z is the depth in the direction of the camera optical axis;
[0065] Step 3.2: Downsample and denoise the initial point cloud to improve data quality; including the following steps:
[0066] Step 3.2.1: Divide the initial point cloud into cubic grids of size s by the voxel grid method, and represent the points in each grid with the centroid of the points in the grid:
[0067] ;
[0068] where, N is the number of points in the grid, is the centroid coordinate of the grid, is the coordinate of each point in the grid;
[0069] Step 3.2.2: Calculate the average distance between each point and its nearest neighbors, and remove the outlier points that do not meet the following conditions:
[0070] ;
[0071] where, is the average distance, is the standard deviation, is the user-set threshold, is the average distance between each point and its nearest neighbors;
[0072] Step 3.3: Input the optimized point cloud data into the Open3D library, generate and render the 3D road scene, add color information to the point cloud data, and map the depth value to color:
[0073] ;
[0074] where, is the depth value of point , and They are the maximum and minimum depth values respectively. The beneficial effects achieved by the present invention are as follows:
[0075] Through the binocular depth estimation combined with the edge constraint algorithm, the present invention improves the accuracy of 3D reconstruction in complex road scenes;
[0076] By utilizing the efficient point cloud processing and rendering capabilities of Open3D, the present invention realizes real-time performance and visualization;
[0077] The present invention uses a low-cost binocular camera to replace the lidar, significantly reducing the hardware cost and providing feasibility for large-scale applications. Description of the Drawings
[0078] Figure 1 is the process framework of a 3D road scene generation method based on binocular depth estimation;
[0079] Figure 2 is the binocular depth estimation network model of a 3D road scene generation method based on binocular depth estimation;
[0080] Figure 3 is the point cloud data processing process of a 3D road scene generation method based on binocular depth estimation. Detailed Embodiments
[0081] The present invention will be further described below in conjunction with specific embodiments, and the advantages and features of the present invention will become clearer as the description progresses. However, these embodiments are merely exemplary and do not constitute any limitation to the scope of the present invention. Those skilled in the art should understand that without departing from the spirit and scope of the present invention, modifications or substitutions can be made to the details and forms of the technical solutions of the present invention, but such modifications and substitutions all fall within the protection scope of the present invention.
[0082] As Figure 1 shown, this embodiment is the process framework of a 3D road scene generation method based on binocular depth estimation. This framework includes the following modules:
[0083] Left and right binocular depth camera sensors: used to capture road scene images in real time, equipped with the internal parameter matrix of the sensors;
[0084] Binocular depth estimation neural network model: used to generate the depth map of the road scene;
[0085] Edge constraint road segmentation module: optimize the features of road edges and scene boundaries;
[0086] Point cloud data optimization module: perform downsampling, noise reduction, and rendering optimization processing based on the generated point cloud data.
[0087] This method specifically includes the following steps:
[0088] Step 1: Capture and preprocess the left and right eye images.
[0089] Dynamically capture the road scene images through the left and right eye depth camera sensors. This step includes the following content:
[0090] Step 1.1: Initialize the left and right vision sensors;
[0091] Configure the left and right eye depth cameras according to the internal parameter matrix of the sensors to make them work synchronously, and set the capture frame rate and resolution to meet the real-time processing requirements.
[0092] Step 1.2: Capture and preprocess the road scene images;
[0093] Collect the left and right road scene images through the left and right eye vision sensors installed on the vehicle and perform preprocessing; specifically, it includes the following steps:
[0094] Step 1.2.1: Image distortion correction;
[0095] Perform geometric correction on the images through the camera calibration parameters (internal parameter matrix and distortion coefficients). The specific formula is as follows:
[0096] , ;
[0097] , ;
[0098] Among them, (x, y, z) are pixel coordinates, (x′, y′) are standard coordinates, , , are distortion coefficients, .
[0099] Step 1.2.2: Region of interest (ROI) extraction;
[0100] Extract the road region of interest from the corrected images, and use edge detection algorithms (such as Canny) combined with the set region masks to retain the road surface and important edge features.
[0101] Step 2: Input and calculation of the binocular depth estimation neural network model;
[0102] As Figure 2 shown, this embodiment shows the architecture of the binocular depth estimation network model, which specifically includes the following modules:
[0103] Disparity estimation module: Generate the initial disparity map;
[0104] Spatial attention module: Highlight the key regions in the scene and enhance the accuracy of depth estimation;
[0105] Depth Edge Constraint Module: Enhance the clarity of the boundary through semantic edge enhancement.
[0106] The specific steps are as follows:
[0107] Step 2.1: Input the left and right eye images.
[0108] Input the preprocessed left and right eye images into the neural network as the basic data of the model.
[0109] Step 2.2: Feature encoding
[0110] Extract three types of features from the left and right eye images through the encoder: low-level features (such as edge information), middle-level features (such as texture information), and high-level features (such as scene semantic information).
[0111] The specific steps are as follows:
[0112] Step 2.2.1: Multi-level feature extraction
[0113] (1) Low-level feature extraction: Extract edge and texture features from the left and right eye images through the initial convolutional layer.
[0114] Use multiple 3×3 convolutional kernels to slide on the image, and the calculation formula is as follows:
[0115] ;
[0116] Where, is the pixel value of the input image, is the low-level convolutional kernel, is the low-level feature output.
[0117] (2) Middle-level feature extraction: Use the convolutional module with a pooling layer to extract geometric and texture information in the scene and reduce redundant data.
[0118] Pooling operation formula:
[0119] ;
[0120] Where, , is the window range, is the middle-level feature output.
[0121] (3) High-level feature extraction: Use stacked convolutional layers to extract scene semantic features, such as roads, obstacles, etc.
[0122] The high-level feature is expressed as:
[0123] ;
[0124] Wherein: is the high-level feature output;
[0125] is the feature map output of the previous layer, and the channel index is i;
[0126] is the convolution kernel, with a size of , the input channel is C, and the output channel is ; is the bias term of the convolution; is the activation function, which is used to introduce non-linearity.
[0127] By stacking multiple convolutional layers, the receptive field is gradually expanded and semantic features are captured. The formula is iteratively carried out until finally is obtained.
[0128] Step 2.2.2, Feature fusion;
[0129] The low, medium, and high-level features are combined through skip connections to form hierarchical features:
[0130] ;
[0131] , , respectively represent the weights of the low, medium, and high-level features in the fusion process.
[0132] Step 2.3, Disparity estimation;
[0133] After feature extraction and fusion are completed, this step aims to generate an initial disparity map through the disparity estimation module, which serves as the basis for depth estimation.
[0134] The disparity map reflects the displacement difference between the left and right images, is directly related to the depth information, and provides input data for subsequent depth edge optimization. The calculation formula of the disparity estimation module is as follows:
[0135] ;
[0136] Wherein, is the pixel coordinate, and are the abscissas of the corresponding points in the left and right images respectively.
[0137] Step 2.4, Spatial attention enhancement;
[0138] To further optimize the accuracy of depth estimation and compensate for the perspective phenomenon that the disparity estimation fails to solve, a spatial attention mechanism is introduced in this step. The spatial attention mechanism is used to enhance the saliency of key regions in the scene (such as lane lines and road edges). The specific steps are as follows:
[0139] Step 2.4.1, Multi-scale feature generation;
[0140] The feature map obtained by feature extraction in Step 2.2.2 is used to extract features of different scales through different convolutional kernels to generate a multi-scale feature map:
[0141] ;
[0142] where, represents the convolution operation with a kernel size of .
[0143] Step 2.4.2, Cross-scale fusion;
[0144] After Step 2.4.1, feature maps of different scales are obtained. The feature maps of different scales are superimposed according to weights, and the formula is:
[0145] ;
[0146] where, is the number of scales, ensuring weight normalization.
[0147] Step 2.4.3, Enhanced output;
[0148] The fused feature map is superimposed on the feature map to obtain the final enhanced feature:
[0149] ;
[0150] Step 2.5, Depth edge constraint;
[0151] After completing the spatial attention enhancement in Step 2.4, to further improve the boundary clarity of the depth map and reduce the blurred area, a depth edge constraint module is introduced in this step. This module optimizes the boundary characteristics of key regions by fusing edge detection information and depth map data, providing more accurate input data for the final 3D reconstruction. The depth edge constraint module is optimized by the following formula:
[0152] ;
[0153] where, is the original depth value, is the edge detection value, and λ is the weight parameter.
[0154] Step 2.6, Depth map fusion;
[0155] Perform weighted fusion on the depth maps output by the disparity estimation module, spatial attention module, and depth edge constraint module:
[0156] ;
[0157] Among them, is the disparity estimation depth map, is the spatially attention-enhanced depth map, is the depth edge-enhanced depth map, , , is the fusion weight.
[0158] Step 3, Point cloud data generation and optimization;
[0159] As Figure 3 shown, this embodiment demonstrates the point cloud data processing flow of this method.
[0160] This step utilizes the optimized depth map output in Step 2 and the camera intrinsic matrix, and specifically includes the following steps:
[0161] Step 3.1, Initial point cloud generation;
[0162] Combine the depth map with the camera intrinsic matrix to generate initial point cloud data, and the formula is as follows:
[0163] ;
[0164] Among them, is the pixel coordinate; ( , ) is the position of the optical center; , is the focal length; is the depth value; X and Y respectively represent the positions of the three-dimensional point in the horizontal and vertical directions of the camera coordinate system, and Z is the depth in the direction of the camera optical axis.
[0165] Step 3.2, Point cloud optimization;
[0166] Perform downsampling and denoising on the initial point cloud to improve data quality.
[0167] The specific implementation steps are as follows:
[0168] Step 3.2.1, Voxel grid downsampling;
[0169] The initial point cloud is divided into cubic grids of size s by the voxel grid method, and the centroid of the points within each grid is used to represent the points of the grid. The formula is as follows:
[0170] ;
[0171] where N is the number of points within the grid, is the centroid coordinate of the grid, are the coordinates of the points within each grid.
[0172] Step 3.2.2, Statistical filtering denoising;
[0173] After the downsampling process by the voxel grid method, calculate the average distance between each point and its nearest neighbors, and remove the outlier points that do not meet the following conditions:
[0174] ;
[0175] where μ is the average distance, σ is the standard deviation, and α is the user-defined threshold.
[0176] Step 3.3, 3D scene construction;
[0177] Input the optimized point cloud data into the Open3D library to generate and render the 3D road scene. Add color information to the point cloud data by mapping the depth value to color. The specific formula is as follows:
[0178] ;
[0179] where, is the depth value of the optimized point , , are the maximum and minimum depth values respectively.
[0180] The above are only the specific steps of the present invention and do not constitute any limitation to the protection scope of the present invention; all technical solutions formed by equivalent transformation or equivalent substitution fall within the scope of the protection of the rights of the present invention; the parts not elaborated in detail in the present invention belong to the well-known technologies of those skilled in the art.
Claims
1. A three-dimensional road scene generation method based on binocular depth estimation, based on left and right eye depth camera sensors, a binocular depth estimation neural network model, an edge constraint road segmentation module and a point cloud data optimization module, characterized in that: The three-dimensional road scene generation method based on binocular depth estimation comprises the following steps: Step 1: Dynamically capture the road scene through the left and right depth camera sensors; Step 2: Calculate the left and right eye depth images through a binocular depth estimation neural network model; Step 3: Generate point cloud data based on the predicted depth map and camera intrinsic parameter matrix; In step 2, the following steps are also included: Step 2.1, input the preprocessed left and right eye images into the neural network as the basic data of the model; Step 2.2: Extract and fuse the multi-layer features of the left and right images through the encoder; Step 2.3, calculating the initial disparity map through the disparity estimation module; ; in, is the pixel coordinate The disparity value on is used to indicate the displacement difference between the left and right images at the pixel position. is the pixel coordinate, and are the horizontal coordinates of the corresponding points in the left and right images respectively; Step 2.4: Use the spatial attention mechanism to increase the saliency of key areas of the scene; Step 2.5, enhancing the boundary information in the depth map through the depth edge constraint module; ; in, Represents the depth value after optimization by the depth edge constraint module. is the original depth value, is the edge detection value, λ is the weight parameter; Step 2.6: weighted fusion of the depth maps output by the disparity estimation module, the spatial attention module, and the depth edge constraint module; The depth maps output by the disparity estimation module, spatial attention module, and depth edge constraint module are weighted fused: ; in, is the final depth map obtained after weighted processing of the three modules. is the depth map for disparity estimation, Enhance the depth map for spatial attention, Enhance the depth map for depth edges, , , is the fusion weight.
2. The method for generating a three-dimensional road scene based on binocular depth estimation according to claim 1, characterized in that: In step 1, the following steps are also included: Step 1.1, configure the left and right depth cameras according to the sensor's intrinsic parameter matrix, make them work synchronously, and set the capture frame rate and resolution to meet real-time processing requirements; Step 1.2: Collect left and right images of the road scene through left and right vision sensors installed on the vehicle and perform preprocessing; including: Step 1.2.1: Use the camera to calibrate the intrinsic parameter matrix and distortion coefficients to perform geometric correction on the image: , ; , ; Among them, (x, y, z) is the pixel coordinate, (x′, y′) is the standard coordinate, ( , ) are the corrected coordinates, , , is the distortion coefficient, ; Step 1.2.2: Extract the road region of interest from the rectified image, use the edge detection algorithm combined with the set region mask to retain the road surface and important edge features.
3. The method for generating a three-dimensional road scene based on binocular depth estimation according to claim 1, characterized in that: In step 2.1, the following steps are also included: Step 2.2.1, multi-level feature extraction, including: Low-level feature extraction: Edge and texture features are extracted from the left and right images through the initial convolution layer: ; in, is the input image pixel value, is the width of the convolution kernel, is the height of the convolution kernel, is the low-level convolution kernel, Output for low-level features; Mid-level feature extraction: Use convolutional modules with pooling layers to extract geometric and texture information in the scene and reduce redundant data: ; in, The pixel coordinate range of the pooling window in the horizontal direction, is the pixel coordinate range of the pooling window in the vertical direction, is the pixel coordinate index within the pooling window, Output for mid-level features; High-level feature extraction: Use stacked convolutional layers to extract scene semantic features: ; in: Output for high-level features; It is the feature map output of the previous layer, in channel The eigenvalues on is the feature position within the local window; is the convolution kernel, and its size is , the input channel is , the output channel is ; is the bias term of output channel c; is the activation function, used to introduce nonlinearity; By stacking multiple convolutional layers, the receptive field is gradually expanded and the semantic features are captured. The formula is iterated until the final result is ; The low, medium and high level features are combined through skip connections to form hierarchical features: ; , , They respectively represent the weights of low-level, medium-level, and high-level features in the fusion process.
4. The method for generating a three-dimensional road scene based on binocular depth estimation according to claim 3, characterized in that: In step 2.4, the following steps are also included: Step 2.4.1, multi-scale feature generation; Input feature map Features of different scales are extracted through different convolution kernels: ; in, The kernel size is The convolution operation, It means that the convolution kernel size is After the convolution operation, at the pixel coordinates The generated Scale feature map; Step 2.4.2: Superimpose feature maps of different scales according to weights: ; in, is the scale quantity, Make sure the weights are normalized, Represented in pixel coordinates Above, the final eigenvalue after integrating all scale features; Step 2.4.3: Superimpose the fused feature map onto the input feature map to obtain the final enhanced feature: ; Represented in pixel coordinates The original hierarchical features are combined and multi-scale fusion features The enhanced feature map value.
5. The method for generating a three-dimensional road scene based on binocular depth estimation according to claim 1, characterized in that: In step 3, the following steps are also included: Step 3.1: Depth map Combined with the camera intrinsic parameter matrix, the initial point cloud data is generated: ; in, is the pixel coordinate; ( ) is the optical center position of the camera; , are the focal lengths of the camera in the horizontal and vertical directions respectively; is the depth value; X, Y represent the horizontal and vertical positions of the three-dimensional point in the camera coordinate system respectively, and Z is the depth in the direction of the camera optical axis; Step 3.2: Downsample and denoise the initial point cloud to improve data quality; including the following steps: Step 3.2.1, divide the initial point cloud into cubic grids of size s by voxel grid method, and use the centroid of each point in the grid to represent the point of the grid: ; Where N is the number of points in the grid, is the centroid coordinate of the grid, is the coordinates of each point in the grid; Step 3.2.2, calculate each point and its The average distance of the nearest neighbors, eliminating outliers that do not meet the following conditions: ; in, is the average distance, is the standard deviation, Set thresholds for users, For each point The average distance of the nearest neighbors; Step 3.3: Input the optimized point cloud data into the Open3D library, generate and render a 3D road scene, add color information to the point cloud data, and map the depth value to color: ; in, The optimized point The depth value of and are the maximum and minimum depth values respectively.
Citation Information
Patent Citations
A method for 3D reconstruction of non-cooperative spacecraft based on scanning lidar
CN110796728B
Unmanned aerial vehicle-based real-time three-dimensional reconstruction method
CN108428255A