A method for dense measurement of lidar based on enhanced imagination computing

By coupling neural networks to fuse the depth and intensity information of lidar, a three-dimensional implicit neural directed distance field is built, which solves the problem of lidar data sparsity, realizes high-resolution data density, reduces hardware costs, and improves application capabilities.

CN116934828BActive Publication Date: 2025-06-20HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310917452.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-25
Publication Date
2025-06-20
Estimated Expiration
2043-07-25

AI Technical Summary

Technical Problem

The original measurement data of existing lidar sensors is sparse, making it difficult to directly apply computer vision technology that relies on dense measurements, and adding hardware configurations to obtain dense data will lead to increased costs, limiting its application and promotion.

Method used

The dense measurement method of lidar based on imagination computing is adopted. By coupling neural networks, the depth information and intensity information provided by lidar are fused to construct a three-dimensional implicit nerve directed distance field to achieve the density of sparse data and provide high-resolution depth and intensity information.

Benefits of technology

The density of lidar data is achieved, high-resolution depth and intensity information is provided, the cost of increasing hardware configuration is reduced, and the application capability of lidar in actual scenarios is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116934828B_ABST
    Figure CN116934828B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for dense measurement of lidar based on enhanced imagination computing. The data collected by the lidar is projected onto a two-dimensional plane to generate a depth map and an intensity map. An N-level sparse voxel octree is constructed to encode the signed distance field (SDF). After sampling and interpolation, a coupled neural network is used to fuse the depth information and intensity information provided by the lidar, mine the spatial information and prior cognitive information jointly contained in the sparse lidar data, construct a three-dimensional implicit neural signed distance field, predict the distance value and intensity value of the scene surface, imagine the continuous three-dimensional spatial distribution, realize the densification of sparse data, and provide high-resolution depth and intensity information. The prediction result is more accurate, and at the same time, the problem of high cost required to increase hardware to obtain dense lidar data is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of lidar, and uses dense high-resolution information to enhance the measurement ability of lidar. Specifically, it relates to a lidar dense measurement method based on imagination calculation enhancement. Background Art

[0002] Common lidar sensors include mechanically rotating lidars and phased array lidars. Due to the characteristic of actively projecting pulsed laser beams and measuring the environment through backscattered echoes, lidar sensors have the advantage of still being able to work well in environments with poor lighting conditions compared to visible light cameras that passively receive ambient reflected light. Lidar sensors can measure the distance from the sensor to the environmental reflection surface, obtain depth and intensity images, and can measure the distance, speed, and direction of surrounding objects in real time. Therefore, they have a wide range of applications in fields such as unmanned driving, topographic mapping, and building quality control.

[0003] However, the original measurement data of lidar sensors is sparse, making it difficult to directly apply many computer vision techniques that rely on dense measurement, such as object segmentation and optical flow. Existing technologies obtain dense environmental data by increasing hardware configuration, using complex rotating structures or phased array phases to ensure that the laser can scan the entire environmental area, or increasing the number of lasers inside the lidar. Whichever method is used, it will increase the usage cost of the lidar, further restricting the application and popularization of lidar in actual scenarios. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention proposes a lidar dense measurement method based on imagination calculation enhancement. It uses a coupled neural network to fuse the depth information and intensity information provided by the lidar, mines the spatial information and prior cognitive information jointly contained in the sparse lidar data, constructs a three-dimensional implicit neural signed distance field, imagines the continuous three-dimensional spatial distribution, realizes the densification of sparse data, and provides high-resolution depth and intensity information.

[0005] Step 1: Use a lidar to scan the scene to be measured, and collect sparse point cloud and corresponding reflected intensity value data. Project the collected original data onto a two-dimensional plane to obtain a depth map and an intensity map. Among them, the pixel value of the depth map represents the measured distance value, and the pixel value of the intensity map represents the reflected intensity value. Generate a signed distance field SDF according to the depth map, where the SDF represents the distance between a voxel in space and the nearest object to it.

[0006] Step 2: Construct a sparse voxel octree with N levels, encode the signed distance field (SDF), store the voxels at the vertices of the spatial grid, and the voxels contain feature vectors of depth information and intensity information. The sparse voxel octree only allocates voxels for the point cloud when the spatial grid contains it. The voxel resolution ranges from the coarsest resolution to the finest resolution [Nmax, Nmin], and is divided into N levels in sequence.

[0007] Step 3: According to the position of the pixel to be generated on the image, the pose of the sensor optical center, and the internal parameters of the sensor, generate a ray passing through the pixel to be generated and sample it sequentially on the ray. For each sampling point in sequence, search for its feature vector in the sparse voxel octree. Traverse the sparse voxel octree to find the voxels containing the sampling point at each level of resolution. By performing trilinear interpolation at the position of the sampling point, obtain the feature vectors at this position at different resolutions, and splice them in order to obtain the feature vectors at multiple levels, which are used as the input of the coupled neural network to estimate the signed distance value SDF and intensity value i at the sampling point x.

[0008] The coupled neural network is:

[0009]

[0010] Among them, represents the observation angle at the position of the sampling point x, z(x) represents the feature vector at the sampling point x, Ω represents the entire voxel space included in the sparse voxel octree, θ represents the parameters of the coupled neural network, and i represents the lidar intensity predicted by the coupled neural network.

[0011] Coupled neural network F θ First, input the feature vector z(x) at the position of the sampling point x into a neural network to obtain the predicted SDF and the state parameters ξ of the hidden layer of the neural network. Then, input the state parameters ξ and the observation angle together into another neural network to obtain the radiance intensity value related to the observation angle Therefore, the signed distance SDF is only related to the position of the sampling point x, while the intensity value i is related to both the position of the sampling point x and the observation angle related.

[0012] Step 4: Use the SDF loss, intensity loss, and semantic feature loss to form an overall loss function By minimizing the overall loss function, jointly optimize the features in the multi-scale sparse voxel octree and the network parameters of the coupled neural network F θ of:

[0013]

[0014] Among them, Represents the semantic consistency feature loss, Represents the intensity loss, Represents the SDF loss at the n-th level resolution of the sparse voxel octree.

[0015] By minimizing the SDF loss to optimize the predicted value SDF(x) of the output of the coupled neural network, minimizing the difference between the reconstructed scene and the rendered 3D model:

[0016]

[0017]

[0018] where S represents the surface of the predicted scene, and Ω\S represents the positions in the entire voxel space that do not contain the surface; Represents the gradient at the predicted surface x, Represents the ground truth normal vector; f(x; θ) represents the SDF value at the sampling point x predicted by the multi-layer perceptron, θ represents the parameters of the coupled neural network, and SDF(x) refers to the ground truth SDF value at the sampling point x; U(x, δ) represents the δ neighborhood of x, including other sampling points whose distance from the sampling point x is less than the distance threshold δ. σ and λ are hyperparameters, is the eikonal equation.

[0019] Use the intensity loss to minimize the intensity residual. According to the Gaussian distribution model of the distance field, the radiation information obtained from multiple sampling points is weighted and accumulated to predict the intensity value of the pixel When there is a corresponding measured ground truth for the predicted pixel, calculate the intensity loss function constructed with the goal of minimizing the intensity residual:

[0020]

[0021] where, Represents the intensity value predicted by the coupled neural network for the pixel to be generated, and i represents the intensity value of the original lidar sampling point.

[0022] Use semantic features to minimize the semantic features between the original lidar data and the predicted data, so that the coupled neural network can better predict the SDF value and the intensity value. Use a feature extractor to extract the semantic features of the original depth image, intensity image, and the depth image and intensity image optimized and completed by the network obtained from the lidar respectively, and minimize the difference between the semantic features of the original image and the network-optimized and completed image to minimize the semantic consistency feature difference.

[0023] Step 5: Generate a set of rays. A ray with the origin at x0 and direction d is represented by r(t) = x0 + td, where t > 0 and t is the coefficient of the direction vector d. Retrieve each ray in the ray set. In the trained sparse voxel octree, retrieve the first voxel V that intersects with the ray at the optimal L-level resolution. L Let its coordinates be x k . Recursively retrieve the parent voxels of voxel V L at different resolutions, and then perform trilinear interpolation of the eigenvector for x k at each level of resolution, and use the coupled neural network F θ to calculate the distance from x k to the nearest surface. Using as the step size, find the next query point x k+1 , until the ray disappears or reaches the object surface.

[0024] Repeat the above retrieval process, perform weighted accumulation on the radiation information obtained at the sampling points, predict the missing intensity values on the original sparse intensity map, and complete the complementation of the intensity map.

[0025] The present invention has the following beneficial effects:

[0026] 1. Jointly optimize the depth map and intensity map of the lidar, make full use of the intensity information to obtain better depth prediction results. Densify the lidar data based on the existing data in software, which solves the problem of high cost required to add hardware to obtain dense lidar data.

[0027] 2. Use the semantic information of the lidar depth map and intensity map as prior cognitive information to optimize the network, making the densification result predicted by the network more in line with the actual situation and more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is the voxel grid model representing the scene in the embodiment;

[0029] Figure 2 is the schematic diagram of coordinate point interpolation in the embodiment;

[0030] Figure 3 is the schematic diagram of multi-scale voxel feature interpolation in the embodiment;

[0031] Figure 4 is the input and output diagram of the encoded MLP network in the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0032] The following further explains and illustrates the present invention with reference to the accompanying drawings;

[0033] The present invention proposes a lidar dense measurement method based on enhanced imagination computing. A coupled neural network is used to fuse the depth information and intensity information provided by the lidar, excavate the spatial information and prior cognitive information jointly contained in the sparse lidar data, construct a three-dimensional implicit neural signed distance field, imagine the continuous three-dimensional spatial distribution, realize the densification of sparse data, and provide high-resolution depth and intensity information. The specific steps are as follows:

[0034] Step 1: Use a handheld lidar terminal device to take a video of the scene to be measured. The terminal device includes a lidar and an inertial measurement device. The lidar is used to provide the original information required to construct the scene model, including the sparse point cloud scanned by the lidar and the reflection intensity value of the point cloud. Among them, the sparse point cloud and its intensity value can be projected onto a two-dimensional plane to obtain another representation form of the original data, namely the depth map and the intensity map. The meaning represented by the pixel value of the depth image is the measured distance value, and the meaning represented by the pixel value of the intensity image is the reflection intensity value. The two image information corresponds one by one according to the pixel coordinates.

[0035] During the shooting process, the device terminal should be kept as stable and slowly translated as possible, so that the captured data has a good coverage range and can show the smallest scene movement. Using the external perception and self-motion information provided by the device, combined with the pre-calibrated internal and external parameters, calculate the accurate position and pose of the device in the world coordinate system, and then use the spatio-temporal registration method to correct, pair, splice, and project the collected images. After fusion, it is used as the dataset for constructing the scene depth and intensity model. The dataset image is an image obtained by projecting single-frame point cloud data onto a two-dimensional plane, or a composite image obtained by projecting multi-frame point cloud data after fusion onto a two-dimensional plane.

[0036] Step 2: Set the voxel resolution from the coarsest resolution to the finest resolution [N max ,N min , construct a sparse voxel octree with N levels, encode the signed distance field, store the voxels at the grid vertices of the space, and the voxels contain the feature vectors of depth information and intensity information. Only when the space contains point clouds, the sparse voxel octree will allocate voxels for it.

[0037] According to the position of the pixel to be generated on the image, the pose of the sensor optical center, and the internal parameters of the sensor, generate a ray passing through the pixel to be generated, and sample sequentially on the ray. For each sampling point in turn, search for its feature information in the sparse voxel octree.

[0038] Step 3: Traverse the sparse voxel octree to find the voxels containing the sampling points at each level of resolution, such as Figure 2As described above, by performing trilinear interpolation at the sampling point position, the feature vectors at this position under different resolutions are obtained, and the feature vectors at multiple levels are obtained by concatenating them in order. As the input of the coupled neural network, the signed distance function (SDF) value SDF and the intensity value i at the sampling point x are estimated, as Figure 3 shown.

[0039] The coupled neural network is composed of two multi-layer perceptron (MLP) neural networks. The first MLP neural network takes the feature vector z(x) at the position of the sampling point x as the input, and outputs the SDF and the state parameter ξ of the hidden layer of the neural network. The second MLP neural network takes the state parameter ξ and the observation angle as the input, and outputs the radiance intensity value i.

[0040] As Figure 4 shown, where N = 0 and N = 1 represent the results of linear interpolation of the voxel grids at two different levels of resolution for the same sampling point x. The results of these two and the results obtained at all other scales are combined into a feature vector as the input of the first MLP neural network to obtain the SDF value at the position of the sampling point x. Then, the state parameter ξ of the hidden layer of the first MLP neural network and the observation ray direction of the sampling point are used as the input of the second MLP neural network to obtain the radiance intensity value i.

[0041] Step 4: Use the SDF loss, intensity loss, and semantic feature loss to form an overall loss function, and directly minimize the overall loss function to jointly optimize the features in the multi-scale sparse voxel octree and the network parameters of the coupled neural network F θ . The overall loss function is:

[0042]

[0043] where represents the semantic consistency feature loss, represents the SDF loss at the nth level of resolution of the sparse voxel octree, represents the intensity loss. By minimizing the SDF loss , the predicted value SDF(x) output by the coupled neural network is optimized.

[0044] Step 5: Input the image to be completed and the observation angle to generate a set of rays passing through the image. The ray with the origin x0 and direction d is represented by r(t) = x0 + td, t > 0, where t is the coefficient of d. Each ray in the ray set is retrieved, and in the trained sparse voxel octree, the optimal voxel V L intersected by the ray at the first level of resolution L is retrieved, and its coordinate is set as x k . Recursively retrieve the voxel V LParent voxels at different resolutions, and then perform trilinear interpolation of the feature vectors for x at each resolution level k and use the coupled neural network F θ to calculate x k the distance to the nearest surface Using as the step size, find the next query point x k+1 , until the ray disappears or reaches the object surface, record the coordinates of the object surface, and use this coordinate as the observation angle and input them into the trained coupled neural network together to predict the SDF value and intensity value i of the object surface. Repeat the above retrieval process to predict the missing intensity values on the original sparse intensity map and complete the completion of the intensity map.

Claims

1. A method for dense measurement of lidar based on enhanced imagination computing, which uses lidar to scan the scene to be measured and collects sparse point cloud and corresponding reflected intensity value data, characterized in that: Densify the collected data using a neural network, and the specific steps are as follows: Step 1: Project the collected original data onto a two-dimensional plane to obtain a depth map and an intensity map; generate a signed distance field (SDF) based on the depth map; Step 2: Construct a sparse voxel octree with N hierarchical resolutions for encoding the signed distance field (SDF) and storing the voxels at the vertices of the spatial grid; the sparse voxel octree only allocates voxels for the point cloud when the spatial grid contains it; the voxel contains a feature vector of depth information and intensity information; Step 3: Generate rays passing through the pixels to be generated in the image and sample them sequentially on the rays; Search for the feature vectors of each sampling point in the sparse voxel octree sequentially; Traverse the sparse voxel octree to find the voxel containing the sampling point at each level of resolution. By performing trilinear interpolation at the sampling point position, obtain the feature vectors at different resolutions at this position, and concatenate them in order to obtain the feature vectors at multiple levels, which serve as the input to the coupled neural network F θ to estimate the signed distance value SDF and intensity value i at the sampling point position; Step 4: By minimizing the overall loss function Jointly optimize the features in the multi-scale sparse voxel octree and the network parameters of the coupled neural network F θ : Among them, represents the semantic consistency feature loss, represents the intensity loss, represents the SDF loss at the n-th level resolution of the sparse voxel octree; Step 5: Generate a set of rays, where a ray with origin \(x_0\) and direction \(d\) is represented as \(r(t)=x_0 + td\), \(t\) is the coefficient of the direction vector \(d\), and \(t>0\); retrieve each ray in the ray set, and in the trained sparse voxel octree, retrieve the first voxel \(V\) at the optimal \(L\)-th level of resolution that intersects with the ray L , and let its coordinate be \(x\) k ; recursively retrieve the parent voxels of voxel \(V\) L at different resolutions, and then perform trilinear interpolation of the feature vectors for \(x\) k at each level of resolution, and use the coupled neural network \(F\) θ to calculate the distance between \(x\) k and the nearest surface Take as the step size to find the next query point \(x\) k+1 , until the ray disappears or reaches the object surface.

2. The method for dense measurement of lidar based on enhanced imagination computing according to claim 1, characterized in that: The coupled neural network is as follows: Among them, represents the observation angle at the position of the sampling point x, z(x) represents the feature vector at the sampling point x, Ω represents the entire voxel space included in the sparse voxel octree, θ represents the coupled neural network parameter, and i represents the lidar intensity predicted by the coupled neural network.

3. The method for dense measurement of lidar based on enhanced imagination computing according to claim 1 or 2, characterized in that: The coupled neural network F θ First, the feature vector z(x) at the sampling point x position is input into an MLP neural network to obtain the predicted SDF and the state parameters ξ of the network hidden layer. Then, the state parameters ξ and the observation angle are input into another MLP neural network together to obtain the radiation intensity value related to the observation angle ​ 4. The method for dense measurement of lidar based on enhanced imagination computing according to claim 1 or 2, characterized in that: By minimizing the SDF loss to optimize the predicted value SDF(x) of the output of the coupled neural network, minimizing the difference between the reconstructed scene and the rendered 3D model: Where, S represents the surface of the predicted scenario, and Ω\S represents the positions in the entire voxel space that do not contain the surface; represents the gradient at the predicted surface x, represents the ground truth normal vector; f(x; θ) represents the SDF value at the position of the sampling point x predicted by the multi-layer perceptron, θ represents the parameters of the coupled neural network, and SDF(x) refers to the ground truth SDF at the position of the sampling point x; U(x, δ) represents the δ neighborhood of x, which contains other sampling points whose distance from the sampling point x is less than the distance threshold δ; σ and λ are hyperparameters, is the eikonal equation.

5. The method for dense measurement of lidar based on enhanced imagination computing according to claim 1, characterized in that: Use the intensity loss to minimize the intensity residual. According to the Gaussian distribution model of the distance field, the radiation information obtained at multiple sampling points is weighted and accumulated to predict the intensity value of the pixel When there is a corresponding measured true value for the predicted pixel, calculate the intensity loss function constructed with the goal of minimizing the intensity residual: Among them, represents the intensity value predicted by the coupled neural network for the pixel to be generated, i represents the intensity value of the original lidar sampling point, and S represents the surface of the predicted scene.

6. The method for dense measurement of lidar based on enhanced imagination computing according to claim 1, characterized in that: Use a feature extractor to extract the semantic features of the original depth image and intensity image obtained by the lidar, as well as the depth image and intensity image optimized and completed by the network, and minimize the difference in semantic consistency features by minimizing the difference between the semantic features of the original image and the image optimized and completed by the network.

Citation Information

Patent Citations

  • Deep learning-based image laser data fusion method for building reconstruction

    CN115423978A

  • Target intelligent positioning control system and method based on infrared imaging, laser radar and sound directional detection

    CN116339337A