Laser radar high-precision surveying and mapping method and system based on multi-source data fusion
Through the multi-source fusion processing of point cloud, image and inertial navigation data, the data loss and inaccuracy of lidar in complex environments are solved, and high-precision lidar surveying and mapping effects are achieved.
Patent Information
- Application Number
- CN202510726775.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-03
AI Technical Summary
Single lidar data is susceptible to occlusion and reflection in complex environments, resulting in missing or inaccurate data. The existing surveying and mapping methods have not fully utilized the advantages of multi-source data, making it difficult to meet the needs of high-precision surveying and mapping.
By acquiring point cloud data, image data and inertial navigation data, denoising, downsampling, geometric correction, error correction and registration processing are performed, a data fusion model is built, and multi-source data fusion is used to fusion using Vision Transformer, 3D sparse convolution and graph attention network to output high-precision fusion results.
It improves the accuracy and reliability of lidar surveying and mapping, and can accurately obtain terrain information in complex environments to achieve high-precision surveying and mapping.
Smart Images

Figure CN120259131A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of surveying and mapping technology, and particularly to a high-precision lidar surveying and mapping method and system based on multi-source data fusion. Background Art
[0002] As an efficient means of spatial data acquisition, lidar surveying and mapping technology has been widely used in fields such as topographic surveying, urban modeling, and intelligent transportation. However, single lidar data has certain limitations. For example, in complex environments, lidar may be affected by factors such as occlusion and reflection, resulting in missing or inaccurate data. In addition, lidar data itself has weak expression ability for information such as the texture and color of objects. At present, although some existing surveying and mapping methods attempt to introduce other data sources for assistance, there are still deficiencies in the way and accuracy of data fusion, and the advantages of multi-source data cannot be fully utilized, making it difficult to meet the requirements of high-precision surveying and mapping. Summary of the Invention
[0003] The purpose of the present invention is to solve the above problems and design a high-precision lidar surveying and mapping method and system based on multi-source data fusion.
[0004] The first aspect of the present invention provides a high-precision lidar surveying and mapping method based on multi-source data fusion, and the method includes the following steps: Obtain the point cloud data and influence data of the target area, and record the attitude and position information during the data acquisition process through an inertial navigation device; Perform denoising and downsampling processing on the point cloud data, perform geometric correction and enhancement processing on the image data, perform error correction on the inertial navigation data, and register the point cloud data with the image data and inertial navigation data respectively; Construct a data fusion model, perform fusion processing on the registered point cloud data, image data and inertial navigation data through the data fusion model, output the fusion result, and perform topographic surveying and analysis according to the fusion result.
[0005] Optionally, in the first implementation manner of the first aspect of the present invention, the performing denoising and downsampling processing on the point cloud data, performing geometric correction and enhancement processing on the image data, and performing error correction on the inertial navigation data includes: Determine the neighborhood range of each point in the point cloud data, and count the distances from all points in the neighborhood of each point in the point cloud data to this point to obtain the preliminarily denoised point cloud data; Determine the voxel grid size according to the spatial distribution range of the preliminarily denoised point cloud data and the expected downsampling degree, and remove the original points in each voxel grid unit and retain the centroid points to complete the downsampling processing; Obtain control points on the image data, establish a transformation model, substitute the image coordinates and actual geographical coordinates of the control points into the transformation model and calculate the transformation parameters, and transform the coordinates of all points on the image data according to the transformation parameters to complete the geometric correction process; Obtain the grayscale histogram of the image data, calculate the cumulative distribution function, and remap the grayscale values according to the cumulative distribution function to complete the histogram equalization enhancement; According to the statistical characteristics of the inertial navigation system noise and the observation noise, use the Kalman filtering algorithm to perform real-time filtering on the inertial navigation data, remove the noise and drift error, and complete the error correction process.
[0006] Optionally, in the second implementation manner of the first aspect of the present invention, the registration of the point cloud data with the image data and the inertial navigation data respectively includes: Project the point cloud data onto the image plane and construct a difference of Gaussian pyramid in the projection area; Find the extreme points in each layer of the difference of Gaussian pyramid, and detect the respective feature points in the image data and the projected point cloud data; Taking the feature point as the center, statistically calculate the gradient direction histogram of the pixel points in the neighborhood, and assign the main direction to the feature point according to the peak direction of the histogram; Taking a fixed-size neighborhood window with each feature point as the center, calculate the gradient amplitude and direction of the pixel points in the neighborhood to form a feature descriptor; Calculate the Euclidean distance between the feature descriptor of a feature point in the image data and all the feature descriptors of the feature points in the point cloud projection area, find the two closest point cloud feature points, and record them as the nearest neighbor point and the second nearest neighbor point respectively; Calculate the distance ratio between the nearest neighbor point and the second nearest neighbor point. If the ratio is less than the set threshold, it is considered that the feature point in the image matches successfully with the nearest neighbor point in the point cloud projection area, where the threshold is 0.8; Summarize the image and point cloud feature point pairs obtained through feature matching to form a set of matching point pairs, and realize the registration of the point cloud data and the image data based on the least squares method.
[0007] Optionally, in the third implementation manner of the first aspect of the present invention, the registration of the point cloud data with the image data and the inertial navigation data respectively further includes: Convert the point cloud data from the lidar device coordinate system to the geographical coordinate system; Select the initial corresponding point pairs, calculate the centroids of the point cloud point sets and the centroids of the inertial navigation position point sets in the corresponding point pairs respectively to obtain the centroid-removed coordinates; Calculate the covariance matrix according to the centroid-removed coordinates, decompose the covariance matrix by the singular value decomposition method to obtain the optimal rotation matrix and translation vector; Rotate and update the points in the point cloud data according to the optimal rotation moment and translation vector, repeat the iterative process, and calculate the average distance between corresponding point pairs after each iteration to complete the registration of the point cloud data and the inertial navigation data.
[0008] Optionally, in the fourth implementation manner of the first aspect of the present invention, the constructing the data fusion model includes: Obtain historical multi-source data of different sensors, and determine the basic architecture parameters of the Vision Transformer based on the historical multi-source data; Construct a 3D sparse convolutional layer, and determine the input and output channels of the sparse convolution according to the point cloud feature dimension in the historical multi-source data; Use the points in the historical multi-source data as nodes, determine the connection relationship of the edges according to the spatial distance between the points, construct the historical multi-source data into a graph structure, and apply a graph attention network to the constructed graph; Integrate the Vision Transformer, 3D sparse convolution, and graph attention network to obtain a multi-modal search space; Use the mean square error after multi-source data fusion to measure the deviation between the fused data and the true value to determine the objective function; Input the historical multi-source data to perform forward propagation according to the network architecture represented by the current parameters, calculate the objective function value, and use the gradient-based optimization algorithm to backpropagate and update the network structure parameter vector according to the objective function value; After multiple rounds of iterative search, find the network structure parameters that make the objective function value optimal, generate the optimal network structure, and obtain the fusion network; Apply the hyperparameters adjusted by the double-layer optimization framework to the fusion network to obtain the final data fusion model.
[0009] Optionally, in the fifth implementation manner of the first aspect of the present invention, the fusing the registered point cloud data, image data, and inertial navigation data through the data fusion model, outputting the fusion result, and performing topographic mapping and analysis according to the fusion result includes: Collect the registered point cloud data, image data, and inertial navigation data, convert the data of different data sources into a unified data format, and input them into the data fusion model; Segment the image into patches of a fixed size, perform linear embedding on each patch to obtain an embedding vector, and extract the global and local features of the image through a multi-layer multi-head self-attention mechanism and a feed-forward neural network; Extract the geometric features and spatial features of the point cloud data through a multi-layer 3D sparse convolutional layer; Use the graph attention network to calculate the feature representation of each node through the attention mechanism, and extract the dynamic features and correlation features of the inertial navigation data; Fuse the image features, point cloud features, and inertial navigation features, and obtain the final fusion result through the non-linear transformation and feature mapping of the data fusion model; Extract terrain-related data from the fusion result to construct a digital elevation model, and perform terrain parameter calculation and terrain analysis based on the digital elevation model.
[0010] The second aspect of the present invention provides a high-precision lidar mapping system based on multi-source data fusion. The system includes: A data acquisition module for acquiring the point cloud data and image data of the target area, and recording the attitude and position information during the data acquisition process through an inertial navigation device; A data processing module for denoising and downsampling the point cloud data, geometrically correcting and enhancing the image data, correcting the error of the inertial navigation data, and registering the point cloud data with the image data and the inertial navigation data respectively; A terrain mapping module for constructing a data fusion model, performing fusion processing on the registered point cloud data, image data, and inertial navigation data through the data fusion model, outputting a fusion result, and performing terrain mapping and analysis according to the fusion result.
[0011] Optionally, in the first implementation manner of the second aspect of the present invention, the data processing module includes: A statistical sub-module for determining the neighborhood range of each point in the point cloud data, and statistically calculating the distance from all points in the neighborhood of each point in the point cloud data to this point, to obtain the preliminarily denoised point cloud data; A division sub-module for determining the voxel grid size according to the spatial distribution range of the preliminarily denoised point cloud data and the desired downsampling degree, and removing the original points in each voxel grid unit and retaining the centroid points to complete the downsampling process; An establishment sub-module for obtaining the control points on the image data, establishing a transformation model, substituting the image coordinates and actual geographical coordinates of the control points into the transformation model and calculating the transformation parameters, and transforming the coordinates of all points on the image data according to the transformation parameters to complete the geometric correction process; A calculation sub-module for obtaining the gray histogram of the image data, calculating the cumulative distribution function, and remapping the gray values according to the cumulative distribution function to complete the histogram equalization enhancement; A filtering sub-module for performing real-time filtering on the inertial navigation data using the Kalman filtering algorithm according to the statistical characteristics of the inertial navigation system noise and the observation noise, removing the noise and drift error, and completing the error correction process.
[0012] Optionally, in the second implementation manner of the second aspect of the present invention, the terrain mapping module includes: A determination sub-module, configured to obtain historical multi-source data of different sensors, and determine basic architecture parameters of a Vision Transformer based on the historical multi-source data; A setting sub-module, configured to construct a 3D sparse convolutional layer, and determine the number of input and output channels of the sparse convolution according to the point cloud feature dimension in the historical multi-source data; An adjustment sub-module, configured to use the points in the historical multi-source data as nodes, determine the connection relationship of edges according to the spatial distance between the points, construct the historical multi-source data into a graph structure, and apply a graph attention network to the constructed graph; An integration sub-module, configured to integrate the Vision Transformer, 3D sparse convolution, and graph attention network to obtain a multi-modal search space; An allocation sub-module, configured to use the mean square error after multi-source data fusion to measure the deviation between the fused data and the true value to determine an objective function; A propagation sub-module, configured to input the historical multi-source data to perform forward propagation according to the network architecture represented by the current parameterization, calculate the objective function value, and update the network structure parameter vector by backpropagation using a gradient-based optimization algorithm according to the objective function value; A generation sub-module, configured to, through multiple rounds of iterative search, find the network structure parameters that optimize the objective function value, generate an optimal network structure, and obtain a fusion network; An application sub-module, configured to apply the hyperparameters adjusted by the double-layer optimization framework to the fusion network to obtain a final data fusion model.
[0013] A second aspect of the present invention provides a high-precision lidar mapping system based on multi-source data fusion, and the system includes: A data acquisition module, configured to acquire point cloud data and influence data of a target area, and record the attitude and position information during the data acquisition process through an inertial navigation device; A data processing module, configured to perform denoising and downsampling processing on the point cloud data, perform geometric correction and enhancement processing on the image data, perform error correction on the inertial navigation data, and register the point cloud data with the image data and the inertial navigation data respectively; A terrain mapping module, configured to construct a data fusion model, perform fusion processing on the registered point cloud data, image data, and inertial navigation data through the data fusion model, output a fusion result, and perform terrain mapping and analysis according to the fusion result.
[0014] Optionally, in a first implementation manner of the second aspect of the present invention, the data processing module includes: A statistics sub-module, configured to determine the neighborhood range of each point in the point cloud data, and statistically calculate the distance from all points in the neighborhood of each point in the point cloud data to the point, to obtain the point cloud data after preliminary denoising; A sub-module for dividing is used to determine the voxel grid size according to the spatial distribution range of the preliminarily denoised point cloud data and the desired downsampling degree, remove the original points within each voxel grid unit and retain the centroid points to complete the downsampling process; A sub-module for establishing is used to obtain control points on the image data, establish a transformation model, substitute the image coordinates and actual geographical coordinates of the control points into the transformation model and calculate the transformation parameters, and transform the coordinates of all points on the image data according to the transformation parameters to complete the geometric correction process; A calculation sub-module is used to obtain the gray histogram of the image data, calculate the cumulative distribution function, and remap the gray values according to the cumulative distribution function to complete the histogram equalization enhancement; A filtering sub-module is used to perform real-time filtering on the inertial navigation data using the Kalman filtering algorithm according to the statistical characteristics of the inertial navigation system noise and the observation noise, remove the noise and drift errors, and complete the error correction process.
[0015] Optionally, in the second implementation manner of the second aspect of the present invention, the topographic surveying and mapping module includes: A determination sub-module is used to obtain the historical multi-source data of different sensors and determine the basic architecture parameters of the Vision Transformer based on the historical multi-source data; A setting sub-module is used to construct a 3D sparse convolution layer and determine the input and output channels of the sparse convolution according to the point cloud feature dimension in the historical multi-source data; An adjustment sub-module is used to take the points in the historical multi-source data as nodes, determine the connection relationship of the edges according to the spatial distance between the points, construct the historical multi-source data into a graph structure, and apply a graph attention network to the constructed graph; An integration sub-module is used to integrate the Vision Transformer, 3D sparse convolution, and graph attention network to obtain a multi-modal search space; An assignment sub-module is used to determine the objective function by measuring the deviation between the fused data and the true value using the mean square error after multi-source data fusion; A propagation sub-module is used to perform forward propagation on the input historical multi-source data according to the network architecture represented by the current parameters, calculate the objective function value, and update the network structure parameter vector by backpropagation using a gradient-based optimization algorithm according to the objective function value; A generation sub-module is used to find the network structure parameters that make the objective function value optimal through multiple rounds of iterative search, generate the optimal network structure, and obtain the fusion network; An application sub-module is used to apply the hyperparameters adjusted by the double-layer optimization framework to the fusion network to obtain the final data fusion model. Description of the Drawings
[0016] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the following detailed description of the preferred embodiments. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present invention.
[0017] Figure 1 Schematic diagram of the first embodiment of the lidar high-precision mapping method based on multi-source data fusion provided by the embodiment of the present invention; Figure 2 Schematic diagram of the second embodiment of the lidar high-precision mapping method based on multi-source data fusion provided by the embodiment of the present invention; Figure 3 Schematic diagram of the third embodiment of the lidar high-precision mapping method based on multi-source data fusion provided by the embodiment of the present invention; Figure 4 Schematic diagram of the structure of the lidar high-precision mapping system based on multi-source data fusion provided by the embodiment of the present invention. Detailed implementation manners
[0018] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0019] For ease of understanding, the following describes the specific process of the embodiment of the present invention. Please refer to Figure 1 Schematic diagram of the first embodiment of the lidar high-precision mapping method based on multi-source data fusion provided by the embodiment of the present invention. The method specifically includes the following steps: Step 101, obtain the point cloud data and influence data of the target area, and record the attitude and position information during the data acquisition process through an inertial navigation device; In this embodiment, a lidar device is used to obtain the point cloud data of the target area, an image acquisition device is used to collect the image data of the target area, and an inertial navigation device is used to record the attitude and position information of the lidar and the image acquisition device during the data acquisition process.
[0020] Step 102: Denoise and downsample the point cloud data, perform geometric correction and enhancement on the image data, correct the errors of the inertial navigation data, and register the point cloud data with the image data and the inertial navigation data respectively; In this embodiment, the image data is convolved with Gaussian kernel functions of different scales to generate a series of images at different scales. The difference between adjacent scale images is used to obtain the difference of Gaussian images and construct a difference of Gaussian pyramid. The point cloud data is projected onto the image plane, and a difference of Gaussian pyramid is constructed within the projection area. At each layer of the difference of Gaussian pyramid, by comparing the gray value changes of pixel points with their neighboring pixel points, extreme points are searched. Extreme points are searched at each layer of the difference of Gaussian pyramid, and respective feature points are detected in the image data and the projected point cloud data. Centering on the feature points, the gradient direction histogram of pixel points in the neighborhood is statistically calculated, and the main direction is assigned to the feature points according to the peak direction of the histogram. For each feature point in the image data and the point cloud projection area, a neighborhood window of a fixed size is taken centered on it. For each feature point, a neighborhood window of a fixed size is taken centered on it, and the gradient amplitude and direction of pixel points in the neighborhood are calculated to form a feature descriptor. The nearest neighbor distance ratio algorithm is adopted to calculate the Euclidean distance between the feature descriptor of a feature point in the image data and the feature descriptors of all feature points in the point cloud projection area, and the two nearest point cloud feature points are found, which are respectively recorded as the nearest neighbor point and the next nearest neighbor point. The distance ratio between the nearest neighbor point and the next nearest neighbor point is calculated. If the ratio is less than the set threshold, it is considered that the feature point in the image matches successfully with the nearest neighbor point in the point cloud projection area, where the threshold is 0.8. The pairs of image and point cloud feature points obtained through feature matching are summarized to form a set of matching point pairs, and the registration of the point cloud data and the image data is realized based on the least square method.
[0021] In this embodiment, a rotation matrix is constructed according to the attitude information of the lidar device recorded in the inertial navigation data. Combining with the position information of the lidar device recorded in the inertial navigation data, the point cloud data is transformed from the lidar device coordinate system to the geographic coordinate system. Between the point cloud data after transformation to the geographic coordinate system and the position information recorded in the inertial navigation data, point pairs within the distance threshold range are selected as the initial corresponding point pairs. The centroids of the point cloud point set and the inertial navigation position point set in the corresponding point pairs are calculated respectively. The point cloud points and the inertial navigation position points in the corresponding point pairs are respectively subtracted from their respective centroid coordinates to obtain the centroid-removed coordinates. The covariance matrix is calculated according to the centroid-removed coordinates, and the covariance matrix is decomposed by the singular value decomposition method to obtain the optimal rotation matrix and translation vector. The points in the point cloud data are rotated and translated according to the optimal rotation matrix and translation vector, and the iterative process is repeated. The average distance between the corresponding point pairs is calculated after each iteration. When the maximum number of iterations is reached, the iteration is stopped, and the registration of the point cloud data and the inertial navigation data is completed.
[0022] Step 103: Construct a data fusion model. Use the data fusion model to perform fusion processing on the registered point cloud data, image data, and inertial navigation data, and output the fusion result. Perform topographic mapping and analysis based on the fusion result. In this embodiment, historical multi-source data of different sensors is obtained, and the basic architecture parameters of Vision Transformer are determined based on the historical multi-source data, including at least the patch size, embedding dimension, number of heads of the multi-head attention mechanism, and number of layers; the size and stride of the sparse convolution kernel are set according to the spatial distribution range of the point cloud in the historical multi-source data to determine the convolution operation method in the three-dimensional space to construct a 3D sparse convolution layer, and the input and output channels of the sparse convolution are determined according to the point cloud feature dimension in the historical multi-source data; taking the points in the historical multi-source data as nodes, the connection relationship of the edges is determined according to the spatial distance between the points, and the historical multi-source data is constructed into a graph structure. Apply a graph attention network to the constructed graph to enable the nodes to adaptively adjust their own feature weights according to the information of the neighboring nodes; integrate Vision Transformer, 3D sparse convolution, and graph attention network to obtain a multi-modal search space; assign a set of learnable parameter vectors to each possible network architecture in the search space, and use the mean square error after multi-source data fusion to measure the deviation between the fused data and the true value to determine the objective function; input the historical multi-source data and perform forward propagation according to the network architecture represented by the current parameterization, calculate the objective function value, and use the gradient-based optimization algorithm to backpropagate and update the network structure parameter vector according to the objective function value; after multiple rounds of iterative search, find the network structure parameters that make the objective function value optimal, so as to automatically generate the optimal network structure and obtain the fusion network.
[0023] In this embodiment, upper layer: construct a multi-dimensional value space of all global hyperparameters to be optimized in the fusion network. Randomly select, through Latin hypercube sampling in the multi-dimensional value space, 20 sets of hyperparameter combinations, for example. Use these hyperparameter combinations to train the fusion network, and record the corresponding performance metric values as the initial data for constructing the Gaussian process model. In each iteration, use the Gaussian process model to predict the objective function values of different points in the hyperparameter space. According to the prediction results, use the acquisition function to determine the combination that can maximize the improvement of the current best performance. Retrain the fusion network with the selected hyperparameter combination, update the performance metric values, and feedback the new hyperparameter combination and the corresponding performance metric values to the Gaussian process model to update the parameters of the Gaussian process model and find the optimal global hyperparameter combination; lower layer: construct a sensor-level parameter space for the parameters of different sensors. In the sensor-level parameter space, randomly generate 50 individuals, and each individual represents a set of sensor parameter value combinations. These individuals form the initial population. At the same time, set the covariance matrix of the initial population. Apply the sensor parameters corresponding to each individual in the population to the optimal global hyperparameter combination generated by the upper layer, and calculate the performance metric of the model. Use this performance metric as the fitness value of the individual. Based on the fitness values of the individuals, perform selection, recombination, and mutation operations using the CMA-ES evolutionary algorithm. Update the covariance matrix according to the fitness changes of the individuals during the evolution process and the parameter distribution of the newly generated individuals. After multiple rounds of evolution, find the optimal sensor-level parameter combination; apply the hyperparameters adjusted by the double-layer optimization framework to the fusion network to obtain the final data fusion model.
[0024] Please refer to Figure 2 , the schematic diagram of the second embodiment of the lidar high-precision mapping method based on multi-source data fusion provided by the embodiment of the present invention. The method includes: Step 201, determine the neighborhood range of each point in the point cloud data, and count the distances from all points in the neighborhood of each point in the point cloud data to this point to obtain the preliminarily denoised point cloud data; In this embodiment, if the distance from this point to the neighborhood mean point is greater than the distance threshold, it is determined as a noise point and removed from the point cloud data to obtain the preliminarily denoised point cloud data.
[0025] Step 202, determine the voxel grid size according to the spatial distribution range of the preliminarily denoised point cloud data and the expected downsampling degree, and remove the original points in each voxel grid unit and retain the centroid point to complete the downsampling process; In this embodiment, divide the three-dimensional space according to the voxel grid, assign each point in the preliminarily denoised point cloud to the corresponding voxel grid unit, calculate the centroid coordinates of all points in each voxel grid unit, represent all points in this voxel grid unit with the centroid coordinates, remove the original points in each voxel grid unit and retain the centroid point to complete the downsampling process.
[0026] Step 203: Obtain the control points on the image data, establish a transformation model, substitute the image coordinates and actual geographic coordinates of the control points into the transformation model and calculate the transformation parameters, and transform the coordinates of all points on the image data according to the transformation parameters to complete the geometric correction process; Step 204: Obtain the gray histogram of the image data, calculate the cumulative distribution function, and remap the gray values according to the cumulative distribution function; In this embodiment, traverse each pixel in the image data, replace its original gray value with a new gray value according to the mapping relationship, and complete the histogram equalization enhancement.
[0027] Step 205: According to the statistical characteristics of the inertial navigation system noise and the observation noise, use the Kalman filtering algorithm to perform real-time filtering on the inertial navigation data, remove the noise and drift error, and complete the error correction process.
[0028] Please refer to Figure 3 , the schematic diagram of the third embodiment of the lidar high-precision mapping method based on multi-source data fusion provided by the embodiment of the present invention, the method includes: Step 301: Collect the registered point cloud data, image data and inertial navigation data, convert the data of different data sources into a unified data format, and input it into the data fusion model; Step 302: Segment the image into fixed-size patches, perform linear embedding on each patch to obtain embedding vectors, and extract the global and local features of the image through multiple layers of multi-head self-attention mechanisms and feed-forward neural networks; In this embodiment, the registered image data is input into the Vision Transformer module.
[0029] Step 303: Extract the geometric features and spatial features of the point cloud data through multiple layers of 3D sparse convolutional layers; In this embodiment, use the 3D sparse convolution module to process the registered point cloud data, and construct a sparse convolution kernel according to the spatial distribution of the point cloud data.
[0030] Step 304: Use the graph attention network to calculate the feature representation of each node through the attention mechanism, and extract the dynamic features and correlation features of the inertial navigation data; In this embodiment, the inertial navigation data is constructed into a graph structure, with different parameters in the inertial navigation data as nodes and the relationships between nodes as edges.
[0031] Step 305: Fuse the image features, point cloud features and inertial navigation features, and obtain the final fusion result through the non-linear transformation and feature mapping of the data fusion model; Step 306: Extract terrain-related data from the fusion result to construct a digital elevation model, and perform terrain parameter calculation and terrain analysis based on the digital elevation model.
[0032] Please refer to Figure 4 , the structural schematic diagram of the lidar high-precision mapping system based on multi-source data fusion provided by the embodiment of the present invention. The system includes: A data acquisition module, configured to acquire point cloud data and influence data of a target area, and record the attitude and position information during the data acquisition process through an inertial navigation device; A data processing module, configured to perform denoising and downsampling processing on the point cloud data, perform geometric correction and enhancement processing on the image data, perform error correction on the inertial navigation data, and register the point cloud data with the image data and the inertial navigation data respectively; A terrain mapping module, configured to construct a data fusion model, perform fusion processing on the registered point cloud data, image data and inertial navigation data through the data fusion model, output a fusion result, and perform terrain mapping and analysis according to the fusion result.
[0033] In this embodiment, the data processing module includes: A statistics sub-module, configured to determine the neighborhood range of each point in the point cloud data, and count the distances from all points in the neighborhood of each point in the point cloud data to this point, so as to obtain the point cloud data after preliminary denoising; A division sub-module, configured to determine the voxel grid size according to the spatial distribution range of the point cloud data after preliminary denoising and the expected downsampling degree, and remove the original points in each voxel grid unit and retain the centroid points to complete the downsampling processing; An establishment sub-module, configured to obtain control points on the image data, establish a transformation model, substitute the image coordinates and actual geographical coordinates of the control points into the transformation model and calculate the transformation parameters, and transform the coordinates of all points on the image data according to the transformation parameters to complete the geometric correction processing; A calculation sub-module, configured to obtain the gray histogram of the image data, calculate the cumulative distribution function, and remap the gray values according to the cumulative distribution function to complete the histogram equalization enhancement; A filtering sub-module, configured to perform real-time filtering on the inertial navigation data using the Kalman filtering algorithm according to the statistical characteristics of the inertial navigation system noise and the observation noise, remove the noise and drift error, and complete the error correction processing.
[0034] In this embodiment, the terrain mapping module includes: A determination sub-module, configured to obtain historical multi-source data of different sensors, and determine the basic architecture parameters of the Vision Transformer based on the historical multi-source data; A setting sub-module for constructing a 3D sparse convolutional layer, and determining the number of input and output channels of the sparse convolution according to the point cloud feature dimension in historical multi-source data; An adjustment sub-module for using the points in historical multi-source data as nodes, determining the connection relationship of edges according to the spatial distance between points, constructing the historical multi-source data into a graph structure, and applying a graph attention network to the constructed graph; An integration sub-module for integrating Vision Transformer, 3D sparse convolution, and graph attention network to obtain a multi-modal search space; An allocation sub-module for determining an objective function by using the mean square error after multi-source data fusion to measure the deviation between the fused data and the true value; A propagation sub-module for inputting historical multi-source data to perform forward propagation according to the network architecture represented by the current parameterization, calculating the objective function value, and using a gradient-based optimization algorithm to perform backpropagation to update the network structure parameter vector according to the objective function value; A generation sub-module for finding the network structure parameters that optimize the objective function value after multiple rounds of iterative search, generating an optimal network structure, and obtaining a fusion network; An application sub-module for applying the hyperparameters adjusted by the double-layer optimization framework to the fusion network to obtain a final data fusion model.
[0035] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A high-precision lidar mapping method based on multi-source data fusion, characterized in that, The method includes the following steps: Obtain the point cloud data and image data of the target area, and record the attitude and position information during the data acquisition process through an inertial navigation device; Perform denoising and downsampling processing on the point cloud data, perform geometric correction and enhancement processing on the image data, perform error correction on the inertial navigation data, and register the point cloud data with the image data and the inertial navigation data respectively; Construct a data fusion model, perform fusion processing on the registered point cloud data, image data and inertial navigation data through the data fusion model, output the fusion result, and perform topographic mapping and analysis according to the fusion result.
2. The high-precision lidar mapping method based on multi-source data fusion according to claim 1, characterized in that The performing denoising and downsampling processing on the point cloud data, performing geometric correction and enhancement processing on the image data, and performing error correction on the inertial navigation data includes: Determine the neighborhood range of each point in the point cloud data, and count the distances from all points in the neighborhood of each point in the point cloud data to this point to obtain the preliminarily denoised point cloud data; Determine the voxel grid size according to the spatial distribution range of the preliminarily denoised point cloud data and the desired downsampling degree, and remove the original points in each voxel grid unit and retain the centroid points to complete the downsampling processing; Obtain the control points on the image data, establish a transformation model, substitute the image coordinates and actual geographical coordinates of the control points into the transformation model and calculate the transformation parameters, and transform the coordinates of all points on the image data according to the transformation parameters to complete the geometric correction processing; Obtain the gray histogram of the image data, calculate the cumulative distribution function, and remap the gray values according to the cumulative distribution function to complete the histogram equalization enhancement; According to the statistical characteristics of the inertial navigation system noise and the observation noise, use the Kalman filtering algorithm to perform real-time filtering on the inertial navigation data to remove noise and drift errors and complete the error correction processing.
3. The high-precision lidar mapping method based on multi-source data fusion according to claim 1, characterized in that The registering the point cloud data with the image data and the inertial navigation data respectively includes: Project the point cloud data onto the image plane, and construct a difference of Gaussian pyramid in the projection area; Find the extreme points in each layer of the difference of Gaussian pyramid, and detect the respective feature points in the image data and the projected point cloud data; Taking the feature point as the center, count the gradient direction histogram of the pixel points in the neighborhood, and assign the main direction to the feature point according to the peak direction of the histogram; Take a neighborhood window of a fixed size centered on each feature point, calculate the gradient amplitude and direction of the pixel points in the neighborhood to form a feature descriptor; Calculate the Euclidean distance between the feature descriptor of a feature point in the image data and the feature descriptors of all feature points in the point cloud projection area, and find the two closest point cloud feature points, which are respectively recorded as the nearest neighbor point and the second nearest neighbor point; Calculate the distance ratio between the nearest neighbor point and the second nearest neighbor point. If the ratio is less than the set threshold, it is considered that the feature point in the image matches successfully with the nearest neighbor point in the point cloud projection area, where the threshold is 0.8; Summarize the image and point cloud feature point pairs obtained through feature matching to form a set of matching point pairs, and realize the registration of the point cloud data and the image data based on the least squares method.
4. The high-precision lidar mapping method based on multi-source data fusion according to claim 1, wherein The registering the point cloud data with the image data and the inertial navigation data respectively further includes: Convert the point cloud data from the lidar device coordinate system to the geographic coordinate system; Select the initial corresponding point pairs, calculate the centroids of the point cloud point sets and the centroids of the inertial navigation position point sets for the corresponding point pairs respectively, and obtain the centroid-removed coordinates; Calculate the covariance matrix based on the centroid-removed coordinates, decompose the covariance matrix by the singular value decomposition method, and obtain the optimal rotation matrix and translation vector; Rotate and translate the points in the point cloud data according to the optimal rotation matrix and translation vector, repeat the iterative process, and calculate the average distance between the corresponding point pairs after each iteration to complete the registration of the point cloud data and the inertial navigation data.
5. The high-precision lidar mapping method based on multi-source data fusion according to claim 1, characterized in that The construction of the data fusion model includes: Obtain the historical multi-source data of different sensors, and determine the basic architecture parameters of the Vision Transformer based on the historical multi-source data; Construct a 3D sparse convolutional layer, and determine the number of input and output channels of the sparse convolution according to the point cloud feature dimension in the historical multi-source data; Use the points in the historical multi-source data as nodes, determine the connection relationship of the edges according to the spatial distance between the points, construct the historical multi-source data into a graph structure, and apply a graph attention network to the constructed graph; Integrate the Vision Transformer, 3D sparse convolution, and graph attention network to obtain a multi-modal search space; Use the mean square error after multi-source data fusion to measure the deviation between the fused data and the true value to determine the objective function; Input the historical multi-source data to perform forward propagation according to the network architecture represented by the current parameters, calculate the objective function value, and use the gradient-based optimization algorithm to backpropagate and update the network structure parameter vector according to the objective function value; After multiple rounds of iterative search, find the network structure parameters that make the objective function value optimal, generate the optimal network structure, and obtain the fusion network; Apply the hyperparameters adjusted by the double-layer optimization framework to the fusion network to obtain the final data fusion model.
6. The high-precision lidar mapping method based on multi-source data fusion according to claim 1, characterized in that, The fusion processing of the registered point cloud data, image data, and inertial navigation data through the data fusion model, and output the fusion result. According to the fusion result, perform topographic mapping and analysis, including: Collect the registered point cloud data, image data, and inertial navigation data, convert the data of different data sources into a unified data format, and input them into the data fusion model; Segment the image into fixed-size patches, perform linear embedding on each patch to obtain embedding vectors, and extract the global and local features of the image through multiple layers of multi-head self-attention mechanisms and feed-forward neural networks; Extract the geometric features and spatial features of the point cloud data through multiple layers of 3D sparse convolutional layers; Use the graph attention network to calculate the feature representation of each node through the attention mechanism, and extract the dynamic features and correlation features of the inertial navigation data; Fuse the image features, point cloud features, and inertial navigation features, and obtain the final fusion result through the non-linear transformation and feature mapping of the data fusion model; Extract terrain-related data from the fusion result to construct a digital elevation model, and perform terrain parameter calculation and terrain analysis based on the digital elevation model.
7. A high-precision lidar mapping system based on multi-source data fusion, characterized in that, The system includes: A data acquisition module for acquiring the point cloud data and image data of the target area, and recording the attitude and position information during the data acquisition process through an inertial navigation device; A data processing module, which is used to perform denoising and downsampling processing on point cloud data, perform geometric correction and enhancement processing on image data, perform error correction on inertial navigation data, and register the point cloud data with the image data and inertial navigation data respectively; A topographic surveying and mapping module, which is used to construct a data fusion model, perform fusion processing on the registered point cloud data, image data and inertial navigation data through the data fusion model, output a fusion result, and perform topographic surveying and analysis according to the fusion result.
8. The lidar high-precision mapping system based on multi-source data fusion according to claim 7, wherein The data processing module includes: A statistical sub-module, which is used to determine the neighborhood range of each point in the point cloud data, and statistically calculate the distances from all points in the neighborhood of each point in the point cloud data to this point, so as to obtain the point cloud data after preliminary denoising; A division sub-module, which is used to determine the voxel grid size according to the spatial distribution range of the point cloud data after preliminary denoising and the expected downsampling degree, remove the original points in each voxel grid unit and retain the center of gravity points to complete the downsampling processing; A establishment sub-module, which is used to obtain the control points on the image data, establish a transformation model, substitute the image coordinates and actual geographical coordinates of the control points into the transformation model and calculate the transformation parameters, and transform the coordinates of all points on the image data according to the transformation parameters to complete the geometric correction processing; A calculation sub-module, which is used to obtain the gray histogram of the image data, calculate the cumulative distribution function, and remap the gray values according to the cumulative distribution function to complete the histogram equalization enhancement; A filtering sub-module, which is used to perform real-time filtering on the inertial navigation data using the Kalman filtering algorithm according to the statistical characteristics of the inertial navigation system noise and observation noise, remove the noise and drift error, and complete the error correction processing.
9. The high-precision lidar mapping system based on multi-source data fusion according to claim 7, characterized in that, The topographic surveying and mapping module includes: A determination sub-module, which is used to obtain the historical multi-source data of different sensors, and determine the basic architecture parameters of the Vision Transformer based on the historical multi-source data; A setting sub-module, which is used to construct a 3D sparse convolutional layer, and determine the number of input and output channels of the sparse convolution according to the point cloud feature dimension in the historical multi-source data; An adjustment sub-module, which is used to take the points in the historical multi-source data as nodes, determine the connection relationship of the edges according to the spatial distance between the points, construct the historical multi-source data into a graph structure, and apply the graph attention network to the constructed graph; An integration sub-module, which is used to integrate the Vision Transformer, 3D sparse convolution, and graph attention network to obtain a multi-modal search space; An allocation sub-module, which is used to determine the objective function by using the mean square error after multi-source data fusion to measure the deviation between the fused data and the true value; A propagation sub-module, which is used to input the historical multi-source data to perform forward propagation according to the currently parameterized network architecture, calculate the objective function value, and update the network structure parameter vector by backpropagation using the gradient-based optimization algorithm according to the objective function value; A generation sub-module, which is used to find the network structure parameters that make the objective function value optimal through multiple rounds of iterative search, generate the optimal network structure, and obtain the fusion network; An application sub-module, which is used to apply the hyperparameters adjusted by the double-layer optimization framework to the fusion network to obtain the final data fusion model.
Citation Information
Patent Citations
Vehicle-mounted multi-sensor fusion positioning method and device, chip and terminal
CN114264301A
Radar point cloud data inertia correction method based on data fusion
CN115421125A
Fusion detection method and device based on multivariate data, equipment and storage medium
CN115909815A
Multi-modal fusion sensing method and device of vehicle, vehicle and storage medium
CN116543361A
Vehicle-mounted video monitoring system and method integrating 360-degree all-round viewing perception and multi-mode intelligent analysis
CN119152462A
Cited By
Cultural relic digital surveying and mapping method based on multi-scale point cloud fusion
CN121582320A
Image-control-free oblique photography surveying and mapping method and system fusing air-ground multi-source data
CN122062635A
A method and system for image-controlled oblique photogrammetry that integrates multi-source air and ground data
CN122062635B