Geographic information surveying and mapping method and device based on laser radar, equipment and medium
Through the spatiotemporal calibration and adaptive filtering algorithm of the lidar, multispectral camera and data measurement unit, the problem of target recognition and data synchronization errors of a single lidar in dynamic scenes was solved, and high-precision multi-scale geographic information mapping was achieved.
Patent Information
- Application Number
- CN202510750593.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing technologies, a single lidar has limited target recognition capabilities in dynamic scenarios and lacks data semantic information. Multi-sensor fusion strategies have large data synchronization errors and make it difficult to distinguish between fixed objects and moving targets, resulting in significant errors in surveying and mapping results.
By performing spatiotemporal calibration on the lidar, multispectral camera, and data measurement unit, we obtain spatiotemporally synchronized multi-source data. We use multispectral images and original point clouds to extract dynamic semantic reference points. We combine IMU measurement data and wheel speed meter speed data to perform an adaptive filtering algorithm, perform pose solution, and perform hierarchical modeling and fusion with satellite geographic data.
It improves the temporal and spatial synchronization accuracy of multi-source sensor data, enhances the ability to distinguish between fixed objects and moving targets in dynamic scenes, improves the modeling accuracy and detail expression ability of digital surface models, and enhances the stability of surveying and mapping systems in complex environments.
Smart Images

Figure CN120652485A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of geographic information surveying and mapping technology, and in particular relates to a geographic information surveying and mapping method, device, equipment and medium based on laser radar. Background Art
[0002] As a core means of acquiring modern spatial information, geographic information surveying and mapping technology plays a vital role in areas such as smart cities, disaster monitoring, and autonomous driving. Traditional geographic information surveying and mapping methods primarily rely on single sensors, such as total stations, for data collection, resulting in low efficiency and limited coverage. In recent years, LiDAR technology, with its high-precision three-dimensional modeling capabilities, has gradually become a key tool for geographic information surveying and mapping. However, single LiDAR suffers from limitations such as limited target recognition and lack of data semantics in dynamic scenarios, making it difficult to meet the mapping needs of complex geographic environments. Therefore, existing technologies typically employ multi-sensor fusion strategies, combining LiDAR with cameras and inertial measurement units (IMUs) to achieve more comprehensive and accurate geographic information surveying and mapping. However, while this multi-sensor fusion strategy can fuse sensor data through a loose coupling approach to achieve data complementarity, it fails to fully consider the spatiotemporal correlations between sensors, resulting in significant data synchronization errors and impacting mapping accuracy. Furthermore, in dynamic scenarios, traditional multi-sensor fusion geographic information surveying and mapping methods struggle to effectively distinguish fixed objects from moving targets, making it difficult to accurately extract geographic reference points, resulting in significant errors in mapping results. Summary of the Invention
[0003] Based on this, it is necessary to provide lidar-based geographic information surveying and mapping methods, devices, equipment and media to address the above technical problems, so as to improve the spatiotemporal synchronization accuracy and fusion efficiency of multi-source sensor data, enhance the ability to distinguish between fixed objects and mobile targets in dynamic scenes, enhance the stability of the surveying and mapping system in complex environments, and improve the modeling accuracy and detail expression ability of digital surface models.
[0004] In a first aspect, the present application provides a geographic information surveying and mapping method based on laser radar, comprising:
[0005] Based on the hardware parameters of the LiDAR, multispectral camera, and data measurement unit, the LiDAR, multispectral camera, and data measurement unit deployed on the mobile platform are subjected to spatiotemporal calibration processing to obtain spatiotemporally synchronized raw point clouds, multispectral images, IMU measurement data, and wheel speed meter speed data. The IMU measurement data includes the angular velocity and acceleration of the mobile platform, and the data measurement unit includes an IMU subunit and a wheel speed meter subunit.
[0006] Based on multispectral images and original point clouds, dynamic semantic reference points are extracted through semantic segmentation and 3D association processing, and a dynamic semantic reference point spatiotemporal constraint database is constructed.
[0007] Based on the dynamic semantic reference points, IMU measurement data and wheel speed meter speed data, the pose is solved through the adaptive filtering algorithm, and the platform pose with error suppression is output;
[0008] Based on the platform posture and original point cloud after error suppression, hierarchical modeling is performed to generate a digital surface model, which is then fused with satellite geographic data to output multi-scale geographic information results. The multi-scale geographic information results include a fused digital surface model and a semantic label distribution map enhanced based on terrain features.
[0009] In one embodiment, a spatiotemporal calibration process is performed on the laser radar, multispectral camera, and data measurement unit deployed on a mobile platform based on their hardware parameters to obtain spatiotemporally synchronized raw point clouds, multispectral images, IMU measurement data, and wheel speed meter speed data, including:
[0010] Deploy lidar, multispectral cameras, and data measurement units based on the preset lidar scanning frequency, multispectral camera band parameters, IMU zero bias stability requirements, and wheel speed meter sampling rate;
[0011] According to the checkerboard calibration plate, the laser radar and multispectral camera are aligned to obtain the extrinsic parameter matrix;
[0012] The Allan variance analysis method is used to calibrate the noise parameters of the data measurement unit and construct the error state Kalman filter model;
[0013] Based on the GNSS pulse-second signal, the laser radar, multispectral camera, and data measurement unit are clock-synchronized to obtain a time-synchronized sensor data stream. The sensor data stream includes the initial raw point cloud, initial multispectral image, initial IMU measurement data, and initial wheel speed meter speed data.
[0014] Based on the time-synchronized sensor data stream, dynamic scene test data is generated by collecting and processing dynamic scene motion trajectories, and reprojection error verification processing is performed on the external parameter matrix and error state Kalman filter model to output spatiotemporally synchronized original point clouds, multispectral images, IMU measurement data, and wheel speed meter speed data.
[0015] In one embodiment, dynamic semantic reference points are extracted from multispectral images and original point clouds through semantic segmentation and three-dimensional association processing, and a dynamic semantic reference point spatiotemporal constraint database is constructed, including:
[0016] Based on the RGB and near-infrared band data of the multispectral camera, the multispectral image is processed by histogram equalization and vegetation index calculation to obtain an enhanced multispectral image;
[0017] The enhanced multispectral image is input into the preset semantic segmentation network for pixel-level label recognition processing to obtain the segmentation result, and the segmentation result is processed by morphological closing operation to output the semantic segmentation mask and target bounding box;
[0018] According to the original point cloud area corresponding to the semantic segmentation mask, the DBSCAN algorithm is used to perform Euclidean clustering processing to obtain a set of point cloud clusters associated with the ground objects;
[0019] The minimum bounding box is calculated based on the point cloud cluster set, and the center point coordinates are extracted. Combined with the global navigation satellite system position data of the mobile platform, the center point coordinates are converted and processed to output dynamic semantic reference points. The semantic label, timestamp and coordinate uncertainty of each dynamic semantic reference point are recorded to construct a dynamic semantic reference point spatiotemporal constraint database.
[0020] In one embodiment, the pose solution is performed based on the dynamic semantic reference points, IMU measurement data, and wheel speed meter speed data through an adaptive filtering algorithm to output the platform pose after error suppression, including:
[0021] According to the availability of the global navigation satellite system signal, the global coordinate system or the local coordinate system is selected to initialize the state vector of the error state Kalman filter. The state vector includes the platform position, velocity, attitude angle and IMU zero bias parameters, and the covariance matrix of the initial pose and IMU zero bias parameters is output;
[0022] Based on the angular velocity of the mobile platform in the IMU measurement data, the attitude quaternion is updated through the preset kinematic model, and the platform velocity and position increment are calculated according to the mobile platform acceleration, and the predicted posture and covariance matrix are output;
[0023] Based on the wheel speed meter speed, the coordinates of the dynamic semantic reference point are used as the position observation value, and the wheel speed meter speed is used as the speed observation value. The observation equation is constructed and the Kalman gain is calculated to correct the predicted posture and obtain the corrected platform posture and corrected covariance matrix.
[0024] Based on the corrected platform pose and the corrected covariance matrix, a multi-source constrained optimization objective function is constructed within a preset sliding window using the benchmark observation data, IMU zero bias records, and historical wheel speedometer speed data. This function is based on IMU pre-integration constraints, benchmark reprojection error constraints, and wheel speedometer odometry constraints. The Levenberg-Marquardt algorithm is used to solve the function and output the platform pose with error suppression.
[0025] In one embodiment, the multi-source constrained optimization objective function is:
[0026]
[0027] Among them, x is the set of state variables to be optimized in the preset sliding window, including the platform posture [p k ,q k ]、IMU bias b k and sensor delay τ, is the IMU pre-integration residual of the kth frame, which is calculated from the IMU measurement values and state variables between adjacent frames. is the reprojection residual of the mth dynamic semantic reference point, which is calculated by projecting the reference point coordinates onto the multispectral image plane. is the wheel speed odometer residual at the nth moment, calculated by the wheel speed meter speed integral and posture change, W IMU 、W L 、W O Both are adaptive weight matrices, which are dynamically adjusted according to the sensor noise parameters to meet W * =(Σ * ) -1 , N is the sliding window length, M and N are the valid reference points and wheel speed meter observation sets in the preset sliding window, respectively.
[0028] In one embodiment, hierarchical modeling is performed based on the platform pose after error suppression and the original point cloud to generate a digital surface model, and the digital surface model is fused with satellite geographic data to output multi-scale geographic information results, including:
[0029] According to the platform pose after error suppression, the original point cloud is converted from the local coordinate system to the global coordinate system, and the point cloud data in the global coordinate system is output. The platform pose after error suppression includes position, attitude angle and timestamp;
[0030] The original point cloud is divided into multiple voxel grids according to the point cloud density distribution of the point cloud data in the global coordinate system, and the center point coordinates and average reflection intensity value of each voxel grid are recorded to output a low-density point cloud;
[0031] The least squares surface fitting method is used to calculate the local slope of each low-density point cloud. Based on the local slope and the semantic label of each low-density point cloud, each low-density point cloud is classified into regions and the key area index is output. The key area index includes flat areas and key areas.
[0032] Based on the original point cloud corresponding to the key area index, the RANSAC algorithm is used to fit the plane or quadratic surface model, the inlier threshold is set to 3cm, and outliers with residuals greater than 3 times the standard deviation are removed, and the surface model is output;
[0033] Based on low-density point clouds and surface models, a coarse-grained digital elevation model is generated for flat areas using voxel center points, and a high-density digital elevation model is generated for key areas using surface model interpolation. The coarse-grained digital elevation model and the high-density digital elevation model are then fused to output a digital surface model.
[0034] The satellite digital elevation model in the satellite geographic data is aligned to the resolution of the digital surface model through bilinear interpolation. The acquisition periods of the satellite digital elevation model and the digital surface model are matched based on the timestamp. Based on the difference in terrain entropy between the satellite digital elevation model and the digital surface model, weighted fusion processing is performed to output multi-scale geographic information results.
[0035] In one embodiment, the preset semantic segmentation network is constructed by the following steps:
[0036] Data augmentation is performed based on the historical RGB band images, near-infrared band images, and manually annotated semantic labels of the mobile platform to generate a training dataset.
[0037] Based on the ResNet-101 network architecture and training dataset, the Focal loss function and cosine annealing learning rate scheduling strategy are used, and supervised training is performed through the Adam optimizer to output the teacher model;
[0038] The intermediate feature map of the teacher model is used as the supervision signal, and the KL divergence is calculated with the corresponding layer output of the MobileNetV3 student network. The student model is optimized using a bidirectional distillation strategy to obtain a compressed student model.
[0039] Based on the output feature map of the teacher model and the MobileNetV3 student network, KL divergence loss is calculated to obtain the compressed student model;
[0040] The student model is converted into ONNX format, and 8-bit integer quantization processing is performed on the NPU using the quantization-aware training method. The activation function threshold is optimized based on the target geographic information surveying and mapping scenario to obtain the preset semantic segmentation network.
[0041] In a second aspect, the present application also provides a geographic information surveying and mapping device based on laser radar, comprising:
[0042] The multi-source spatiotemporal collaborative acquisition module is used to perform spatiotemporal calibration processing on the LiDAR, multispectral camera, and data measurement unit deployed on the mobile platform based on their hardware parameters, thereby obtaining spatiotemporally synchronized raw point clouds, multispectral images, IMU measurement data, and wheel speed meter speed data. The IMU measurement data includes the angular velocity and acceleration of the mobile platform, and the data measurement unit includes an IMU subunit and a wheel speed meter subunit.
[0043] The dynamic semantic reference point extraction module is used to extract dynamic semantic reference points based on multispectral images and original point clouds through semantic segmentation and three-dimensional association processing, and to build a dynamic semantic reference point spatiotemporal constraint database;
[0044] The multi-source data pose calculation module is used to calculate the pose based on dynamic semantic reference points, IMU measurement data, and wheel speed meter speed data through an adaptive filtering algorithm, and output the platform pose after error suppression;
[0045] The multi-scale geographic information generation module is used to perform hierarchical modeling based on the platform posture and original point cloud after error suppression, generate a digital surface model, fuse the digital surface model with satellite geographic data, and output multi-scale geographic information results. The multi-scale geographic information results include a fused digital surface model and a semantic label distribution map enhanced based on terrain features.
[0046] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the first aspect when executing the computer program.
[0047] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the first aspect when executed by a processor.
[0048] The aforementioned LiDAR-based geographic information surveying and mapping method, device, equipment, and medium, through spatiotemporal calibration of the LiDAR, multispectral camera, and data measurement unit, obtains spatiotemporally synchronized multi-source data. This effectively eliminates time delays and spatial deviations between sensors, ensures the consistency of raw point clouds, multispectral images, and other data in both temporal and spatial dimensions, and lays a solid foundation for subsequent high-precision surveying and mapping analysis. Secondly, by extracting dynamic semantic reference points from multispectral images and raw point clouds and constructing a spatiotemporal constraint database, key geographic information in dynamic scenes can be identified and located, effectively distinguishing fixed objects from mobile targets, and enhancing the ability to extract geographic information in complex dynamic environments. Furthermore, by combining dynamic semantic reference points, IMU measurement data, and wheel speedometer speed data, and performing pose calculations through an adaptive filtering algorithm, the pose of the mobile platform can be accurately and in real time, suppressing error accumulation and significantly improving the accuracy and robustness of pose calculations. Finally, hierarchical modeling is performed based on the platform posture after error suppression, and the generated digital surface model is fused with satellite geographic data to output multi-scale geographic information results. This not only improves the detail expression ability and modeling accuracy of the digital surface model, but also realizes the efficient fusion of geographic information at different scales.
[0049] Compared with traditional geographic information surveying and mapping methods, this method significantly improves the accuracy and efficiency of geographic information surveying and mapping through technical means such as multi-source sensor spatiotemporal collaborative acquisition, dynamic semantic benchmark extraction, multi-source data posture correction, and multi-scale geographic information fusion, enhances adaptability to dynamic and complex environments, and provides more reliable and efficient geographic information data support for smart city construction, terrain monitoring, environmental assessment and other fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0051] Figure 1 A flowchart of a laser radar-based geographic information surveying and mapping method according to an exemplary embodiment of the present invention is provided;
[0052] Figure 2 A schematic structural diagram of a laser radar-based geographic information surveying and mapping device provided as an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0054] In one embodiment, Figure 1 As shown, a method for geographic information mapping based on laser radar is provided. This embodiment uses the method applied to a terminal as an example for illustration. It is understandable that the method can also be applied to a server, or to a system including a terminal and a server, and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0055] S101: Based on the hardware parameters of the lidar, multispectral camera, and data measurement unit, the lidar, multispectral camera, and data measurement unit deployed on the mobile platform are subjected to spatiotemporal calibration processing to obtain spatiotemporally synchronized original point clouds, multispectral images, IMU measurement data, and wheel speed meter speed data. The IMU measurement data includes the angular velocity and acceleration of the mobile platform. The data measurement unit includes an IMU subunit and a wheel speed meter subunit.
[0056] Specifically, the LiDAR is used to acquire 3D point cloud data of the target area, the multispectral camera is used to obtain multispectral image data, and the data measurement unit includes an inertial measurement subunit and a wheel speed meter subunit, which are used to measure the angular velocity, acceleration, and speed of the mobile platform, respectively. By performing spatiotemporal calibration on the LiDAR, multispectral camera, and data measurement unit, the spatiotemporal deviations between the different sensors can be eliminated, resulting in spatiotemporally synchronized raw point clouds, multispectral images, IMU measurement data, and wheel speed meter speed data.
[0057] S102: Extract dynamic semantic reference points based on the multispectral image and the original point cloud through semantic segmentation and three-dimensional association processing, and construct a dynamic semantic reference point spatiotemporal constraint database.
[0058] Specifically, semantic segmentation is an image processing technology that can classify and mark different landforms in an image, such as roads, vegetation, and buildings. And by combining three-dimensional point cloud data, the semantic information in the two-dimensional image can be further mapped into three-dimensional space to achieve semantic annotation in three-dimensional space. Through this process, reference points with specific semantics can be identified, such as road edges, building corners, etc. The reference points can remain relatively stable in dynamic environments such as changes in traffic flow and vegetation growth. In addition, by constructing a dynamic semantic reference point spatiotemporal constraint database, the position information of the reference points under different time and space conditions can be recorded, providing a reliable reference point for subsequent pose solution, thereby improving the accuracy and stability of pose solution.
[0059] S103: Based on the dynamic semantic reference points, IMU measurement data, and wheel speed meter speed data, an adaptive filtering algorithm is used to perform posture calculation and output the platform posture after error suppression.
[0060] Specifically, the adaptive filtering algorithm is a data processing technique that dynamically adjusts filtering parameters based on the characteristics of the input data and environmental conditions. In this embodiment, the algorithm combines the position information provided by dynamic semantic reference points, the angular velocity and acceleration data measured by the IMU, and the speed data of the wheel tachometer to calculate the pose (position and attitude) of the mobile platform in real time. Furthermore, adaptive filtering can effectively suppress error accumulation in dynamic environments or complex terrain, significantly improving the accuracy and stability of pose calculations.
[0061] S104: Based on the platform posture after error suppression and the original point cloud, hierarchical modeling is performed to generate a digital surface model, and the digital surface model is fused with satellite geographic data to output a multi-scale geographic information result. The multi-scale geographic information result includes a fused digital surface model and a semantic label distribution map enhanced based on terrain features.
[0062] Specifically, a digital surface model is a three-dimensional geographic information model that accurately represents the undulations of the terrain surface and the distribution of landforms. Satellite geographic data typically includes large-scale terrain, landforms, and landform information. By fusing it with the generated digital surface model, it can complement the shortcomings of lidar and multispectral camera data in terms of large-scale coverage and macro-geographic information. Furthermore, the fused multi-scale geographic information results include not only a high-precision digital surface model, but also a semantic label distribution map enhanced based on terrain features. This semantic label distribution map can intuitively display the distribution of different landforms, providing richer and more intuitive data support for the application of geographic information.
[0063] In the above method, by performing spatiotemporal calibration processing based on the hardware parameters of the lidar, multispectral camera, and data measurement unit, not only is the precise alignment of multi-source heterogeneous sensor data in the temporal and spatial dimensions achieved, but the inherent bias between sensors is also eliminated by utilizing the characteristics of the hardware parameters, providing a reliable foundation for the subsequent fusion and analysis of surveying and mapping data. Secondly, semantic segmentation and three-dimensional association processing are performed based on the multispectral image and the original point cloud, dynamic semantic reference points are extracted, and a spatiotemporal constraint database is constructed. This not only fully utilizes the complementary advantages of image semantic information and point cloud three-dimensional structural information, but also improves the recognition and positioning accuracy of geographic information by introducing dynamic semantic reference points. Furthermore, this method also uses dynamic semantic reference points, IMU measurement data, and wheel speed meter speed data to perform pose solution through an adaptive filtering algorithm. This combines the advantages of different sensor data, effectively suppresses the error accumulation in the pose solution process, and significantly improves the real-time and accuracy of the calculation process. Finally, hierarchical modeling is performed based on the platform posture and original point cloud after error suppression, and the digital surface model is fused with satellite geographic data to output multi-scale geographic information results. Not only is the fine depiction of terrain details achieved through hierarchical modeling, but the coverage and scale of geographic information are expanded by using satellite geographic data, further improving the accuracy and completeness of geographic information drawing.
[0064] In one embodiment, a spatiotemporal calibration process is performed on the laser radar, multispectral camera, and data measurement unit deployed on a mobile platform based on their hardware parameters to obtain spatiotemporally synchronized raw point clouds, multispectral images, IMU measurement data, and wheel speed meter speed data, including:
[0065] Deploy lidar, multispectral cameras, and data measurement units based on the preset lidar scanning frequency, multispectral camera band parameters, IMU zero bias stability requirements, and wheel speed meter sampling rate;
[0066] According to the checkerboard calibration plate, the laser radar and multispectral camera are aligned to obtain the extrinsic parameter matrix;
[0067] The Allan variance analysis method is used to calibrate the noise parameters of the data measurement unit and construct the error state Kalman filter model;
[0068] Based on the GNSS pulse-second signal, the laser radar, multispectral camera, and data measurement unit are clock-synchronized to obtain a time-synchronized sensor data stream. The sensor data stream includes the initial raw point cloud, initial multispectral image, initial IMU measurement data, and initial wheel speed meter speed data.
[0069] Based on the time-synchronized sensor data stream, dynamic scene test data is generated by collecting and processing dynamic scene motion trajectories, and reprojection error verification processing is performed on the external parameter matrix and error state Kalman filter model to output spatiotemporally synchronized original point clouds, multispectral images, IMU measurement data, and wheel speed meter speed data.
[0070] Specifically, the preset lidar scanning frequency determines the amount and density of point cloud data acquired per unit time. For example, a high scanning frequency can more accurately depict the outline of the target object. The band parameters of the multispectral camera include the spectral range it can capture. Different bands can be used to identify different ground object information. For example, the near-infrared band helps distinguish vegetation from other ground objects. The IMU zero-bias stability requirement specifies the accuracy of the IMU measurement. Among them, low zero-bias stability can reduce measurement errors. The wheel speed meter sampling rate specifies the frequency at which the wheel speed meter obtains speed data. The above parameters can ensure that the installation position and posture of each sensor meet the data collection requirements to achieve comprehensive and accurate perception of geographic information. After the sensor deployment is completed, the coordinate system alignment of the lidar and multispectral camera is required.
[0071] Furthermore, a checkerboard calibration plate can be placed within the sensor's field of view to ensure that both the LiDAR and the multispectral camera can clearly capture the feature points of the calibration plate. The checkerboard's corner points can then be extracted as feature points in the multispectral image, and the three-dimensional feature points corresponding to the checkerboard can be extracted from the LiDAR point cloud data. Using algorithms such as Zhang's calibration method, the two-dimensional feature points in the multispectral image and the three-dimensional feature points in the LiDAR point cloud are matched, and the rotation matrix and translation vector between the LiDAR coordinate system and the multispectral camera coordinate system are calculated to obtain the extrinsic parameter matrix. This extrinsic parameter matrix can realize the conversion between the two coordinate systems, so that the point cloud data collected by the LiDAR and the image data collected by the multispectral camera can accurately correspond in space, providing a basis for subsequent data fusion.
[0072] Furthermore, Allan variance is a noise analysis method that can accurately estimate the random noise characteristics of sensors, such as white noise and random walk noise. Therefore, Allan variance analysis can be used to evaluate the noise characteristics of the angular velocity and acceleration data collected by the data measurement unit through long-term statistical analysis. Based on the analysis results, noise parameters such as the random walk parameters of the IMU can be determined, and these parameters can be used to construct an error-state Kalman filter model. This model can estimate and correct errors in IMU measurement data, improving the accuracy and reliability of IMU measurement data and providing more accurate inertial measurement information for mobile platform pose calculation. Furthermore, to ensure temporal consistency of data collected by each sensor, the LiDAR, multispectral camera, and data measurement unit can be connected to the pulse-per-second signal provided by the global navigation satellite system through a hardware interface to synchronize the clocks of each sensor. Accurate timestamps are then applied to the data collected by each sensor, thereby generating a time-synchronized sensor data stream containing the initial raw point cloud, initial multispectral image, initial IMU measurement data, and initial wheel speed data.
[0073] Schematically, after completing spatiotemporal calibration, the spatiotemporal synchronization of the sensor data can be verified. For example, based on the time-synchronized sensor data stream, the motion trajectory data of the mobile platform in a dynamic scene is collected to generate dynamic scene test data. Using this test data, the lidar point cloud is projected onto the multispectral image, and the reprojection error between the projected points and the actual image feature points is calculated. Based on the reprojection error, the external parameter matrix and the error state Kalman filter model can be optimized and adjusted, and the parameters can be iteratively corrected. Ultimately, the spatiotemporally synchronized raw point cloud, multispectral camera image, IMU measurement data, and wheel speed meter speed data are output, thereby ensuring the spatiotemporal consistency and accuracy of the data.
[0074] In one embodiment, dynamic semantic reference points are extracted from multispectral images and original point clouds through semantic segmentation and three-dimensional association processing, and a dynamic semantic reference point spatiotemporal constraint database is constructed, including:
[0075] Based on the RGB and near-infrared band data of the multispectral camera, the multispectral image is processed by histogram equalization and vegetation index calculation to obtain an enhanced multispectral image;
[0076] The enhanced multispectral image is input into the preset semantic segmentation network for pixel-level label recognition processing to obtain the segmentation result, and the segmentation result is processed by morphological closing operation to output the semantic segmentation mask and target bounding box;
[0077] According to the original point cloud area corresponding to the semantic segmentation mask, the DBSCAN algorithm is used to perform Euclidean clustering processing to obtain a set of point cloud clusters associated with the ground objects;
[0078] The minimum bounding box is calculated based on the point cloud cluster set, and the center point coordinates are extracted. Combined with the global navigation satellite system position data of the mobile platform, the center point coordinates are converted and processed to output dynamic semantic reference points. The semantic label, timestamp and coordinate uncertainty of each dynamic semantic reference point are recorded to construct a dynamic semantic reference point spatiotemporal constraint database.
[0079] Specifically, based on the RGB and near-infrared band data from the multispectral camera, histogram equalization can be performed on each band of the multispectral image. Histogram equalization is an image enhancement technique that adjusts the image's grayscale distribution, enhancing contrast and making details clearer. It can also calculate the vegetation index, effectively distinguishing vegetation from other features, further improving semantic segmentation. The semantic segmentation network is a deep learning model that classifies each pixel in the image and identifies different feature categories, such as roads, vegetation, and buildings. The enhanced multispectral image is then input into a pre-set semantic segmentation network for pixel-level label recognition. Using a convolutional neural network architecture, the network performs pixel-by-pixel classification on the input multispectral image, generating a segmentation result. This segmentation result is a label map of the same size as the input image, where the value of each pixel represents its corresponding feature category. To eliminate noise and small holes in the segmentation result, a morphological closing operation can be performed on the segmentation result. Morphological closing is an image processing technique that uses dilation followed by erosion to fill small holes in the segmentation results, connect broken parts, and make feature boundaries more complete and continuous. After morphological closing, a semantic segmentation mask and a target bounding box are output. The semantic segmentation mask is a binary image that identifies the segmented feature area; the target bounding box is a rectangular bounding box that embodies the feature area and facilitates subsequent point cloud association processing.
[0080] Furthermore, based on the semantic segmentation mask, point cloud regions corresponding to the segmented features can be extracted from the original point cloud, ensuring that the point cloud data is consistent with the semantic information of the multispectral image. The DBSCAN algorithm is a density-based clustering algorithm that automatically identifies clusters in point clouds and distinguishes noise points. By setting an appropriate neighborhood radius and minimum point count threshold, the DBSCAN algorithm can be used to perform Euclidean clustering on the extracted point cloud regions, resulting in a set of point cloud clusters associated with the features. Each point cloud cluster represents an independent feature instance, such as a building or an area of vegetation. For each point cloud cluster, its minimum bounding box (i.e., the smallest axis-aligned rectangular box) is then calculated, providing the feature's boundary information. The center point coordinates are extracted from the minimum bounding box and used as the initial position information for the dynamic semantic fiducial. By combining the mobile platform's global navigation satellite system position data, the center point coordinates can be converted from the sensor coordinate system to the global coordinate system, thereby ensuring that the coordinate information of the dynamic semantic fiducial is consistent with the actual geographic location. Furthermore, the semantic label indicates the feature category corresponding to the benchmark; the timestamp records the acquisition time of the benchmark; and the coordinate uncertainty reflects the accuracy of the benchmark's location. By recording this information, a dynamic semantic benchmark spatiotemporal constraint database can be constructed.
[0081] In one embodiment, an adaptive filtering algorithm is used to perform pose calculation based on dynamic semantic reference points, IMU measurement data, and wheel speed meter speed data, and output the platform pose after error suppression, including:
[0082] According to the availability of the global navigation satellite system signal, the global coordinate system or the local coordinate system is selected to initialize the state vector of the error state Kalman filter. The state vector includes the platform position, velocity, attitude angle and IMU zero bias parameters, and the covariance matrix of the initial pose and IMU zero bias parameters is output;
[0083] Based on the angular velocity of the mobile platform in the IMU measurement data, the attitude quaternion is updated through the preset kinematic model, and the platform velocity and position increment are calculated according to the mobile platform acceleration, and the predicted posture and covariance matrix are output;
[0084] Based on the wheel speed meter speed, the coordinates of the dynamic semantic reference point are used as the position observation value, and the wheel speed meter speed is used as the speed observation value. The observation equation is constructed and the Kalman gain is calculated to correct the predicted posture and obtain the corrected platform posture and corrected covariance matrix.
[0085] Based on the corrected platform pose and the corrected covariance matrix, a multi-source constrained optimization objective function is constructed within a preset sliding window using the benchmark observation data, IMU zero bias records, and historical wheel speedometer speed data. This function is based on IMU pre-integration constraints, benchmark reprojection error constraints, and wheel speedometer odometry constraints. The Levenberg-Marquardt algorithm is used to solve the function and output the platform pose with error suppression.
[0086] Specifically, when the GNSS signal is strong, the global coordinate system provides a highly accurate position reference. The initial pose is then determined by initializing the platform's position, velocity, attitude angle, and IMU bias parameters. Simultaneously, the covariance matrix of the IMU bias parameters is calculated based on the IMU's noise characteristics and measurement data. This matrix reflects the uncertainty of the IMU bias parameters and provides an initial error estimate for subsequent pose calculations.
[0087] Furthermore, the platform's angular velocity can be used to update the attitude quaternion using a pre-defined kinematic model, such as a posture update model based on quaternion differential equations, to determine the platform's attitude change. The platform's velocity and position increments are then calculated using an integral operation based on the platform's acceleration. This integral operation takes into account the temporal variation of acceleration, enabling accurate calculation of the platform's velocity and position changes over time. Finally, a predicted pose and covariance matrix are output. The predicted pose reflects the platform's estimated position and attitude at the current moment, while the covariance matrix represents the uncertainty of the predicted pose. Furthermore, the wheel speed provides the platform's velocity observations, while the coordinates of the dynamic semantic reference points serve as position observations. Based on these observations, an observation equation can be constructed. This equation describes the relationship between the sensor observations and the predicted pose. The Kalman gain can then be calculated based on the predicted covariance matrix and the observation noise covariance matrix. This Kalman gain determines the weighting of the observation data when revising the predicted state, balancing the uncertainty between prediction and observation. The predicted pose is then corrected using this Kalman gain, resulting in a corrected platform pose and a corrected covariance matrix. The corrected platform pose is closer to the true pose of the platform, and the corrected covariance matrix also reflects the uncertainty of the pose more accurately.
[0088] Specifically, based on the corrected platform pose and covariance matrix, a multi-source constrained optimization objective function can be constructed using historical fiducial observation data, IMU bias records, and historical wheel speed data. The IMU pre-integration constraint calculates the IMU pre-integration residual from the IMU measurements and state variables between adjacent frames to reflect the IMU measurement error. The fiducial reprojection error constraint calculates the reprojection residual by projecting the fiducial coordinates onto the multispectral image plane, reflecting the fiducial observation error. The wheel speedometer odometry constraint measures the wheel speedometer measurement error by calculating the wheel speed odometry residual from the wheel speed integral and pose change. The adaptive weight matrix can be dynamically adjusted based on sensor noise parameters, allowing the objective function to appropriately allocate weights based on the measurement accuracy of different sensors. Furthermore, by setting the sliding window length, the state variables within the sliding window can be optimized and solved, using the Levenberg-Marquardt algorithm for iterative calculations, ultimately outputting the platform pose after error suppression. The Levenberg-Marquardt algorithm is a nonlinear optimization algorithm that combines the advantages of gradient descent and Gauss-Newton methods to quickly converge to the optimal solution. Through multi-objective optimization, error accumulation can be further suppressed, improving the accuracy and stability of pose calculations.
[0089] In one embodiment, the multi-source constrained optimization objective function is:
[0090]
[0091] Among them, x is the set of state variables to be optimized in the preset sliding window, including the platform posture [p k ,q k ]、IMU bias b k and sensor delay τ, is the IMU pre-integration residual of the kth frame, which is calculated from the IMU measurement values and state variables between adjacent frames. is the reprojection residual of the mth dynamic semantic reference point, which is calculated by projecting the reference point coordinates onto the multispectral image plane. is the wheel speed odometer residual at the nth moment, calculated by the wheel speed meter speed integral and posture change, W IMU 、W L 、W O Both are adaptive weight matrices, which are dynamically adjusted according to the sensor noise parameters to meet W * =(Σ * ) -1 , N is the sliding window length, M and N are the valid reference points and wheel speed meter observation sets in the preset sliding window, respectively.
[0092] In one embodiment, a hierarchical modeling process is performed based on the platform pose after error suppression and the original point cloud to generate a digital surface model. The digital surface model is then fused with satellite geographic data to output multi-scale geographic information results, including:
[0093] Based on the platform pose after error suppression, the original point cloud is converted from the local coordinate system to the global coordinate system, and the point cloud data in the global coordinate system is output. The platform pose after error suppression includes position, attitude angle, and timestamp. The original point cloud is divided into multiple voxel grids based on the point cloud density distribution in the global coordinate system, and the center point coordinates and average reflection intensity value of each voxel are recorded to output a low-density point cloud.
[0094] The least squares surface fitting method is used to calculate the local slope of each low-density point cloud. Based on the local slope and the semantic label of each low-density point cloud, each low-density point cloud is classified into regions and the key area index is output. The key area index includes flat areas and key areas.
[0095] Based on the original point cloud corresponding to the key area index, the RANSAC algorithm is used to fit a plane or quadratic surface model. The inlier threshold is set to 3 cm, and outliers with residuals greater than 3 times the standard deviation are removed to output the surface model. Based on the low-density point cloud and surface model, a coarse-grained digital elevation model is generated using the voxel center point for flat areas, and a high-density digital elevation model is generated for key areas using surface model interpolation. The coarse-grained digital elevation model and the high-density digital elevation model are fused to output the digital surface model.
[0096] The satellite digital elevation model in the satellite geographic data is aligned to the resolution of the digital surface model through bilinear interpolation. The acquisition periods of the satellite digital elevation model and the digital surface model are matched based on the timestamp. Based on the difference in terrain entropy between the satellite digital elevation model and the digital surface model, weighted fusion processing is performed to output multi-scale geographic information results.
[0097] Specifically, the platform's position determines the coordinates of the mobile platform in space, the attitude angle describes the platform's rotational state, and the timestamp marks the moment of data acquisition. Using this information, the coordinate transformation matrix can be used to transform the original point cloud from the mobile platform's local coordinate system to the global coordinate system through rotation and translation operations. This eliminates spatial deviations caused by platform motion and installation posture, generating point cloud data in the global coordinate system and providing a unified benchmark for subsequent modeling and analysis. Furthermore, in the density distribution of point cloud data, higher-density areas typically correspond to detailed features such as building surfaces, while lower-density areas may represent open terrain such as squares and open spaces. For example, the original point cloud space can be divided into multiple regular voxel grids based on point cloud density characteristics. For each voxel grid, the coordinates of its center point are calculated as the representative position of the point cloud within the grid, and the average reflectance intensity value of all points within the grid is calculated. This value reflects the surface material characteristics of the feature. This voxelization operation can remove redundant data while preserving the main features of the terrain, generating a low-density point cloud and reducing the computational complexity of subsequent processing.
[0098] Furthermore, the local slope in the point cloud reflects the degree of terrain undulation. For example, a slope close to 0 indicates a flat area, while a larger slope corresponds to a sloping or steep terrain. By combining the local slope with the semantic labels of the corresponding low-density point cloud, each point cloud can be classified into regions. For example, areas with a slope less than a set threshold (e.g., 5°) and semantic labels such as roads or other features can be classified as flat, generating a key region index. This index identifies the location and type of different characteristic regions in the terrain, providing a basis for differentiated modeling. The RANSAC algorithm can be used to fit a plane or quadratic surface model to the original point cloud corresponding to the key region index. The inlier threshold is set to 3 cm, meaning that any point within 3 cm of the fitted surface is considered an inlier. Subsequently, through multiple random sampling and model fitting, the model with the largest number of inliers is selected as the final fitting result. The residuals between each point cloud and the fitted surface are calculated, and outliers with residuals greater than three times the standard deviation are removed to further improve the accuracy and reliability of the surface model.
[0099] Specifically, for key areas, interpolation operations can be performed using a surface model, such as using inverse distance weighted interpolation (IDWI). This generates dense elevation sampling points at regular intervals on the surface model to construct a high-density digital elevation model (DEM). A coarse-grained DEM quickly depicts the general shape of flat terrain with lower precision. Combining the two creates a DSM. Bilinear interpolation, an image scaling method, effectively reduces errors during interpolation. Bilinear interpolation can be used to adjust the resolution of the satellite DEM to match that of the DEM. Furthermore, based on data acquisition timestamps, the acquisition periods of the two can be matched, and data with similar time periods can be selected for fusion to reduce errors caused by temporal differences in terrain changes. Furthermore, terrain entropy is a metric used to measure terrain complexity; higher entropy values indicate more complex terrain. By calculating the differences in terrain entropy between satellite digital elevation models and high-density digital elevation models in different regions, the fusion weights can be dynamically adjusted to fuse the two, generating multi-scale geographic information results that include a fused digital surface model and a semantic label distribution map enhanced based on terrain features, providing rich and accurate data support for applications such as geographic information analysis, urban planning, and disaster monitoring.
[0100] In one embodiment, the preset semantic segmentation network is constructed by the following steps:
[0101] Data augmentation is performed based on the historical RGB band images, near-infrared band images, and manually annotated semantic labels of the mobile platform to generate a training dataset.
[0102] Based on the ResNet-101 network architecture and training dataset, the Focal loss function and cosine annealing learning rate scheduling strategy are used, and supervised training is performed through the Adam optimizer to output the teacher model;
[0103] The intermediate feature map of the teacher model is used as the supervision signal, and the KL divergence is calculated with the corresponding layer output of the MobileNetV3 student network. The student model is optimized using a bidirectional distillation strategy to obtain a compressed student model.
[0104] Based on the output feature map of the teacher model and the MobileNetV3 student network, the KL divergence loss is calculated to obtain the compressed student model; the student model is converted to ONNX format, and 8-bit integer quantization is performed on the NPU using the quantization-aware training method. The activation function threshold is optimized based on the target geographic information surveying and mapping scenario to obtain the preset semantic segmentation network.
[0105] Specifically, for historical RGB and near-infrared images, random cropping, rotation, flipping, and color adjustments can be performed to generate more diverse training samples, increasing the model's adaptability to different scenes and lighting conditions. The enhanced images and their corresponding semantic labels are then combined to form a training dataset. ResNet-101, a deep residual network, boasts powerful feature extraction capabilities, making it suitable for complex semantic segmentation tasks. Using the ResNet-101 network architecture as the foundational framework for the teacher model, this architecture, through a multi-layer residual block design, effectively extracts deep semantic features from images. During training, the Focal loss function can be used instead of the traditional cross-entropy loss function. This introduces a modulation factor to reduce the weight of easily classified samples and focus learning on difficult-to-classify samples. This can address model bias caused by imbalanced object classification in geographic information mapping scenarios. Furthermore, a cosine annealing learning rate scheduling strategy accelerates convergence by using a higher learning rate at the beginning of training. As training progresses, the learning rate gradually decays according to a cosine function curve, preventing the model from becoming trapped in local optima. Finally, the Adam optimizer is used to iteratively update the network parameters, which enables the teacher model to accurately identify and segment various types of objects in the image and output high-precision semantic segmentation results.
[0106] Specifically, intermediate-layer feature maps from the trained teacher model, such as those output by the convolutional layer, are used as supervisory signals. KL divergence is calculated with the output of the corresponding layer in the MobileNetV3 student network. By minimizing the KL divergence between the two, the student network is guided to learn the teacher network's feature representation capabilities. A bidirectional distillation strategy is also employed to dynamically adjust weights based on the importance of features at different layers, prioritizing features that are more critical to the semantic segmentation task. This ultimately yields a compressed student model. This model significantly reduces the number of parameters and significantly improves inference speed, meeting the computational resource constraints and real-time requirements of mobile platforms. Furthermore, the compressed student model is converted to the ONNX (Open Neural Network Exchange) format, ensuring cross-platform and cross-framework compatibility and facilitating deployment across various hardware devices and deep learning frameworks. Furthermore, a quantization-aware training approach is employed to perform 8-bit integer quantization on the NPU (neural network processing unit). This converts model parameters and calculations from traditional 32-bit floating-point numbers to 8-bit integers, significantly reducing model memory usage and computational complexity while maintaining manageable precision loss. Finally, combined with the characteristics of the target geographic information mapping scene, such as terrain complexity and object category distribution, the threshold parameters of activation functions such as ReLU can be adjusted experimentally to enhance the model's responsiveness to key object features, and ultimately obtain a preset semantic segmentation network suitable for real-time geographic information semantic segmentation tasks.
[0107] like Figure 2As shown, based on the same inventive concept, the embodiment of the present application further provides a laser radar-based geographic information surveying and mapping device 200 for implementing the laser radar-based geographic information surveying and mapping method involved above. The implementation solution provided by this device is similar to the implementation solution described in the above method. Therefore, the specific limitations of one or more laser radar-based geographic information surveying and mapping device embodiments provided below can be found in the above limitations of a laser radar-based geographic information surveying and mapping method, and will not be repeated here.
[0108] The multi-source spatiotemporal collaborative acquisition module 201 is used to perform spatiotemporal calibration processing on the laser radar, multispectral camera, and data measurement unit deployed on the mobile platform based on their hardware parameters, thereby obtaining spatiotemporally synchronized raw point clouds, multispectral images, IMU measurement data, and wheel speed meter speed data. The IMU measurement data includes the angular velocity and acceleration of the mobile platform, and the data measurement unit includes an IMU subunit and a wheel speed meter subunit.
[0109] A dynamic semantic reference point extraction module 202 is used to extract dynamic semantic reference points based on the multispectral image and the original point cloud through semantic segmentation and three-dimensional association processing, and to construct a dynamic semantic reference point spatiotemporal constraint database;
[0110] The multi-source data pose calculation module 203 is used to perform pose calculation based on the dynamic semantic reference points, IMU measurement data and wheel speed meter speed data through an adaptive filtering algorithm, and output the platform pose after error suppression;
[0111] The multi-scale geographic information generation module 204 is used to perform hierarchical modeling based on the platform posture after error suppression and the original point cloud to generate a digital surface model, and to fuse the digital surface model with the satellite geographic data to output a multi-scale geographic information result. The multi-scale geographic information result includes a fused digital surface model and a semantic label distribution map enhanced based on terrain features.
[0112] In the above-mentioned device, the multi-source spatiotemporal collaborative acquisition module 201 can effectively solve the temporal and spatial deviation problem of multi-sensor data by accurately calibrating the hardware parameters of the lidar, multispectral camera, and data measurement unit, providing high-quality, high-precision synchronized data for subsequent data processing. The dynamic semantic reference point extraction module 202 uses multispectral images and raw point clouds to extract semantic reference points in dynamic environments through semantic segmentation and three-dimensional association processing. It can deeply mine the semantic information in multi-source data and organically combine two-dimensional image information with three-dimensional point cloud data, thereby achieving accurate identification and positioning of objects in complex dynamic environments, enhancing adaptability to dynamic environments, and providing stable and reliable reference points for subsequent geographic information generation. The multi-source data pose solution module 203, based on dynamic semantic reference points, IMU measurement data, and wheel speed data, can dynamically optimize the pose solution process according to different environmental conditions and data characteristics through an adaptive filtering algorithm, effectively suppressing error accumulation and improving the accuracy and stability of pose solution, thereby making the positioning of the mobile platform more accurate in complex terrain and dynamic environments. The multi-scale geographic information generation module 204 performs hierarchical modeling based on the platform posture and original point cloud after error suppression, and can generate digital surface models with different resolutions and levels of detail, and fuse the digital surface models with satellite geographic data, further improving the integrity and accuracy of geographic information, and outputting richer and more comprehensive multi-scale geographic information results, providing data support for various geographic information applications.
[0113] In an exemplary embodiment, the present invention further provides a computer device comprising a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the lidar-based geographic information mapping method of the present application. A multi-core processor is preferred to improve the system's parallel processing capabilities. The memory provides sufficient temporary storage space to support program execution and data processing. The memory capacity should be large enough to accommodate large amounts of supply information and computing tasks.
[0114] In an exemplary embodiment, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the laser radar-based geographic information mapping method of the present application.
[0115] The above-described embodiments merely represent several implementation methods of the embodiments of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the concept of the embodiments of the present application, and these modifications and improvements fall within the scope of protection of the embodiments of the present application.
Claims
1. A geographic information surveying and mapping method based on laser radar, characterized in that: The method comprises: Performing spatiotemporal calibration processing on the laser radar, the multispectral camera, and the data measurement unit deployed on the mobile platform based on their hardware parameters to obtain spatiotemporally synchronized original point clouds, multispectral images, IMU measurement data, and wheel speed meter speed data, wherein the IMU measurement data includes the angular velocity and acceleration of the mobile platform, and the data measurement unit includes an IMU subunit and a wheel speed meter subunit; Extracting dynamic semantic reference points based on the multispectral image and the original point cloud through semantic segmentation and three-dimensional association processing, and constructing a dynamic semantic reference point spatiotemporal constraint database; Performing posture calculation based on the dynamic semantic reference points, the IMU measurement data, and the wheel speed meter speed data through an adaptive filtering algorithm, and outputting the platform posture after error suppression; A hierarchical modeling process is performed based on the platform pose after error suppression and the original point cloud to generate a digital surface model, and the digital surface model is fused with satellite geographic data to output a multi-scale geographic information result, wherein the multi-scale geographic information result includes a fused digital surface model and a semantic label distribution map enhanced based on terrain features.
2. The method according to claim 1, characterized in that The method includes performing spatiotemporal calibration processing on the laser radar, the multispectral camera, and the data measurement unit deployed on the mobile platform according to the hardware parameters of the laser radar, the multispectral camera, and the data measurement unit to obtain spatiotemporally synchronized original point clouds, multispectral images, IMU measurement data, and wheel speed meter speed data, including: Deploy the lidar, the multispectral camera, and the data measurement unit according to the preset lidar scanning frequency, the multispectral camera band parameters, the IMU zero bias stability requirements, and the wheel speed meter sampling rate; Performing coordinate system alignment processing on the laser radar and the multispectral camera according to a checkerboard calibration plate to obtain an extrinsic parameter matrix; The Allan variance analysis method is used to calibrate the noise parameters of the data measurement unit and construct an error state Kalman filter model; performing clock synchronization processing on the laser radar, the multispectral camera, and the data measurement unit according to a global navigation satellite system pulse-second signal to obtain a time-synchronized sensor data stream, the sensor data stream including an initial raw point cloud, an initial multispectral image, an initial IMU measurement data, and an initial wheel speed meter speed data; According to the time-synchronized sensor data stream, dynamic scene test data is generated by dynamic scene motion trajectory acquisition and processing, and the extrinsic parameter matrix and the error state Kalman filter model are subjected to reprojection error verification processing, and the spatiotemporally synchronized original point cloud, the multispectral image, the IMU measurement data, and the wheel speed meter speed data are output.
3. The method according to claim 1, characterized in that The method extracts dynamic semantic reference points based on the multispectral image and the original point cloud through semantic segmentation and three-dimensional association processing, and constructs a dynamic semantic reference point spatiotemporal constraint database, including: According to the RGB and near-infrared band data of the multispectral camera, the multispectral image is subjected to histogram equalization and vegetation index calculation processing to obtain an enhanced multispectral image; Inputting the enhanced multispectral image into a preset semantic segmentation network for pixel-level label recognition processing to obtain a segmentation result, and performing morphological closing operation processing on the segmentation result to output a semantic segmentation mask and a target bounding box; According to the original point cloud area corresponding to the semantic segmentation mask, the DBSCAN algorithm is used to perform Euclidean clustering processing to obtain a point cloud cluster set associated with the ground object; The minimum bounding box is calculated based on the point cloud cluster set, the center point coordinates are extracted, and the center point coordinates are converted into data in combination with the global navigation satellite system position data of the mobile platform to output dynamic semantic reference points. The semantic label, timestamp and coordinate uncertainty of each dynamic semantic reference point are recorded to construct the dynamic semantic reference point spatiotemporal constraint database.
4. The method according to claim 2, characterized in that The method of performing posture calculation processing based on the dynamic semantic reference point, the IMU measurement data, and the wheel speed meter speed data by an adaptive filtering algorithm and outputting the platform posture after error suppression includes: Selecting a global coordinate system or a local coordinate system to perform state vector initialization processing of the error state Kalman filter according to the signal availability status of the global navigation satellite system, wherein the state vector includes the platform position, velocity, attitude angle and IMU zero bias parameters, and outputting the covariance matrix of the initial pose and the IMU zero bias parameters; Based on the angular velocity of the mobile platform in the IMU measurement data, the attitude quaternion is updated through a preset kinematic model, and the platform velocity and position increment are calculated according to the acceleration of the mobile platform, and the predicted posture and covariance matrix are output; According to the wheel speed meter speed, the coordinates of the dynamic semantic reference point are used as position observation values, and the wheel speed meter speed is used as speed observation value, an observation equation is constructed, and a Kalman gain is calculated to correct the predicted posture to obtain a corrected platform posture and a corrected covariance matrix; Based on the corrected platform pose and the corrected covariance matrix, a multi-source constrained optimization objective function based on IMU pre-integration constraints, reference point reprojection error constraints, and wheel speedometer odometry constraints is constructed within a preset sliding window through the benchmark point observation data at historical moments, IMU zero bias records, and historical wheel speedometer speed data. The function is solved using the Levenberg-Marquardt algorithm and the error-suppressed platform pose is output.
5. The method according to claim 4, characterized in that The multi-source constrained optimization objective function is: Wherein, x is the set of state variables to be optimized in the preset sliding window, including the platform posture [p k ,q k ]、IMU bias b k and sensor delay τ, is the IMU pre-integration residual of the kth frame, which is calculated from the IMU measurement values and state variables between adjacent frames. is the reprojection residual of the mth dynamic semantic reference point, which is calculated by projecting the reference point coordinates onto the multispectral image plane. is the wheel speed odometer residual at the nth moment, calculated by the wheel speed meter speed integral and posture change, W IMU 、W L 、W O Both are adaptive weight matrices, which are dynamically adjusted according to the sensor noise parameters to meet W * =(Σ * ) -1 , N is the sliding window length, M and N are the valid reference points and wheel speed meter observation sets in the preset sliding window respectively.
6. The method according to claim 1, characterized in that The method includes performing hierarchical modeling processing based on the platform posture after error suppression and the original point cloud to generate a digital surface model, and fusing the digital surface model with satellite geographic data to output a multi-scale geographic information result, including: According to the platform posture after error suppression, the original point cloud is converted from the local coordinate system to the global coordinate system, and point cloud data in the global coordinate system is output, wherein the platform posture after error suppression includes position, attitude angle and timestamp; Dividing the original point cloud into a plurality of voxel grids according to the point cloud density distribution of the point cloud data in the global coordinate system, recording the center point coordinates and the average reflection intensity value of each voxel grid, and outputting a low-density point cloud; The least squares surface fitting method is used to calculate the local slope of each low-density point cloud, and based on the local slope and the semantic label of each low-density point cloud, the low-density point cloud is regionally classified, and a key area index is output, where the key area index includes a flat area and a key area; According to the original point cloud corresponding to the key area index, the RANSAC algorithm is used to fit the plane or quadratic surface model, the inlier threshold is set to 3 cm, and outliers with residuals greater than 3 times the standard deviation are removed, and the surface model is output; Based on the low-density point cloud and the curved surface model, a coarse-grained digital elevation model is generated for the flat area using voxel grid center points, a high-density digital elevation model is generated for the key area using the curved surface model interpolation, and the coarse-grained digital elevation model and the high-density digital elevation model are fused to output the digital surface model; The satellite digital elevation model in the satellite geographic data is aligned to the resolution of the digital surface model through bilinear interpolation, the acquisition period of the satellite digital elevation model and the digital surface model are matched based on the timestamp, and based on the difference in terrain entropy between the satellite digital elevation model and the digital surface model, a weighted fusion process is performed to output the multi-scale geographic information result.
7. The method according to claim 3, characterized in that The preset semantic segmentation network is constructed by the following steps: Performing data augmentation processing on the historical RGB band images, near-infrared band images, and manually annotated semantic labels of the mobile platform to generate a training data set; Based on the ResNet-101 network architecture and the training dataset, the Focal loss function and the cosine annealing learning rate scheduling strategy are used, and supervised training is performed through the Adam optimizer to output a teacher model; The intermediate feature map of the teacher model is used as a supervisory signal, and the KL divergence is calculated with the corresponding layer output of the MobileNetV3 student network. The student model is optimized using a bidirectional distillation strategy to obtain a compressed student model. According to the output feature map of the teacher model and the MobileNetV3 student network, KL divergence loss calculation is performed to obtain a compressed student model; the student model is converted into ONNX format, 8-bit integer quantization is performed on the NPU using a quantization-aware training method, and the activation function threshold is optimized based on the target geographic information surveying and mapping scenario to obtain the preset semantic segmentation network.
8. A geographic information surveying and mapping device based on laser radar, characterized in that: The device comprises: A multi-source spatiotemporal collaborative acquisition module is configured to perform spatiotemporal calibration processing on the laser radar, multispectral camera, and data measurement unit deployed on the mobile platform based on their hardware parameters, thereby obtaining spatiotemporally synchronized original point clouds, multispectral images, IMU measurement data, and wheel speed meter speed data. The IMU measurement data includes the angular velocity and acceleration of the mobile platform, and the data measurement unit includes an IMU subunit and a wheel speed meter subunit. A dynamic semantic reference point extraction module is used to extract dynamic semantic reference points based on the multispectral image and the original point cloud through semantic segmentation and three-dimensional association processing, and to construct a dynamic semantic reference point spatiotemporal constraint database; A multi-source data pose calculation module is used to perform pose calculation based on the dynamic semantic reference points, the IMU measurement data, and the wheel speed meter speed data through an adaptive filtering algorithm, and output the platform pose after error suppression; The multi-scale geographic information generation module is used to perform hierarchical modeling processing based on the platform posture after error suppression and the original point cloud to generate a digital surface model, and to fuse the digital surface model with satellite geographic data to output a multi-scale geographic information result, wherein the multi-scale geographic information result includes a fused digital surface model and a semantic label distribution map enhanced based on terrain features.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Three-dimensional measurement data fusion device and method based on Beidou positioning and laser scanning
CN121522656A