A method and system for road asset inventory based on three-dimensional Gaussian splashing
Through a three-dimensional Gaussian splashing method, a multi-source data training model is used to conduct road assets inventory, which solves the problems of inefficiency and large errors in the existing technology, and realizes efficient and accurate asset management and display, supporting scientific decision-making and maintenance of the road.
Patent Information
- Application Number
- CN202510429225.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The inventory technology of existing road assets is inefficient, has large errors, complex data processing and poor display results, making it difficult to meet the efficient, accurate inventory and display needs of large-scale road assets.
Using a three-dimensional Gaussian splashing method, the three-dimensional Gaussian splashing model is trained by collecting multi-source data (image data, IMU data, lidar data), pre-processing, and combining semantic segmentation and sparse point cloud matching to achieve automated inventory and display of road assets.
It improves the efficiency and accuracy of road asset inventory, reduces labor costs, realizes real-time updates and accurate asset information management, and supports scientific decision-making and maintenance of roads.
Smart Images

Figure CN119940745B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road engineering, and particularly to a method and system for road asset inventory based on three-dimensional Gaussian splash. Background Art
[0002] In the fields of road engineering and highway engineering, the inventory and display of road assets are crucial for road maintenance, management, and planning. Road assets include infrastructure such as signs, streetlights, gantries, guardrails, and signal lights. Their accurate inventory and visual display can provide important decision-making support for road management departments. Currently, road asset inventory is mainly carried out by staff through on-site visits, manually measuring and recording road assets using measuring tools (such as tape measures, rangefinders, etc.). In addition, there are also individual cases using lidar point cloud scanning technology to scan the road environment through vehicle-mounted or airborne lidar equipment to generate high-precision three-dimensional point cloud data, and then extract road asset information.
[0003] Manual inventory with static measurement: Staff conduct on-site visits and manually measure and record road assets using measuring tools (such as tape measures, rangefinders, etc.). Low efficiency: Manual inventory is time-consuming and laborious, and it is difficult to meet the needs of large-scale road asset inventory. Human error: The measurement results are easily affected by human factors, resulting in inaccurate data. Low safety: Staff need to conduct measurements on the road, posing certain safety risks. Poor display effect: The display effect is not good.
[0004] Lidar point cloud scanning technology: Scan the road environment through vehicle-mounted or airborne lidar equipment to generate high-precision three-dimensional point cloud data, and then extract road asset information. Difficult dynamic update: Lack of lightweight representation methods to support incremental updates. Low efficiency: Traditional point cloud processing requires complex registration and post-processing, and the point cloud data occupies a large amount of storage space, and the display real-time performance is poor. Summary of the Invention
[0005] The purpose of the present invention is to solve the problems of low efficiency, large errors, complex data processing, and poor display effect existing in the existing road asset inventory and display technologies, and provide a method and system for road asset inventory and display based on three-dimensional Gaussian splash, which can efficiently and accurately inventory and display road assets, improve the efficiency and quality of road management, and provide strong support for road maintenance, management, and planning.
[0006] To achieve the above purpose, the present invention provides the following solutions:
[0007] A method for road asset inventory based on three-dimensional Gaussian splash, comprising:
[0008] Collect multi-source data of road assets; among them, the multi-source data includes: image data, IMU data, and lidar data;
[0009] Preprocess the multi-source data;
[0010] Use the pre-trained data for three-dimensional Gaussian splash model training;
[0011] Based on the trained three-dimensional Gaussian splash model, conduct road asset inventory.
[0012] Optionally, preprocessing the multi-source data includes:
[0013] Perform semantic segmentation on the image data;
[0014] Convert IMU data to camera pose;
[0015] Match the lidar data and the image data to obtain sparse point clouds with color and position information.
[0016] Optionally, performing semantic segmentation on the image data includes:
[0017] Use a semantic segmentation neural network model based on the Transformer architecture to perform semantic segmentation on the image data; among them, the semantic segmentation neural network model based on the Transformer architecture is trained using historical image data with manual annotation and data augmentation, and the historical image data includes various road scenes and weather conditions;
[0018] Save the result of model segmentation as a grayscale mask image, and generate a masked grayscale mask image with the same size as the grayscale mask image; among them, the masked grayscale mask image only includes: vehicles, pedestrians, and sky.
[0019] Optionally, converting IMU data to camera pose includes:
[0020] Based on the IMU data, estimate the attitude information by angular velocity integration and estimate the position information by acceleration integration.
[0021] Optionally, matching the lidar data and the image data includes:
[0022] Through external parameter calibration, convert the point cloud in the lidar data to the camera coordinate system;
[0023] Use the camera internal parameter matrix to project the point cloud in the camera coordinate system onto the image plane;
[0024] Convert the projected plane coordinates to normalized image coordinates, and filter out the points projected outside the image;
[0025] Extract image key points from the image using the SIFT feature extraction algorithm; wherein, the image key points are pixel points or regions with significant features in the image;
[0026] For each image key point, search for the nearest point in the projected point cloud;
[0027] Based on these nearest points, form a sparse point cloud.
[0028] Optionally, using the pre-trained data for three-dimensional Gaussian splash model training includes:
[0029] Use the masked grayscale mask image, camera intrinsics, camera pose, and sparse point cloud to train the three-dimensional Gaussian splash model;
[0030] Using the masked grayscale mask image, camera intrinsics, camera pose, and sparse point cloud to train the three-dimensional Gaussian splash model includes:
[0031] (1). Initialize each point in the sparse point cloud as a 3D Gaussian distribution;
[0032] (2). For each camera pose, project the Gaussian distribution onto the image plane;
[0033] (3). Render the image using the projected Gaussian distribution;
[0034] (4). Based on the rendered image, calculate the loss according to the masked grayscale mask image;
[0035] (5). According to the calculated loss, use the gradient descent method to optimize the parameters of the Gaussian distribution;
[0036] (6). Repeat steps (2) to (5) until the loss converges or reaches the maximum number of iterations to obtain the trained model.
[0037] Optionally, calculate the loss according to the masked grayscale mask image;
[0038] Ignore the spatial region according to the masked grayscale mask image and only train the static region. For each region of the image, introduce a weight w alpha , for the weight w of the dynamic region alpha = 0, for the weight w of the static region alpha = 1; wherein, the dynamic regions include: vehicles, pedestrians, sky, and the static regions include: road surface, road assets, buildings, greenery;
[0039] For w alpha = 0 region, update C gt :
[0040]
[0041] Loss calculation:
[0042] Calculate the loss between the rendered image and the ground truth image:
[0043] Use the L1 loss function:
[0044] 1
[0045] where C render is the rendered image, and C gt is the ground truth image;
[0046] Use SSIM Loss:
[0047] The SSIM Loss is defined as:
[0048]
[0049] Combine the L1 Loss and the SSIM Loss with weights:
[0050]
[0051] where: L is the final loss, and λ dssim is the weight of the SSIM Loss, and its value ranges from [0, 1].
[0052] Optionally, based on the trained 3D Gaussian splash model, perform road asset inventory including:
[0053] Based on the trained 3D Gaussian splash model, perform road asset extraction, asset status analysis, asset size measurement, and GPS matching;
[0054] Performing road asset extraction includes:
[0055] Based on the trained 3D Gaussian splash model, train an encoder to map the features of Gaussian points to high-dimensional feature vectors;
[0056] For each Gaussian kernel, project the Gaussian kernel onto the corresponding grayscale mask image through the camera pose to obtain the class label of the Gaussian kernel; if the Gaussian kernel projects to different class labels under multiple viewpoints, a voting mechanism is used to determine its final class; assign a 32-bit class feature vector to each Gaussian kernel;
[0057] Use the grayscale mask image as the supervision signal and train the class features through an optimization algorithm;
[0058] Design a neural network with 32-bit categorical features as input and categorical labels as output. Train a discriminator using labeled Gaussian kernel data with cross-entropy loss as the loss function and the optimization objective being to minimize the categorical prediction error. Among them, the Gaussian kernel data includes categorical features and corresponding categorical labels.
[0059] Input the 32-bit categorical features of each Gaussian kernel into the trained discriminator to obtain the categorical label of each Gaussian kernel. Screen out the Gaussian kernels of the specified category according to the categorical labels to generate a segmented three-dimensional scene.
[0060] Perform spatial clustering on the Gaussian kernels of the extracted road assets according to the segmented three-dimensional scene to obtain a three-dimensional Gaussian splash model of independent assets and number each asset. Among the extracted road assets are: signs, gantries, street lights.
[0061] Perform asset status analysis including:
[0062] Project the Gaussian cluster of a single asset onto k camera views, use a classification model to classify the rendered images to obtain k classification results, determine the final asset status analysis result according to the voting mechanism, and record the asset number and asset status. Among them, the classification model uses a fine-tuned convolutional neural network classification model.
[0063] Perform GPS matching including:
[0064] According to the information of the extracted three-dimensional Gaussian splash model of the asset, calculate the average GPS coordinates of the Gaussian kernels of each asset, which are the real-world coordinates of the asset.
[0065] Perform asset size measurement including:
[0066] Fit the minimum bounding cube according to the world coordinate distribution of each asset, and the length, width, and height of the cube are considered as the length, width, and height information of the asset.
[0067] A system for inventorying road assets based on three-dimensional Gaussian splash, the system includes: a data acquisition module, a data preprocessing module, a three-dimensional Gaussian splash model training module, and a three-dimensional Gaussian splash model analysis module.
[0068] The data acquisition module is used to collect multi-source data of road assets.
[0069] The data preprocessing module is used to preprocess the multi-source data.
[0070] The three-dimensional Gaussian splash model training module is used to train the three-dimensional Gaussian splash model using pre-trained data.
[0071] The three-dimensional Gaussian splash model analysis module is used to conduct road asset inventory based on the trained three-dimensional Gaussian splash model.
[0072] Optionally, the data acquisition module includes: a 360 panoramic camera, a professional real-time kinematic surveying device with an inertial measurement unit, and a lidar sensor; each sensor is assembled and installed on the roof of the vehicle through a rigid structure; an industrial industrial computer is installed inside the vehicle to supply power to the equipment and control the acquisition work. The industrial industrial computer is equipped with GPU computing power and is also equipped with a 5G module, which can upload the acquired data to the cloud server.
[0073] The beneficial effects of the present invention are as follows:
[0074] Through the automated data acquisition and processing technology, the present invention greatly improves the efficiency of road asset inventory. Compared with the traditional manual inventory method, it can complete the inventory work of large-scale road assets in a short time. It reduces the dependence on a large number of manual measurements and data processing personnel and lowers the labor cost.
[0075] It can obtain the change information of road assets in real time and update the inventory data in a timely manner. When there are damages, replacements, or additions to road assets, the system can respond quickly, update the asset information, ensure the timeliness and accuracy of the inventory data, and provide timely and accurate basis for road maintenance and management.
[0076] Accurate road asset inventory data helps to reasonably plan and allocate resources, and avoid resource waste caused by inaccurate data.
[0077] It provides comprehensive, accurate, and real-time road asset information for road management departments, helping managers make more scientific and reasonable decisions.
[0078] Through an efficient inventory and display system, it can better manage and maintain road assets, extend the service life of assets, and enhance the overall value of road assets. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0080] Figure 1 It is a schematic flow chart of a method for road asset inventory based on three-dimensional Gaussian splash according to an embodiment of the present invention;
[0081] Figure 2 It is a schematic diagram of the system installation according to an embodiment of the present invention. Detailed Implementation Manner
[0082] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0083] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0084] Three-dimensional Gaussian splatting (3DGS) is an advanced 3D modeling and visualization technology. It constructs a 3D scene by representing an image as tiny elliptical spots, i.e., Gaussian kernels. The Gaussian kernels can be stretched, compressed, and given color and transparency, and finally fused to form a continuous surface. The core of Gaussian splatting lies in using 3D Gaussian functions as the basic elements of the scene and optimizing the parameters of these Gaussian functions to achieve high-quality 3D scene reconstruction and new view synthesis. 3DGS has the advantages of high efficiency, flexibility, high-quality rendering, and real-time performance.
[0085] As Figure 1 shown, this embodiment proposes a method for inventorying road assets based on three-dimensional Gaussian splatting, including:
[0086] Collect multi-source data of road assets; wherein, the multi-source data includes: image data, IMU data, and lidar data;
[0087] Preprocess the multi-source data;
[0088] Use the pre-trained data for training the three-dimensional Gaussian splatting model;
[0089] Based on the trained three-dimensional Gaussian splatting model, conduct an inventory of road assets.
[0090] Specifically, in this embodiment, in the data collection stage:
[0091] Install data collection devices such as a 360 camera, IMU, RTK, and lidar on the roof of the vehicle to ensure the stability of the devices and the accuracy of data collection. The vehicle travels at a constant speed (30 km / h) to collect data on the target road. During the collection process, record data such as original pictures, camera poses, and laser point clouds, and obtain accurate geographical location information through RTK.
[0092] An acquisition program is deployed on the industrial control computer. After the acquisition program is started, it receives the PPS (Pulse Per Second) pulse signal from the RTK-GPS module and records the current system time T as the synchronization reference time. For sensors that support hardware triggering (such as cameras and lidar), the PPS pulse is directly used as the trigger signal, and the sensor immediately starts to acquire a frame of data. For sensors that do not support hardware triggering (such as RTK), after the acquisition program detects the PPS pulse, it immediately sends a software instruction to require the sensor to start acquiring data. Each sensor starts to acquire data after being triggered and stamps each frame of data with a timestamp. Align the data of all sensors according to the timestamp T to ensure that the data frames at the same moment can be matched. Store the acquired data frame by frame into a file or a database.
[0093] Furthermore, the preprocessing of the multi-source data includes:
[0094] Performing semantic segmentation on the image data;
[0095] Converting IMU data into camera poses;
[0096] Matching lidar data and image data to obtain sparse point clouds with color and position information.
[0097] Furthermore, performing semantic segmentation on the image data includes:
[0098] Using a semantic segmentation neural network model based on the Transformer architecture to perform semantic segmentation on the image data; among them, the semantic segmentation neural network model based on the Transformer architecture is trained using manually labeled and data-augmented historical image data, and the historical image data includes various road scenes and weather conditions;
[0099] Saving the result of the model segmentation as a grayscale mask image and generating a masked grayscale mask image with the same size as the grayscale mask image; among them, the masked grayscale mask image only includes: vehicles, pedestrians, and the sky.
[0100] Furthermore, converting IMU data into camera poses includes:
[0101] Based on the IMU data, estimating the attitude information by angular velocity integration and estimating the position information by acceleration integration.
[0102] Furthermore, matching lidar data and image data includes:
[0103] Through external parameter calibration, converting the point cloud in the lidar data into the camera coordinate system;
[0104] Using the camera internal parameter matrix to project the point cloud in the camera coordinate system onto the image plane;
[0105] Convert the projected planar coordinates into normalized image coordinates and filter out the points projected outside the image;
[0106] Extract image key points from the image using the SIFT feature extraction algorithm;
[0107] For each image key point, search for the nearest point in the projected point cloud;
[0108] Based on these nearest points, form a sparse point cloud.
[0109] Among them, the image key points are pixel points or regions with significant features in the image; the local regions around the key points have unique textures, shapes, or structures in the image, which can be distinguished from other regions; the key points are robust to the rotation, scaling, illumination changes, etc. of the image and can be stably detected under different conditions. The same key points can be repeatedly detected in different images.
[0110] Specifically, in this embodiment, after the acquisition program acquires multi-sensor data, the analysis program immediately preprocesses the data. The data preprocessing specifically includes:
[0111] (1) Perform semantic segmentation on the original image
[0112] Use a semantic segmentation neural network model (SegFormer) based on the Transformer architecture to perform semantic segmentation on the acquired original image. The model is fine-tuned using manually labeled and data-augmented data from previous acquisitions. The training data includes various road scenes and weather conditions, etc., and can effectively segment road assets and objects such as pedestrians, vehicles, sky, signs, signal lights, gantries, etc. in various scenarios. The model's mean intersection over union (mIoU) > 95%.
[0113] The result of the model segmentation is saved in the form of a grayscale mask image of size H*W, where H and W are the width and height of the original image, and the value of the mask image represents the segmentation category ID of this pixel coordinate. At the same time, a masked grayscale mask image needs to be generated, with the same size as the grayscale mask image. The masked grayscale mask image only includes categories such as vehicles, pedestrians, and sky. Because vehicles and pedestrians are dynamic objects during data acquisition and have a relative motion relationship with the acquisition vehicle, and the 3DGS reconstruction of road assets is mainly for the static background, so the masked mask image is generated here to avoid the interference of dynamic objects during subsequent training and affect the modeling quality. The sky is not within the scope of 3DGS reconstruction, but occupies a relatively large proportion in the image and will also affect the reconstruction quality, so it is filtered together with dynamic objects.
[0114] (2) Convert IMU data into camera pose information
[0115] IMU data can provide the acceleration a(t) and angular velocity ω(t) of the camera at each time point. The initial position of the camera is obtained through extrinsic calibration during calibration, i.e., the initial attitude q(0) and initial position p(0).
[0116] (a)Estimate the attitude by integrating the angular velocity.
[0117]
[0118] q(t) is the attitude at the current time t, represents the multiplication of quaternions, and Δq(t) is the incremental quaternion obtained by integrating the angular velocity, which can be calculated by the following formula:
[0119]
[0120] where qω(t) is the quaternion calculated from the angular velocity ω(t) and the time interval Δt.
[0121]
[0122] Δt is the time interval, i.e., the data acquisition interval.
[0123] (b)Estimate the position by integrating the acceleration.
[0124] The position is the integral of velocity with respect to time, and the position p(t) can be updated through the velocity v(t).
[0125]
[0126] where p(t) is the camera position at the current time t, and v(t) is calculated by the following formula:
[0127]
[0128] a(t) is the acceleration collected by the IMU;
[0129] Through the above, the position p(t) and orientation q(t) of the camera are obtained from the IMU data (acceleration, angular velocity).
[0130] (3)Construct a sparse point cloud
[0131] Obtain the camera image synchronized with the lidar, and project the point cloud onto the image plane to obtain color information:
[0132] The original coordinates of the point cloud P lidar =(x lidar ,y lidar ,z lidar ).
[0133] Camera coordinate system: Through external parameter calibration, the point cloud is transformed into the camera coordinate system:
[0134]
[0135] Where: is the rotation matrix from the lidar to the camera; is the translation vector from the lidar to the camera; P cam =(x cam , y cam , z cam ) is the point cloud coordinate in the camera coordinate system.
[0136] Use the camera intrinsic matrix K to project the point cloud in the camera coordinate system onto the image plane:
[0137] p img = K P cam
[0138] Where K is the camera intrinsic matrix:
[0139]
[0140] P cam is the point cloud coordinate (homogeneous coordinate) in the camera coordinate system:
[0141]
[0142] p img is the projected image plane coordinate (homogeneous coordinate):
[0143]
[0144] Where (u, v) are the normalized image coordinates:
[0145]
[0146] Convert the projected homogeneous coordinate to the normalized image coordinate:
[0147]
[0148]
[0149] Where (u, v) are the pixel coordinates of the point cloud on the image plane.
[0150] Filter out the points projected outside the image:
[0151] 0 ≤ u < w, 0 ≤ v < h
[0152] Where w and h are the width and height of the image respectively.
[0153] Extract key points from the image using the SIFT feature extraction algorithm:
[0154]
[0155] where (u i , v i ) are the pixel coordinates of the i-th key point.
[0156] After projecting the lidar point cloud onto the image plane, find the points closest to the image key points:
[0157] For each image key point (u i , v i ), search for the closest points in the projected point cloud:
[0158]
[0159] Retain these closest points to form a sparse point cloud.
[0160] Each point in the sparse point cloud contains the following information:
[0161] 3D coordinates: P cam = (x cam , y cam , z cam ).
[0162] Color information: RGB values (r, g, b) extracted from the image.
[0163] Furthermore, using the pre-trained data for three-dimensional Gaussian splash model training includes:
[0164] Using the mask grayscale mask image, camera intrinsics, camera pose, and sparse point cloud to train the three-dimensional Gaussian splash model;
[0165] Using the mask grayscale mask image, camera intrinsics, camera pose, and sparse point cloud to train the three-dimensional Gaussian splash model includes:
[0166] (1). Initialize each point in the sparse point cloud as a 3D Gaussian distribution;
[0167] For each camera pose, project the Gaussian distribution onto the image plane;
[0168] (3). Render the image using the projected Gaussian distribution;
[0169] (4). Based on the rendered image, calculate the loss according to the mask grayscale mask image;
[0170] (5). Optimize the parameters of the Gaussian distribution using the gradient descent method according to the calculated loss;
[0171] (6). Repeat steps (2) to (5) until the loss converges or the maximum number of iterations is reached to obtain the trained model.
[0172] Specifically, in this embodiment, the three-dimensional Gaussian splash (3DGS) training includes the following steps:
[0173] Use the mask grayscale mask map, camera intrinsics, camera pose, and sparse point cloud obtained by data preprocessing to train the 3DGS model.
[0174] (1) Initialize the Gaussian distribution
[0175] Initialize each point in the sparse point cloud as a 3D Gaussian distribution.
[0176] Each Gaussian distribution is defined by the following parameters:
[0177] Position μ=(x,y,z): Point cloud coordinates.
[0178] Covariance matrix Σ: Initialized as a small diagonal matrix.
[0179] Color c=(r,g,b): Color information extracted from the point cloud.
[0180] Opacity α: Initialized as 1.
[0181] (2) Project the Gaussian distribution onto the image plane
[0182] For each camera pose, project the 3D Gaussian distribution onto the image plane:
[0183] a) Convert the position μ of the Gaussian distribution to the camera coordinate system:
[0184] μ cam =R μ+t
[0185] b) Project μ cam onto the image plane:
[0186] μ img =K μ cam
[0187] Convert the covariance matrix Σ to the image plane:
[0188] Σ img =J Σ J T
[0189] Among them, J is the Jacobian matrix of the projection.
[0190] (3)Render the image
[0191] Render the image using the projected Gaussian distribution:
[0192] For each pixel (u, v), calculate its color value:
[0193]
[0194] Where: ci is the color of the i-th Gaussian distribution; αi is the opacity of the i-th Gaussian distribution; is the probability density function of the i-th Gaussian distribution on the image plane.
[0195] (4)Dynamic object masking and loss calculation
[0196] Ignore the spatial region according to the mask grayscale mask map to ensure that only the static region is trained. For each region of the image, introduce a weight w alpha , for the dynamic regions (vehicles, pedestrians, sky), the weight w alpha = 0, and for the static regions, the weight w alpha = 1. For the regions where w alpha = 0, update C gt before calculating the loss:
[0197]
[0198] Calculating the dynamic object mask is equivalent to labeling the dynamic regions of the rendered image, so that the dynamic object part is ignored when calculating the loss;
[0199] Here, the dynamic part in the real image is modified to the dynamic part in the rendered image; w alpha is the weight in the image used to distinguish dynamic and static targets. This weight is obtained after semantic segmentation of the image by a semantic segmentation neural network model based on the Transformer architecture. This Transformer model is specially trained. After processing, for the dynamic regions in the image, Walpha = 0 (dynamic regions include: vehicles, pedestrians, sky); for the static regions in the image, Walpha = 1 (static regions include: road surface, road assets, buildings, greenery).
[0200] Loss calculation:
[0201] Calculate the loss between the rendered image and the real image:
[0202] Use the L1 loss function:
[0203] 1
[0204] Among them, C render is the rendered image, and C gt is the real image.
[0205] Using SSIM Loss:
[0206] The calculation formula of SSIM (Structural Similarity Index Measure):
[0207]
[0208] Among them: , are the means of the rendered image and the real image respectively; , are the variances of the rendered image and the real image respectively; is the covariance of the rendered image and the real image; c1 and c2 are constants used to stabilize the calculation.
[0209] The SSIM Loss is defined as:
[0210]
[0211] Combining the L1 Loss and the SSIM Loss with weights:
[0212]
[0213] Among them: L is the final loss, and λ dssim is the weight of the SSIM Loss, and its value is between [0, 1].
[0214] L is the final loss. The L1 loss is used to measure the difference between the rendered image and the target image at the pixel level. It is not sensitive to outliers and can provide a smoother optimization process, which is suitable for capturing the detailed information of the image. The SSIM loss is used to measure the difference in perceptual quality between the rendered image and the target image. It can better reflect the perceptual characteristics of the human visual system and is suitable for capturing the overall structure and perceptual quality of the image.
[0215] The final loss L is a weighted combination of the L1 loss and the SSIM loss, which is used to optimize the detailed information and perceptual quality of the image simultaneously. By adjusting λdssim, the contributions of the two can be flexibly balanced according to the task requirements.
[0216] (5) Optimize the Gaussian distribution parameters
[0217] Use the gradient descent method to optimize the parameters of the Gaussian distribution (mean μ, covariance Σ, color c, opacity α):
[0218] Update formula:
[0219]
[0220] where θ is the parameter of the Gaussian distribution and η is the learning rate.
[0221] (6) Iterative training
[0222] Repeat steps (2) to (5) until the loss converges or the maximum number of iterations is reached. Store the 3DGS model as a.ply file.
[0223] Furthermore, based on the trained three-dimensional Gaussian splash model, road asset inventory includes:
[0224] Based on the trained three-dimensional Gaussian splash model, road asset extraction, asset status analysis, asset size measurement, and GPS matching are performed;
[0225] (1) Road asset extraction
[0226] Train an encoder E based on the trained 3DGS model to map the features of Gaussian points to high-dimensional feature vectors.
[0227]
[0228] where is the feature vector of the Gaussian point , D is the feature dimension, and here D is taken as 32; the encoder E adopts a multi-layer perceptron (MLP) structure.
[0229] For each Gaussian kernel, project it onto the corresponding grayscale mask map through the camera pose to obtain its class label. If the Gaussian kernel projects to different class labels under multiple viewpoints, a voting mechanism is used to determine its final class. Assign a 32-bit class feature vector to each Gaussian kernel.
[0230] Use the grayscale mask map as the supervision signal and train the class features through an optimization algorithm (cross-entropy loss) so that they can accurately represent the class of the Gaussian kernel.
[0231] Design a simple neural network (such as MLP), with the 32-bit class feature as the input and the class label (0 - 255) as the output. Use the labeled Gaussian kernel data (class features and corresponding class labels) to train the discriminator, with the cross-entropy loss as the loss function and the optimization goal being to minimize the class prediction error.
[0232] Input the 32-bit class features of each Gaussian kernel into the trained discriminator to obtain its class label. Screen out the Gaussian kernels of the specified class according to the class label to generate the segmented three-dimensional scene.
[0233] Perform spatial clustering on the Gaussian kernels of road assets such as extracted signboards, gantries, and streetlights to obtain a 3D GS model of independent assets (such as a single streetlight), and number each asset.
[0234] (2)Asset status analysis
[0235] Using a fine-tuned convolutional neural network (CNN) classification model, the status of an image can be classified as intact, damaged, deformed, dirty, occluded, etc.
[0236] Project the Gaussian clusters of individual assets onto k camera views, use the classification model to classify the rendered images to obtain k classification results, and determine the final asset status analysis result according to the voting mechanism (e.g., guardrail - intact, guardrail - deformed, guardrail - dirty, etc.), and record the asset number and asset status.
[0237] (3)GPS matching
[0238] The center point (x, y, z) of each Gaussian kernel in the 3D GS model is the spatial coordinate of the lidar point cloud. Since the lidar has been calibrated with GPS, the spatial coordinates can be mapped into GPS coordinates through an affine transformation.
[0239]
[0240] where is the rotation matrix from lidar to GPS.
[0241] is the translation vector from lidar to GPS.
[0242] is the world coordinate of the point cloud in the GPS coordinate system.
[0243] According to the extracted asset 3D GS information, calculate the average GPS coordinate of each Gaussian kernel of the asset, which is the real-world coordinate of the asset.
[0244] (4)Asset size measurement
[0245] According to the world coordinate distribution of each asset, fit the minimum bounding cube, and the length, width, and height of the cube are considered as the length, width, and height information of the asset.
[0246] (5)3D GS noise reduction
[0247] To improve the aesthetics and accuracy of the rendering effect, use the KNN algorithm to remove outlier Gaussian kernels to ensure the smoothness and consistency of the segmentation results.
[0248] After completing the above steps, the category, world coordinates, dimension information, and asset status of each asset are obtained. Based on this information, the quantity of each type of asset is counted, and the analysis results and the 3DGS model are uploaded to the cloud server.
[0249] The method of this embodiment further includes: data rendering and display:
[0250] Upload the analysis results and the modeling results to the cloud server, and then the rasterization technology can be used on the display platform to perform real-time rendering and display of the 3DGS data. At the same time, the discrete Gaussian points are removed through the KNN algorithm to improve the aesthetics and accuracy of the rendering effect. On the display platform, the 3DGS rendering effect, road asset statistics information, asset numbers, and real coordinates are displayed, and an asset query function is provided to facilitate users to view and analyze specific assets in detail.
[0251] This embodiment also discloses a system for inventorying road assets based on three-dimensional Gaussian splash, which includes: a data acquisition module, a data preprocessing module, a three-dimensional Gaussian splash model training module, and a three-dimensional Gaussian splash model analysis module;
[0252] The data acquisition module is used to acquire multi-source data of road assets;
[0253] The data preprocessing module is used to preprocess the multi-source data;
[0254] The three-dimensional Gaussian splash model training module is used to train the three-dimensional Gaussian splash model by using the pre-trained data;
[0255] The three-dimensional Gaussian splash model analysis module is used to conduct an inventory of road assets based on the trained three-dimensional Gaussian splash model.
[0256] Furthermore, the data acquisition module of this embodiment adopts a set of comprehensive data acquisition devices, including a 360 panoramic camera, a professional real-time kinematic (RTK) system with an inertial measurement unit (IMU), a lidar, and other sensors. Each sensor is assembled and installed on the roof through a rigid structure. An industrial industrial personal computer is installed in the vehicle to supply power to the device and control the acquisition work. This industrial personal computer is equipped with GPU computing power, which can accelerate data processing. At the same time, it is equipped with a 5G module and can upload the acquired data to the cloud server.
[0257] In this embodiment, sensor calibration and coordinate alignment:
[0258] Since the installation positions of the various sensors are fixed and the relative poses are known, the following can be obtained before installation: (1) the camera internal parameters (focal length f x , f y and the principal point offset c x , c y, distortion coefficient (dist);
[0259] (2) Coordinate transformation from lidar to camera: rotation matrix and translation vector are obtained through calibration measurement.
[0260] After the device is installed, it is also necessary to calibrate the extrinsic parameters of the camera and RTK-GPS:
[0261] Use a 1m * 1m * 1m cube calibration board, and install high-reflectivity markers (such as reflective stickers or reflective balls) at each corner point for easy lidar recognition. Draw a checkerboard pattern on each face of the cube for easy camera recognition. There is a fixed point at the top of the cube for installing the GPS antenna to ensure stable signals.
[0262] Calibration process:
[0263] (1) Place the calibration object in the scene to ensure good GPS signals.
[0264] (2) Use RTK-GPS to measure the coordinates of the GPS antenna at the top of the calibration object.
[0265] (3) Use the lidar to scan the calibration object and extract the point cloud coordinates of the reflective markers. Combine the GPS data of the calibration object and use the least squares method or SVD (singular value decomposition) to calculate the rotation matrix and translation vector , and use an optimization algorithm (such as Levenberg-Marquardt) to further improve the accuracy, that is, obtain the transformation relationship from the lidar coordinate system to the GPS coordinate system.
[0266] (4) Use a 360° camera to take pictures of the calibration object, extract the checkerboard corner point coordinates, and calculate the extrinsic parameters of the camera.
[0267] In the data acquisition stage, install data acquisition devices such as 360 cameras, IMUs, RTKs, and lidars on the roof of the vehicle, as Figure 2 shown, to ensure the stability of the devices and the accuracy of data acquisition. The vehicle travels at a constant speed (30 km / h) to collect data on the target road. During the collection process, record data such as original pictures, camera poses, and laser point clouds, and obtain accurate geographical location information through RTK.
[0268] The system architecture of this embodiment includes a data acquisition module, a data preprocessing module, a 3DGS training module, a 3DGS model analysis module, an asset database and management platform, and a user interaction module. Each module works together through an efficient data transmission and communication mechanism to ensure the stable operation and efficient processing of the entire system.
[0269] Asset Database and Management Platform:
[0270] After the cloud server obtains the Gaussian splash model and attribute metadata, it uses WebGL to perform real-time rendering and display of 3DGS data. At the same time, differential updates can be performed according to the GPS coordinate neighborhood matching of assets (updating the Gaussian parameters of some areas), that is, ensuring that the rendering result of the platform is the latest inventory modeling result. Finally, on the display platform, not only can the 3DGS rendering effect be displayed, but also the statistical information of road assets can be displayed, number each asset, and display its real coordinates. In addition, queries can be made according to the asset number, and the area near the asset can be rendered in real time after the query, facilitating users to view, analyze, and mark maintenance records for specific assets in detail.
[0271] This embodiment has the following beneficial effects:
[0272] (1) Execution benefits
[0273] Improve the inventory efficiency: The present invention greatly improves the efficiency of road asset inventory through automated data collection and processing technologies. Compared with the traditional manual inventory method, it can complete the inventory work of large-scale road assets in a short time. For example, for a 100-kilometer-long highway, the present invention can complete data collection and processing within a few hours, while manual inventory may take several weeks.
[0274] Reduce labor costs: Reduce the dependence on a large number of manual measurement and data processing personnel, and reduce labor costs. The automated operation of data collection equipment and intelligent processing algorithms make the entire inventory process more efficient and economical.
[0275] Real-time update and maintenance: Can obtain the change information of road assets in real time and update the inventory data in a timely manner. When road assets are damaged, replaced, or newly added, etc., the system can respond quickly, update the asset information, ensure the timeliness and accuracy of the inventory data, and provide timely and accurate basis for road maintenance and management.
[0276] (2) Economic benefits
[0277] Reduce resource waste: Accurate road asset inventory data helps to reasonably plan and allocate resources, avoiding resource waste caused by inaccurate data. For example, when carrying out road maintenance and updates, according to accurate asset information, the maintenance plan and material procurement can be reasonably arranged to reduce unnecessary expenses.
[0278] Improve the scientific nature of management decisions: Provide comprehensive, accurate, and real-time road asset information for road management departments to help managers make more scientific and reasonable decisions. For example, when formulating road development plans, the layout and quantity of facilities such as signs and streetlights can be reasonably planned based on the results of asset inventories, improving the usage efficiency and safety of roads.
[0279] Enhance the value of road assets: Through an efficient inventory and display system, road assets can be better managed and maintained, extending the service life of assets and enhancing the overall value of road assets. A good condition of road assets can not only improve the traffic capacity and service level of roads but also create good conditions for the economic development of surrounding areas.
[0280] (3)Social benefits
[0281] Improve road safety: Accurate road asset information helps to promptly discover and handle potential safety hazards, such as damaged signs and malfunctioning streetlights. Through rapid response and repair, the occurrence of traffic accidents can be effectively reduced, ensuring the life and property safety of road users.
[0282] Enhance public satisfaction: Provide an intuitive and accurate road asset display platform for the public to facilitate their understanding of the condition and distribution of road facilities. For example, the public can query information about nearby signs, bus stops, etc. through the platform, improving the convenience and satisfaction of travel.
[0283] Promote the construction of smart cities: As an innovative technology in the field of road engineering, the present invention provides strong support for the construction of smart cities. Through integration and collaboration with other intelligent systems in the city, the intelligent management and operation of urban infrastructure can be realized, enhancing the overall intelligent level of the city.
[0284] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for road asset inventory based on three-dimensional Gaussian splashing, characterized in that, Including: Collecting multi-source data of road assets; wherein, the multi-source data includes: image data, IMU data, and lidar data; Preprocessing the multi-source data to obtain a mask grayscale mask image, camera intrinsics, camera pose, and sparse point cloud; Training a three-dimensional Gaussian splash model using the preprocessed data; including: Initializing the sparse point cloud as a 3D Gaussian distribution, projecting it onto the image plane for rendering, calculating the loss based on the mask grayscale mask image, and using the gradient descent method to optimize the Gaussian distribution parameters. Through iterative optimization until the loss converges or reaches the maximum number of iterations, finally obtaining the trained model; Based on the trained three-dimensional Gaussian splash model, conducting an inventory of road assets; including: extracting road assets, analyzing asset status, GPS matching, and measuring asset dimensions; Conducting road asset extraction includes: Based on the trained three-dimensional Gaussian splash model, mapping Gaussian point features to high-dimensional feature vectors through an encoder, determining the Gaussian kernel class label using the grayscale mask image and a neural network discriminator, and performing spatial clustering on the same-class Gaussian kernels to obtain an independent asset model and numbering; Conducting asset status analysis includes: Projecting the Gaussian cluster of a single asset in the trained three-dimensional Gaussian splash model onto multiple camera views, classifying the rendered images, and determining the asset status in combination with a voting mechanism; Conducting GPS matching includes: Converting the center point coordinates of each Gaussian kernel in the trained three-dimensional Gaussian splash model into GPS coordinates, and calculating the average GPS coordinates of the asset, which is the real-world coordinates of the asset; Conducting asset dimension measurement includes: Obtaining the length, width, and height information of the asset by fitting the minimum bounding cube of the world coordinate distribution of each asset in the trained three-dimensional Gaussian splash model; Preprocessing the multi-source data includes: Performing semantic segmentation on the image data; Converting IMU data into camera pose; Matching lidar data and image data to obtain a sparse point cloud with color and position information; Performing semantic segmentation on the image data includes: Using a semantic segmentation neural network model based on the Transformer architecture to perform semantic segmentation on the image data; wherein, the semantic segmentation neural network model based on the Transformer architecture is trained using manually annotated and data-augmented historical image data, and the historical image data includes various road scenes and weather conditions; Saving the result of the model segmentation as a grayscale mask image, and generating a mask grayscale mask image with the same size as the grayscale mask image; wherein, the mask grayscale mask image only includes: vehicles, pedestrians, and sky.
2. The method for inventorying road assets based on three-dimensional Gaussian splashing according to claim 1, wherein Converting IMU data into camera pose includes: Based on IMU data, estimating attitude information through angular velocity integration and estimating position information through acceleration integration.
3. The method for inventorying road assets based on three-dimensional Gaussian splashing according to claim 1, wherein Matching lidar data and image data includes: Through extrinsic calibration, converting the point cloud in lidar data into the camera coordinate system; Using the camera intrinsic matrix to project the point cloud in the camera coordinate system onto the image plane; Converting the projected plane coordinates into normalized image coordinates and filtering out the points projected outside the image; Extract image key points from the image using the SIFT feature extraction algorithm; wherein, the image key points are pixel points or regions with significant features in the image; For each image key point, search for the nearest point in the projected point cloud; Based on these nearest points, form a sparse point cloud.
4. The method for road asset inventory based on three-dimensional Gaussian splash according to claim 1, characterized in that According to the mask grayscale mask image, calculate the loss including: Ignore the spatial area according to the mask grayscale mask map, and only train the static area. For each area of the image, introduce a weight w alpha , for the weight w of the dynamic area alpha = 0, for the weight w of the static area alpha = 1; among them, the dynamic areas include: vehicles, pedestrians, sky, and the static areas include: road surface, road assets, buildings, greenery; For the region where w alpha = 0, update C before calculating the loss gt : Loss calculation: Calculate the loss between the rendered image and the real image: Use the L1 loss function: Among them, C render is the rendered image, and C gt is the real image; Use SSIM Loss: SSIM Loss is defined as: The L1 Loss and the SSIM Loss are combined with weights to form the final loss: Where: L is the final loss, and λ dssim is the weight of the SSIM Loss, and its value ranges between [0, 1].
5. The method for road asset inventory based on three-dimensional Gaussian splash according to claim 1, wherein, Use the preprocessed data to train the 3D Gaussian splash model, specifically including: (1) Initialize each point in the sparse point cloud as a 3D Gaussian distribution; (2) For each camera pose, project the Gaussian distribution onto the image plane; (3) Render the image using the projected Gaussian distribution; (4) Based on the rendered image, calculate the loss according to the mask grayscale mask image; (5) According to the calculated loss, use the gradient descent method to optimize the parameters of the Gaussian distribution; (6) Repeat steps (2) to (5) until the loss converges or reaches the maximum number of iterations to obtain the trained model.
6. The method for road asset inventory based on 3D Gaussian splash according to claim 1, wherein, Performing road asset extraction includes: Based on the trained 3D Gaussian splash model, train an encoder to map the features of Gaussian points into high-dimensional feature vectors; For each Gaussian kernel, project the Gaussian kernel onto the corresponding grayscale mask image through the camera pose to obtain the class label of the Gaussian kernel; if the Gaussian kernel projects to different class labels under multiple perspectives, use a voting mechanism to determine its final class; assign a 32-bit class feature vector to each Gaussian kernel; Use the grayscale mask image as a supervision signal to train the class features through an optimization algorithm; Design a neural network with the input being the 32-bit class feature and the output being the class label; use the labeled Gaussian kernel data to train the discriminator, the loss function being the cross-entropy loss, and the optimization goal being to minimize the class prediction error; wherein, the Gaussian kernel data includes: class features and corresponding class labels; Input the 32-bit class features of each Gaussian kernel into the trained discriminator to obtain the class label of each Gaussian kernel; screen out the Gaussian kernels of the specified class according to the class label to generate a segmented 3D scene; According to the segmented 3D scene, perform spatial clustering on the Gaussian kernels of the extracted road assets to obtain the 3D Gaussian splash model of independent assets and number each asset; wherein, the extracted road assets include: signs, gantries, street lamps; Performing asset status analysis includes: Project the Gaussian clusters of individual assets in the trained 3D Gaussian splash model onto multiple camera perspectives, perform image classification on the rendered images using a classification model to obtain multiple classification results, and determine the final asset status analysis result according to the voting mechanism, and record the asset number and asset status; wherein, the classification model uses a fine-tuned convolutional neural network classification model.
7. A system for road asset inventory based on three-dimensional Gaussian splashing, characterized in that, For implementing the method according to any one of claims 1-6, the system includes: a data acquisition module, a data preprocessing module, a three-dimensional Gaussian splash model training module, and a three-dimensional Gaussian splash model analysis module; The data acquisition module is used to acquire multi-source data of road assets; The data preprocessing module is used to preprocess the multi-source data, and obtain a mask grayscale mask image, camera internal parameters, camera poses, and sparse point clouds; The three-dimensional Gaussian splash model training module is used to train a three-dimensional Gaussian splash model using the preprocessed data; it includes: Initializing the sparse point cloud as a 3D Gaussian distribution, projecting it onto the image plane for rendering, calculating the loss based on the mask grayscale mask image, and using the gradient descent method to optimize the Gaussian distribution parameters. Through iterative optimization until the loss converges or reaches the maximum number of iterations, finally obtain the trained model; The three-dimensional Gaussian splash model analysis module is used to conduct an inventory of road assets based on the trained three-dimensional Gaussian splash model; it includes: Conducting an inventory of road assets includes: extracting road assets, analyzing asset status, GPS matching, and measuring asset dimensions; Conducting road asset extraction includes: Based on the trained three-dimensional Gaussian splash model, mapping the Gaussian point features to high-dimensional feature vectors through an encoder, determining the Gaussian kernel class labels using the grayscale mask image and a neural network discriminator, and performing spatial clustering on the same-class Gaussian kernels to obtain an independent asset model and number it; Conducting asset status analysis includes: Projecting the Gaussian cluster of a single asset in the trained three-dimensional Gaussian splash model onto multiple camera views, classifying the rendered images, and combining the voting mechanism to determine the asset status; Conducting GPS matching includes: Converting the center point coordinates of each Gaussian kernel in the trained three-dimensional Gaussian splash model into GPS coordinates, and calculating the average GPS coordinates of the asset, which is the real-world coordinates of the asset; Conducting asset dimension measurement includes: Obtaining the length, width, and height information of the asset by fitting the minimum bounding cube of the world coordinate distribution of each asset in the trained three-dimensional Gaussian splash model.
8. The system for road asset inventory based on three-dimensional Gaussian splashing according to claim 7, wherein The data acquisition module includes: a 360 panoramic camera, a professional real-time kinematic measurement device with an inertial measurement unit, and a lidar sensor; each sensor is assembled and installed on the roof through a rigid structure; an industrial industrial computer is installed in the vehicle to be responsible for powering the equipment and controlling the acquisition work. The industrial industrial computer is equipped with GPU computing power and at the same time is equipped with a 5G module to upload the acquired data to the cloud server.
Citation Information
Patent Citations
One-map road asset detection method based on image and laser point cloud
CN118967959A
Dynamic real-time rendering method for large assembly scene based on three-dimensional Gaussian splashing
CN119229031A