Road asset checking method and system based on three-dimensional Gaussian splashing

Through the three-dimensional Gaussian splashing method and system, the problems of low inventory efficiency, large error, complex processing and poor display effects in the existing technology are solved, and efficient and accurate inventory and display are achieved, and scientific decision-making in road management is supported.

CN119940745AActive Publication Date: 2025-05-06SHANGHAI TONGLU CLOUD TRANSPORTATION TECH CO LTD

Patent Information

Application Number
CN202510429225.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The existing road asset inventory technology is inefficient, has large errors, complex data processing and poor display effect.

Method used

Using a three-dimensional Gaussian splashing method and system, by collecting multi-source data (image data, IMU data, lidar data), pre-processing and three-dimensional Gaussian splashing model training, to achieve efficient inventory and display of road assets.

Benefits of technology

It improves the efficiency and accuracy of road asset inventory, reduces labor costs, realizes real-time updates and displays, and supports road maintenance, management and planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940745A_ABST
    Figure CN119940745A_ABST
Patent Text Reader

Abstract

The invention relates to a road asset checking method and system based on three-dimensional Gaussian splashing, and the method is characterized in that the method comprises the steps: collecting the multi-source data of road assets; wherein the multi-source data comprises image data, IMU (Inertial Measurement Unit) data and laser radar data; performing preprocessing on the multi-source data; performing three-dimensional Gaussian splash model training by using the pre-trained data; and performing road asset checking based on the trained three-dimensional Gaussian splashing model. The road assets can be efficiently and accurately checked and displayed, the road management efficiency and quality are improved, and powerful support is provided for road maintenance, management and planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of road engineering, and in particular to a method and system for road asset inventory based on three-dimensional Gaussian splashing. Background Art

[0002] In the field of road engineering and highway engineering, the inventory and display of road assets are crucial for the maintenance, management and planning of roads. Road assets include infrastructure such as signs, street lights, gantries, guardrails, and traffic lights. Their accurate inventory and visual display can provide important decision-making support for road management departments. At present, road asset inventories are mainly carried out by staff through field visits, using measuring tools (such as tape measures, rangefinders, etc.) to manually measure and record road assets. In addition, there are individual cases that use LiDAR point cloud scanning technology to scan the road environment through vehicle-mounted or airborne LiDAR equipment to generate high-precision three-dimensional point cloud data, and then extract road asset information.

[0003] Manual inventory static measurement: Workers conduct on-site visits and use measuring tools (such as tape measures, distance meters, etc.) to manually measure and record road assets. Inefficiency: Manual inventory is time-consuming and labor-intensive, and it is difficult to meet the needs of large-scale road asset inventory. Human error: The measurement results are easily affected by human factors, resulting in inaccurate data. Low safety: Workers need to measure on the road, which poses certain safety risks. Poor display effect: The display effect is not good.

[0004] LiDAR point cloud scanning technology: Scan the road environment through vehicle-mounted or airborne LiDAR equipment to generate high-precision three-dimensional point cloud data, and then extract road asset information. Difficult to update dynamically: Lack of lightweight representation methods to support incremental updates. Low efficiency: Traditional point cloud processing requires complex registration and post-processing, point cloud data occupies a large storage space, and the real-time display is poor. Summary of the invention

[0005] The purpose of the present invention is to solve the problems of low efficiency, large errors, complex data processing and poor display effect existing in the existing road asset inventory and display technology, and to provide a method and system for road asset inventory and display based on three-dimensional Gaussian splashing, which can efficiently and accurately inventory and display road assets, improve the efficiency and quality of road management, and provide strong support for road maintenance, management and planning.

[0006] To achieve the above object, the present invention provides the following solutions:

[0007] A method for road asset inventory based on three-dimensional Gaussian splashing, comprising:

[0008] Collect multi-source data of road assets; wherein the multi-source data includes: image data, IMU data, and LiDAR data;

[0009] Preprocessing the multi-source data;

[0010] Use pre-trained data to train the 3D Gaussian splash model;

[0011] Road asset inventory is performed based on the trained 3D Gaussian splash model.

[0012] Optionally, preprocessing the multi-source data includes:

[0013] Performing semantic segmentation on the image data;

[0014] Convert IMU data to camera pose;

[0015] Match the lidar data and image data to obtain a sparse point cloud with color and location information.

[0016] Optionally, performing semantic segmentation on the image data includes:

[0017] The semantic segmentation neural network model based on the Transformer architecture is used to perform semantic segmentation on the image data. The semantic segmentation neural network model based on the Transformer architecture uses historical image data with manual annotation and data enhancement for model training. The historical image data includes a variety of road scenes and weather conditions.

[0018] The result of the model segmentation is saved as a grayscale mask image, and a mask grayscale mask image is generated, the size of which is the same as that of the grayscale mask image; wherein the mask grayscale mask image only includes: vehicles, pedestrians, and the sky.

[0019] Optionally, converting IMU data to camera pose includes:

[0020] Based on the IMU data, the attitude information is estimated by integrating the angular velocity and the position information is estimated by integrating the acceleration.

[0021] Optionally, matching the laser radar data with the image data includes:

[0022] Through external parameter calibration, the point cloud in the lidar data is converted to the camera coordinate system;

[0023] Use the camera intrinsic parameter matrix to project the point cloud in the camera coordinate system to the image plane;

[0024] Convert the projected plane coordinates into normalized image coordinates and filter out the points projected outside the image;

[0025] The image key points are extracted from the image using the SIFT feature extraction algorithm; the image key points are pixels or areas with significant features in the image;

[0026] For each image keypoint, search for the nearest point in the projected point cloud;

[0027] Based on these nearest points, a sparse point cloud is formed.

[0028] Optionally, using the pre-trained data to train the three-dimensional Gaussian splash model includes:

[0029] Use the mask grayscale mask image, camera intrinsic parameters, camera pose, and sparse point cloud to train the 3D Gaussian splash model;

[0030] The three-dimensional Gaussian splash model is trained using the mask grayscale mask image, camera intrinsic parameters, camera pose, and sparse point cloud, including:

[0031] (1) Initialize each point in the sparse point cloud as a 3D Gaussian distribution;

[0032] (2) For each camera pose, project the Gaussian distribution onto the image plane.

[0033] (3) Render the image using the projected Gaussian distribution;

[0034] (4) Based on the rendered image, calculate the loss according to the mask grayscale mask map;

[0035] (5) Based on the calculated loss, use the gradient descent method to optimize the parameters of the Gaussian distribution;

[0036] (6) Repeat steps (2) to (5) until the loss converges or the maximum number of iterations is reached to obtain the trained model.

[0037] Optionally, calculating the loss based on the mask grayscale mask image;

[0038] According to the grayscale mask image, the spatial area is ignored and only the static area is trained. For each area of ​​the image, a weight w is introduced alpha , for the weight w of the dynamic area alpha =0, weight w for static area alpha =1; dynamic areas include vehicles, pedestrians, and the sky, and static areas include road surfaces, road assets, buildings, and greenery;

[0039] For w alpha = 0, update C before calculating the loss gt :

[0040]

[0041] Loss calculation:

[0042] Calculate the loss between the rendered image and the real image:

[0043] Using L1 loss function:

[0044] 1

[0045] Among them, C render is the rendered image, C gt is a real image;

[0046] Using SSIM Loss:

[0047] SSIM Loss is defined as:

[0048]

[0049] Weighted combination of L1 Loss and SSIM Loss:

[0050]

[0051] Where: L is the final loss, λ dssim is the weight of SSIM Loss, and its value is between [0,1].

[0052] Optionally, based on the trained three-dimensional Gaussian splash model, performing a road asset inventory includes:

[0053] Based on the trained 3D Gaussian splash model, road asset extraction, asset status analysis, asset size measurement, and GPS matching are performed;

[0054] Road asset extraction includes:

[0055] Based on the trained 3D Gaussian splash model, an encoder is trained to map the features of Gaussian points into high-dimensional feature vectors.

[0056] For each Gaussian kernel, project the Gaussian kernel onto the corresponding grayscale mask image through the camera pose to obtain the category label of the Gaussian kernel; if the Gaussian kernel is projected to different category labels under multiple viewing angles, a voting mechanism is used to determine its final category; a 32-bit category feature vector is assigned to each Gaussian kernel;

[0057] Use the grayscale mask image as a supervisory signal and train the category features through the optimization algorithm;

[0058] Design a neural network with 32-bit category features as input and category labels as output; use labeled Gaussian kernel data to train the discriminator, the loss function is cross entropy loss, and the optimization goal is to minimize the category prediction error; the Gaussian kernel data includes: category features and corresponding category labels;

[0059] The 32-bit category feature of each Gaussian kernel is input into the trained discriminator to obtain the category label of each Gaussian kernel; the Gaussian kernel of the specified category is filtered out according to the category label to generate the segmented 3D scene;

[0060] According to the segmented 3D scene, the Gaussian kernel of the extracted road assets is spatially clustered to obtain a 3D Gaussian splash model of independent assets, and each asset is numbered; the extracted road assets include signs, gantries, and street lights;

[0061] Performing asset status analysis includes:

[0062] The Gaussian cluster of a single asset is projected onto k camera perspectives, and the rendered image is classified using a classification model to obtain k classification results. The final asset status analysis result is determined based on a voting mechanism, and the asset number and asset status are recorded. The classification model uses a fine-tuned convolutional neural network classification model.

[0063] GPS matching includes:

[0064] Based on the extracted three-dimensional Gaussian splash model information of the asset, the average GPS coordinates of the Gaussian kernel of each asset are calculated, i.e., the real-world coordinates of the asset;

[0065] Taking asset size measurements includes:

[0066] According to the world coordinate distribution of each asset, the smallest circumscribed cube is fitted, and the length, width and height of the cube are considered to be the length, width and height information of the asset.

[0067] A system for road asset inventory based on three-dimensional Gaussian splashing, the system comprising: a data acquisition module, a data preprocessing module, a three-dimensional Gaussian splashing model training module, and a three-dimensional Gaussian splashing model analysis module;

[0068] The data acquisition module is used to collect multi-source data of road assets;

[0069] The data preprocessing module is used to preprocess the multi-source data;

[0070] The three-dimensional Gaussian splash model training module is used to perform three-dimensional Gaussian splash model training using pre-trained data;

[0071] The three-dimensional Gaussian splash model analysis module is used to perform a road asset inventory based on the trained three-dimensional Gaussian splash model.

[0072] Optionally, the data acquisition module includes: a 360 panoramic camera, a professional-grade real-time motion measurement device with an inertial measurement unit, and a lidar sensor; each sensor is assembled and installed on the roof through a rigid structure; an industrial computer is installed in the vehicle to be responsible for powering the equipment and controlling the acquisition work. The industrial computer is equipped with GPU computing power and a 5G module, which can upload the collected data to a cloud server.

[0073] The beneficial effects of the present invention are:

[0074] The present invention greatly improves the efficiency of road asset inventory through automated data collection and processing technology. Compared with the traditional manual inventory method, it can complete the inventory of large-scale road assets in a short time. It reduces the dependence on a large number of manual measurement and data processing personnel and reduces labor costs.

[0075] It can obtain the change information of road assets in real time and update the inventory data in time. When road assets are damaged, replaced or added, the system can respond quickly and update the asset information to ensure the timeliness and accuracy of the inventory data, providing a timely and accurate basis for road maintenance and management.

[0076] Accurate road asset inventory data helps to plan and allocate resources rationally and avoid waste of resources due to inaccurate data.

[0077] Provide comprehensive, accurate and real-time road asset information to road management departments, helping managers make more scientific and reasonable decisions.

[0078] Through an efficient inventory and display system, road assets can be better managed and maintained, the service life of assets can be extended, and the overall value of road assets can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0080] Figure 1 A schematic diagram of a method flow for road asset inventory based on three-dimensional Gaussian splashing according to an embodiment of the present invention;

[0081] Figure 2 Schematic diagram of system installation according to an embodiment of the present invention. DETAILED DESCRIPTION

[0082] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0083] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0084] 3D Gaussian splatting (3DGS) is an advanced 3D modeling and visualization technology that constructs 3D scenes by representing images as tiny elliptical spots, namely Gaussian kernels. Gaussian kernels can be stretched, compressed, and given color and transparency, and finally merged to form a continuous surface. The core of Gaussian splatting is to use 3D Gaussian functions as the basic elements of the scene, and to achieve high-quality 3D scene reconstruction and new perspective synthesis by optimizing the parameters of these Gaussian functions. 3DGS has the advantages of high efficiency, flexibility, high-quality rendering, and real-time performance.

[0085] like Figure 1 As shown, this embodiment proposes a method for road asset inventory based on three-dimensional Gaussian splashing, including:

[0086] Collect multi-source data of road assets; wherein the multi-source data includes: image data, IMU data, and LiDAR data;

[0087] Preprocessing the multi-source data;

[0088] Use pre-trained data to train the 3D Gaussian splash model;

[0089] Road asset inventory is performed based on the trained 3D Gaussian splash model.

[0090] Specifically, in this embodiment, during the data collection phase:

[0091] Data acquisition equipment such as 360 cameras, IMU, RTK, and laser radar are installed on the roof to ensure the stability of the equipment and the accuracy of data acquisition. The car collects data on the target road at a constant speed (30km / h). During the collection process, the original pictures, camera posture, laser point cloud and other data are recorded, and accurate geographic location information is obtained through RTK.

[0092] An acquisition program is deployed on the industrial computer. After the acquisition program is started, the PPS (Pulse PerSecond) pulse signal of the RTK-GPS module is received, and the current system time T is recorded as the synchronization reference time. For sensors that support hardware triggering (such as cameras and lidars), the PPS pulse is directly used as a trigger signal, and the sensor immediately starts to collect a frame of data. For sensors that do not support hardware triggering (such as RTK), the acquisition program immediately sends a software instruction after detecting the PPS pulse, requiring the sensor to start collecting data. Each sensor starts collecting data after being triggered and timestamps each frame of data. Align the data of all sensors according to the timestamp T to ensure that the data frames at the same time can match. Store the collected data in files or databases by frame.

[0093] Further, preprocessing the multi-source data includes;

[0094] Performing semantic segmentation on the image data;

[0095] Convert IMU data to camera pose;

[0096] Match the lidar data and image data to obtain a sparse point cloud with color and location information.

[0097] Furthermore, performing semantic segmentation on the image data includes:

[0098] The semantic segmentation neural network model based on the Transformer architecture is used to perform semantic segmentation on the image data. The semantic segmentation neural network model based on the Transformer architecture uses historical image data with manual annotation and data enhancement for model training. The historical image data includes a variety of road scenes and weather conditions.

[0099] The result of the model segmentation is saved as a grayscale mask image, and a mask grayscale mask image is generated, the size of which is the same as that of the grayscale mask image; wherein the mask grayscale mask image only includes: vehicles, pedestrians, and the sky.

[0100] Furthermore, converting IMU data into camera pose includes:

[0101] Based on the IMU data, the attitude information is estimated by integrating the angular velocity and the position information is estimated by integrating the acceleration.

[0102] Furthermore, matching the laser radar data with the image data includes:

[0103] Through external parameter calibration, the point cloud in the lidar data is converted to the camera coordinate system;

[0104] Use the camera intrinsic parameter matrix to project the point cloud in the camera coordinate system to the image plane;

[0105] Convert the projected plane coordinates into normalized image coordinates and filter out the points projected outside the image;

[0106] Use SIFT feature extraction algorithm to extract image key points from the image;

[0107] For each image keypoint, search for the nearest point in the projected point cloud;

[0108] Based on these nearest points, a sparse point cloud is formed.

[0109] Among them, the key points of an image are pixels or regions with significant features in the image; the local area around the key points has a unique texture, shape or structure in the image and can be distinguished from other areas; the key points are robust to image rotation, scaling, lighting changes, etc., and can be stably detected under different conditions. The same key points can be detected repeatedly in different images.

[0110] Specifically, in this embodiment, after the acquisition program acquires the multi-sensor data, the analysis program immediately pre-processes the data. The data pre-processing specifically includes:

[0111] (1) Perform semantic segmentation on the original image

[0112] The semantic segmentation neural network model (SegFormer) based on the Transformer architecture is used to perform semantic segmentation on the collected original images. The model uses previously collected data for manual annotation and data enhancement to fine-tune the model. The training data includes a variety of road scenes and weather conditions, etc. It can effectively segment road assets and objects such as pedestrians, vehicles, sky, signs, traffic lights, gantries, etc. in various scenes. The average intersection over union (mIoU) of the model is >95%.

[0113] The result of model segmentation is saved in the form of a grayscale mask image of size H*W, where H and W are the width and height of the original image, and the value of the mask image represents the segmentation category ID of this pixel coordinate. At the same time, a mask grayscale mask image needs to be generated, with the same size as the grayscale mask image. The mask grayscale mask image only contains categories such as vehicles, pedestrians, and the sky, because vehicles and pedestrians are dynamic objects during data collection and have a relative motion relationship with the collection vehicle, while the 3DGS reconstruction of road assets is mainly for static backgrounds. Therefore, the mask image is generated here to avoid interference from dynamic objects during subsequent training and affect the modeling quality. The sky is not within the scope of 3DGS reconstruction, but it accounts for a large proportion of the image, which will also affect the quality of reconstruction, so it is filtered together with dynamic objects.

[0114] (2) Convert IMU data into camera pose information

[0115] IMU data can provide the acceleration a(t) and angular velocity ω(t) of the camera at each time point. The initial position of the camera is obtained through external parameter calibration during calibration, namely the initial attitude q(0) and initial position p(0).

[0116] (a) Estimation of attitude by integrating angular velocity.

[0117]

[0118] q(t) is the posture at the current time t, represents the multiplication of quaternions, Δq(t) is the incremental quaternion obtained by integrating the angular velocity, which can be calculated by the following formula:

[0119]

[0120] Here, qω(t) is the quaternion calculated from the angular velocity ω(t) and the time interval Δt.

[0121]

[0122] Δt is the time interval, that is, the data collection interval.

[0123] (b) Estimation of position by integrating acceleration.

[0124] The position is the integral of the velocity over time, and the position p(t) can be updated by the velocity v(t).

[0125]

[0126] Where p(t) is the camera position at the current time t, and v(t) is calculated as follows:

[0127]

[0128] a(t) is the acceleration collected by IMU;

[0129] Through the above, the camera's position p(t) and orientation q(t) can be obtained from the IMU data (acceleration, angular velocity).

[0130] (3) Constructing sparse point cloud

[0131] Get a camera image synchronized with the lidar and project the point cloud onto the image plane to get the color information:

[0132] Point cloud original coordinates P lidar =(x lidar ,y lidar ,z lidar ).

[0133] Camera coordinate system: Convert the point cloud to the camera coordinate system through external parameter calibration:

[0134]

[0135] in: is the rotation matrix from lidar to camera; is the translation vector from the laser radar to the camera; P cam =(x cam ,y cam ,z cam ) is the point cloud coordinate in the camera coordinate system.

[0136] Use the camera intrinsic parameter matrix K to project the point cloud in the camera coordinate system to the image plane:

[0137] p img =K P cam

[0138] Among them, K is the camera intrinsic parameter matrix:

[0139]

[0140] P cam is the point cloud coordinates in the camera coordinate system (homogeneous coordinates):

[0141]

[0142] p img are the image plane coordinates after projection (homogeneous coordinates):

[0143]

[0144] Where (u,v) is the normalized image coordinate:

[0145]

[0146] Convert the projected homogeneous coordinates to normalized image coordinates:

[0147]

[0148]

[0149] Among them, (u,v) is the pixel coordinate of the point cloud on the image plane.

[0150] Filter out points that are projected outside the image:

[0151] 0≤u <w,0≤v<h

[0152] Where w and h are the width and height of the image respectively.

[0153] Use SIFT feature extraction algorithm to extract key points from the image:

[0154]

[0155] Among them, (u i ,v i ) is the pixel coordinate of the i-th key point.

[0156] After projecting the lidar point cloud onto the image plane, find the closest point to the image keypoint:

[0157] For each image key point (u i ,v i ), search for the nearest point in the projected point cloud:

[0158]

[0159] These closest points are retained to form a sparse point cloud.

[0160] Each point in a sparse point cloud contains the following information:

[0161] 3D coordinates: P cam =(x cam ,y cam ,z cam ).

[0162] Color information: RGB values ​​(r,g,b) extracted from the image.

[0163] Further, using the pre-trained data to train the three-dimensional Gaussian splash model includes:

[0164] Use the mask grayscale mask image, camera intrinsic parameters, camera pose, and sparse point cloud to train the 3D Gaussian splash model;

[0165] The three-dimensional Gaussian splash model is trained using the mask grayscale mask image, camera intrinsic parameters, camera pose, and sparse point cloud, including:

[0166] (1) Initialize each point in the sparse point cloud as a 3D Gaussian distribution;

[0167] (2) For each camera pose, project the Gaussian distribution onto the image plane.

[0168] (3) Render the image using the projected Gaussian distribution;

[0169] (4) Based on the rendered image, calculate the loss according to the mask grayscale mask map;

[0170] (5) Based on the calculated loss, use the gradient descent method to optimize the parameters of the Gaussian distribution;

[0171] (6) Repeat steps (2) to (5) until the loss converges or the maximum number of iterations is reached to obtain the trained model.

[0172] Specifically, in this embodiment, three-dimensional Gaussian splashing (3DGS) training includes the following steps:

[0173] The 3DGS model is trained using the mask grayscale mask image, camera intrinsic parameters, camera pose, and sparse point cloud obtained through data preprocessing.

[0174] (1) Initialize Gaussian distribution

[0175] Initialize each point in the sparse point cloud as a 3D Gaussian distribution.

[0176] Each Gaussian distribution is defined by the following parameters:

[0177] Position μ=(x,y,z): point cloud coordinates.

[0178] Covariance matrix Σ: Initialized to a small diagonal matrix.

[0179] Color c=(r,g,b): Color information extracted from the point cloud.

[0180] Opacity α: Initialized to 1.

[0181] (2) Projecting Gaussian distribution onto the image plane

[0182] For each camera pose, project the 3D Gaussian distribution onto the image plane:

[0183] a) Convert the Gaussian distribution position μ to the camera coordinate system:

[0184] μ cam =R μ+t

[0185] b) μ cam Project onto the image plane:

[0186] μ img =K μ cam

[0187] Transform the covariance matrix Σ to the image plane:

[0188] Σ img =J Σ J T

[0189] where J is the Jacobian matrix of the projection.

[0190] (3) Rendering images

[0191] Render an image using a projected Gaussian distribution:

[0192] For each pixel (u,v), calculate its color value:

[0193]

[0194] Where: ci is the color of the i-th Gaussian distribution; αi is the opacity of the i-th Gaussian distribution; is the probability density function of the i-th Gaussian distribution on the image plane.

[0195] (4) Dynamic object masking and loss calculation

[0196] According to the grayscale mask image, the spatial area is ignored to ensure that only the static area is trained. For each area of ​​the image, a weight w is introduced alpha , the weight w for dynamic areas (vehicles, pedestrians, sky) alpha =0, weight w for static area alpha = 1. For w alpha = 0, update C before calculating the loss gt :

[0197]

[0198] Calculating the dynamic object mask is equivalent to labeling the dynamic area of ​​the rendered image, so that the dynamic object part is ignored when calculating the loss;

[0199] Here, the dynamic part of the real image is modified into the dynamic part of the rendered image; alpha It is the weight used to distinguish dynamic and static objects in the image. The weight is obtained by semantic segmentation of the image by the Transformer-based semantic segmentation neural network model. The Transformer model is specially trained. After processing, the dynamic area in the image has Walpha=0 (dynamic areas include: vehicles, pedestrians, and the sky); the static area in the image has Walpha=1 (static areas include: road surface, road assets, buildings, and greenery).

[0200] Loss calculation:

[0201] Calculate the loss between the rendered image and the real image:

[0202] Using L1 loss function:

[0203] 1

[0204] Among them, C render is the rendered image, C gt is a real image.

[0205] Using SSIM Loss:

[0206] SSIM (Structural Similarity Index Measure) calculation formula:

[0207]

[0208] in: , are the means of the rendered image and the real image respectively; , are the variances of the rendered image and the real image, respectively; is the covariance of the rendered image and the real image; c1, c2 are constants used to stabilize the calculation.

[0209] SSIM Loss is defined as:

[0210]

[0211] Weighted combination of L1 Loss and SSIM Loss:

[0212]

[0213] Where: L is the final loss, λ dssim is the weight of SSIM Loss, and its value is between [0,1].

[0214] L is the final loss. L1 loss is used to measure the difference between the rendered image and the target image at the pixel level. It is insensitive to outliers, can provide a smoother optimization process, and is suitable for capturing image details. SSIM loss is used to measure the difference in perceptual quality between the rendered image and the target image. It can better reflect the perceptual characteristics of the human visual system and is suitable for capturing the overall structure and perceptual quality of the image.

[0215] The final loss L is a weighted combination of L1 loss and SSIM loss, which is used to optimize both the detail information and the perceptual quality of the image. By adjusting λdssim, the contribution of the two can be flexibly balanced according to the task requirements.

[0216] (5) Optimizing Gaussian distribution parameters

[0217] Use gradient descent to optimize the parameters of the Gaussian distribution (mean μ, covariance Σ, color c, opacity α):

[0218] Update formula:

[0219]

[0220] Here, θ is the parameter of the Gaussian distribution and η is the learning rate.

[0221] (6) Iterative training

[0222] Repeat steps (2) to (5) until the loss converges or the maximum number of iterations is reached. Store the 3DGS model as a .ply file.

[0223] Furthermore, based on the trained 3D Gaussian splash model, the road asset inventory includes:

[0224] Based on the trained 3D Gaussian splash model, road asset extraction, asset status analysis, asset size measurement, and GPS matching are performed;

[0225] (1) Road asset extraction

[0226] Based on the trained 3DGS model, an encoder E is trained to map the features of Gaussian points into high-dimensional feature vectors.

[0227]

[0228] in It is a Gaussian point The feature vector of , D is the feature dimension, here D is 32; the encoder E adopts a multi-layer perceptron (MLP) structure.

[0229] For each Gaussian kernel, project it onto the corresponding grayscale mask image through the camera pose to obtain its category label. If the Gaussian kernel is projected onto different category labels under multiple viewing angles, a voting mechanism is used to determine its final category. A 32-bit category feature vector is assigned to each Gaussian kernel.

[0230] Using the grayscale mask image as a supervisory signal, the category feature is trained through an optimization algorithm (cross entropy loss) so that it can accurately represent the category of the Gaussian kernel.

[0231] Design a simple neural network (such as MLP) with 32-bit category features as input and category labels (0-255) as output. Use labeled Gaussian kernel data (category features and corresponding category labels) to train the discriminator, with the loss function being cross entropy loss, and the optimization goal being to minimize the category prediction error.

[0232] The 32-bit category feature of each Gaussian kernel is input into the trained discriminator to obtain its category label. The Gaussian kernel of the specified category is filtered out according to the category label to generate the segmented 3D scene.

[0233] The Gaussian kernel of the extracted road assets such as signs, gantries, and street lights is spatially clustered to obtain the 3DGS model of independent assets (such as a single street light), and each asset is numbered.

[0234] (2) Asset status analysis

[0235] Using a fine-tuned convolutional neural network (CNN) classification model, images can be classified into intact, damaged, deformed, dirty, occluded, etc.

[0236] The Gaussian cluster of a single asset is projected onto k camera perspectives, and the rendered image is classified using a classification model to obtain k classification results. The final asset status analysis result (such as guardrail-intact, guardrail-deformed, guardrail-dirty, etc.) is determined based on a voting mechanism, and the asset number and asset status are recorded.

[0237] (3) GPS matching

[0238] The center point (x, y, z) of each Gaussian kernel in the 3DGS model is the spatial coordinate of the LiDAR point cloud. Since the LiDAR has been calibrated with GPS, the spatial coordinates can be mapped to GPS coordinates through affine transformation.

[0239]

[0240] in is the rotation matrix from lidar to GPS.

[0241] is the translation vector from LiDAR to GPS.

[0242] It is the world coordinate of the point cloud in the GPS coordinate system.

[0243] Based on the extracted 3DGS information of the assets, the average GPS coordinates of the Gaussian kernel of each asset are calculated, that is, the real-world coordinates of the assets.

[0244] (4) Asset size measurement

[0245] According to the world coordinate distribution of each asset, the smallest circumscribed cube is fitted, and the length, width and height of the cube are considered to be the length, width and height information of the asset.

[0246] (5) 3DGS noise reduction

[0247] In order to improve the aesthetics and accuracy of the rendering effect, the KNN algorithm is used to remove outlier Gaussian kernels to ensure the smoothness and consistency of the segmentation results.

[0248] After completing the above steps, the category, world coordinates, size information and asset status of each asset are obtained. Based on this information, the number of assets of each type is counted, and the analysis results and 3DGS model are uploaded to the cloud server.

[0249] The method of this embodiment also includes: data rendering and display:

[0250] By uploading the analysis and modeling results to the cloud server, the 3DGS data can be rendered in real time on the display platform using rasterization technology, and the discrete Gaussian points can be removed through the KNN algorithm to improve the aesthetics and accuracy of the rendering effect. On the display platform, the 3DGS rendering effect, road asset statistics, asset number and real coordinates are displayed, and an asset query function is provided to facilitate users to view and analyze specific assets in detail.

[0251] This embodiment also discloses a system for road asset inventory based on three-dimensional Gaussian splashing, the system comprising: a data acquisition module, a data preprocessing module, a three-dimensional Gaussian splashing model training module, and a three-dimensional Gaussian splashing model analysis module;

[0252] The data acquisition module is used to collect multi-source data of road assets;

[0253] The data preprocessing module is used to preprocess the multi-source data;

[0254] The three-dimensional Gaussian splash model training module is used to perform three-dimensional Gaussian splash model training using pre-trained data;

[0255] The three-dimensional Gaussian splash model analysis module is used to perform a road asset inventory based on the trained three-dimensional Gaussian splash model.

[0256] Furthermore, the data acquisition module of this embodiment adopts a comprehensive set of data acquisition equipment, including a 360-degree panoramic camera, a professional-grade real-time motion measurement system (RTK) with an inertial measurement unit (IMU), a laser radar and other sensors. Each sensor is assembled and installed on the roof through a rigid structure. The industrial computer installed in the car is responsible for powering the equipment and controlling the acquisition work. The industrial computer is equipped with GPU computing power, which can accelerate data processing. At the same time, it is equipped with a 5G module, which can upload the collected data to the cloud server.

[0257] In this embodiment, sensor calibration and coordinate alignment:

[0258] Since the installation positions of the sensors are fixed and their relative positions are known, the following can be obtained before installation: (1) camera intrinsic parameters (focal length f x , f y and principal point offset c x , c y, distortion coefficient dist);

[0259] (2) Coordinate transformation from lidar to camera: rotation matrix and translation vector Obtained through calibration measurements.

[0260] After the equipment is installed, the camera external parameters and RTK-GPS need to be calibrated:

[0261] Use a 1m*1m*1m cube calibration board, and install high-reflectivity markers (such as reflective stickers or reflective balls) at each corner to facilitate lidar recognition. Draw a checkerboard pattern on each face of the cube to facilitate camera recognition. There is a fixed point on the top of the cube for installing the GPS antenna to ensure a stable signal.

[0262] Calibration process:

[0263] (1) Place the calibration object in the scene and ensure that the GPS signal is good.

[0264] (2) Use RTK-GPS to measure the coordinates of the GPS antenna on the top of the calibration object.

[0265] (3) Use LiDAR to scan the calibration object and extract the point cloud coordinates of the reflective mark. Combined with the GPS data of the calibration object, use the least squares method or SVD (singular value decomposition) to calculate the rotation matrix and translation vector , and use optimization algorithms (such as Levenberg-Marquardt) to further improve the accuracy, that is, to obtain the conversion relationship from the lidar coordinate system to the GPS coordinate system.

[0266] (4) Use a 360° camera to photograph the calibration object, extract the coordinates of the checkerboard corner points, and calculate the camera's external parameters.

[0267] During the data collection phase, 360 cameras, IMU, RTK, LiDAR and other data collection equipment are installed on the roof. Figure 2 As shown, the stability of the equipment and the accuracy of data collection are ensured. The car collects data on the target road at a constant speed (30km / h). During the collection process, the original image, camera position, laser point cloud and other data are recorded, and accurate geographic location information is obtained through RTK.

[0268] The system architecture of this embodiment includes a data acquisition module, a data preprocessing module, a 3DGS training module, a 3DGS model analysis module, an asset database and management platform, and a user interaction module. The modules work together through efficient data transmission and communication mechanisms to ensure stable operation and efficient processing of the entire system.

[0269] Asset database and management platform:

[0270] After the cloud server obtains the Gaussian splash model and attribute metadata, it uses WebGL to render and display the 3DGS data in real time. At the same time, it can perform differential updates (update Gaussian parameters in some areas) based on the nearest matching of the asset's GPS coordinates, which ensures that the platform rendering results are the latest inventory modeling results. Finally, the display platform can not only display the 3DGS rendering effect, but also display the statistical information of road assets, number each asset, and display its real coordinates. In addition, it can also query based on the asset number, and after the query, the area near the asset is rendered in real time, which is convenient for users to view, analyze and mark maintenance records of specific assets in detail.

[0271] This embodiment has the following beneficial effects:

[0272] (1) Implementation benefits

[0273] Improve inventory efficiency: The present invention greatly improves the efficiency of road asset inventory through automated data collection and processing technology. Compared with the traditional manual inventory method, it can complete the inventory of large-scale road assets in a short time. For example, for a highway with a length of 100 kilometers, the present invention can complete data collection and processing in a few hours, while manual inventory may take several weeks.

[0274] Reduce labor costs: Reduce the reliance on a large number of manual measurement and data processing personnel, and reduce labor costs. The automated operation and intelligent processing algorithms of data acquisition equipment make the entire inventory process more efficient and economical.

[0275] Real-time update and maintenance: It can obtain the change information of road assets in real time and update the inventory data in time. When road assets are damaged, replaced or added, the system can respond quickly and update the asset information to ensure the timeliness and accuracy of the inventory data, providing a timely and accurate basis for road maintenance and management.

[0276] (2) Economic benefits

[0277] Reduce resource waste: Accurate road asset inventory data helps to rationally plan and allocate resources, avoiding resource waste caused by inaccurate data. For example, when repairing and renewing roads, you can reasonably arrange maintenance plans and material procurement based on accurate asset information, reducing unnecessary expenses.

[0278] Improve the scientific nature of management decisions: Provide comprehensive, accurate and real-time road asset information to road management departments to help managers make more scientific and reasonable decisions. For example, when formulating road development plans, the layout and quantity of signs, street lights and other facilities can be reasonably planned based on the results of asset inventory to improve the efficiency and safety of road use.

[0279] Improve the value of road assets: Through an efficient inventory and display system, road assets can be better managed and maintained, the service life of assets can be extended, and the overall value of road assets can be improved. Good road asset conditions can not only improve the road's traffic capacity and service level, but also create good conditions for the economic development of surrounding areas.

[0280] (3) Social benefits

[0281] Improve road safety: Accurate road asset information helps to timely discover and deal with potential safety hazards, such as damaged signs and faulty street lights. Through rapid response and repair, it can effectively reduce the occurrence of traffic accidents and protect the lives and property of road users.

[0282] Improve public satisfaction: Provide the public with an intuitive and accurate road asset display platform to facilitate the public to understand the status and distribution of road facilities. For example, the public can query nearby signs, bus stops and other information through the platform to improve travel convenience and satisfaction.

[0283] Promoting the construction of smart cities: As an innovative technology in the field of road engineering, this invention provides strong support for the construction of smart cities. Through integration and coordination with other smart systems in the city, it can realize the intelligent management and operation of urban infrastructure and improve the overall intelligence level of the city.

[0284] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.

Claims

1. A method for road asset inventory based on three-dimensional Gaussian splashing, characterized in that: include: Collect multi-source data of road assets; wherein the multi-source data includes: image data, IMU data, and LiDAR data; Preprocessing the multi-source data to obtain a mask grayscale mask image, camera intrinsic parameters, camera pose, and sparse point cloud; Use the preprocessed data to train the 3D Gaussian splash model; including: Initialize the sparse point cloud as a 3D Gaussian distribution, project it onto the image plane for rendering, calculate the loss based on the grayscale mask map, and use the gradient descent method to optimize the Gaussian distribution parameters. Iterate the optimization until the loss converges or the maximum number of iterations is reached, and finally obtain the trained model. Based on the trained 3D Gaussian splash model, road asset inventory is conducted, including: road asset extraction, asset status analysis, GPS matching, and asset size measurement; Road asset extraction includes: Based on the trained three-dimensional Gaussian splash model, the encoder maps the Gaussian point features into high-dimensional feature vectors, and uses the grayscale mask image and neural network discriminator to determine the Gaussian kernel category label. The same type of Gaussian kernels are spatially clustered to obtain independent asset models and number them. Performing asset status analysis includes: The Gaussian cluster of a single asset in the trained 3D Gaussian splash model is projected to multiple camera viewpoints, the rendered images are classified, and the asset status is determined by combining a voting mechanism; GPS matching includes: Convert the coordinates of the center point of each Gaussian kernel in the trained three-dimensional Gaussian splash model to GPS coordinates, and calculate the average GPS coordinates of the asset, i.e., the real-world coordinates of the asset; Taking asset size measurements includes: The length, width and height information of the asset are obtained by fitting the minimum circumscribed cube of the world coordinate distribution of each asset in the trained three-dimensional Gaussian splash model.

2. The method for road asset inventory based on three-dimensional Gaussian splashing according to claim 1 is characterized in that: Preprocessing the multi-source data includes: Performing semantic segmentation on the image data; Convert IMU data to camera pose; Match the lidar data and image data to obtain a sparse point cloud with color and location information.

3. The method for road asset inventory based on three-dimensional Gaussian splashing according to claim 2 is characterized in that: Performing semantic segmentation on the image data includes: The semantic segmentation neural network model based on the Transformer architecture is used to perform semantic segmentation on the image data. The semantic segmentation neural network model based on the Transformer architecture uses historical image data with manual annotation and data enhancement for model training. The historical image data includes a variety of road scenes and weather conditions. The result of the model segmentation is saved as a grayscale mask image, and a mask grayscale mask image is generated, the size of which is the same as that of the grayscale mask image; wherein the mask grayscale mask image only includes: vehicles, pedestrians, and the sky.

4. The method for road asset inventory based on three-dimensional Gaussian splashing according to claim 2 is characterized in that: Converting IMU data to camera pose involves: Based on the IMU data, the attitude information is estimated by integrating the angular velocity and the position information is estimated by integrating the acceleration.

5. The method for road asset inventory based on three-dimensional Gaussian splashing according to claim 2 is characterized in that: Matching lidar data with image data includes: Through external parameter calibration, the point cloud in the lidar data is converted to the camera coordinate system; Use the camera intrinsic parameter matrix to project the point cloud in the camera coordinate system to the image plane; Convert the projected plane coordinates into normalized image coordinates and filter out the points projected outside the image; The image key points are extracted from the image using the SIFT feature extraction algorithm; the image key points are pixels or areas with significant features in the image; For each image keypoint, search for the nearest point in the projected point cloud; Based on these nearest points, a sparse point cloud is formed.

6. The method for road asset inventory based on three-dimensional Gaussian splashing according to claim 1, characterized in that: According to the mask grayscale mask image, the loss is calculated including: According to the grayscale mask image, the spatial area is ignored and only the static area is trained. For each area of ​​the image, a weight w is introduced alpha , for the weight w of the dynamic area alpha =0, weight w for static area alpha =1; dynamic areas include vehicles, pedestrians, and the sky, and static areas include road surfaces, road assets, buildings, and greenery; For w alpha = 0, update C before calculating the loss gt : Loss calculation: Calculate the loss between the rendered image and the real image: Using L1 loss function: 1 Among them, C render is the rendered image, C gt is a real image; Using SSIM Loss: SSIM Loss is defined as: The L1 Loss and SSIM Loss are weighted together to form the final loss: Where: L is the final loss, λ dssim is the weight of SSIM Loss, and its value is between [0,1].

7. The method for road asset inventory based on three-dimensional Gaussian splashing according to claim 1, characterized in that: The preprocessed data is used to train the three-dimensional Gaussian splash model, including: (1) Initialize each point in the sparse point cloud as a 3D Gaussian distribution; (2) For each camera pose, project the Gaussian distribution onto the image plane; (3) Render the image using the projected Gaussian distribution; (4) Based on the rendered image, calculate the loss according to the mask grayscale mask map; (5) Based on the calculated loss, use the gradient descent method to optimize the parameters of the Gaussian distribution; (6) Repeat steps (2) to (5) until the loss converges or the maximum number of iterations is reached to obtain the trained model.

8. The method for road asset inventory based on three-dimensional Gaussian splashing according to claim 1, characterized in that: Road asset extraction includes: Based on the trained 3D Gaussian splash model, an encoder is trained to map the features of Gaussian points into high-dimensional feature vectors. For each Gaussian kernel, project the Gaussian kernel onto the corresponding grayscale mask image through the camera pose to obtain the category label of the Gaussian kernel; if the Gaussian kernel is projected to different category labels under multiple viewing angles, a voting mechanism is used to determine its final category; a 32-bit category feature vector is assigned to each Gaussian kernel; Use the grayscale mask image as a supervisory signal and train the category features through the optimization algorithm; Design a neural network with 32-bit category features as input and category labels as output; use labeled Gaussian kernel data to train the discriminator, the loss function is cross entropy loss, and the optimization goal is to minimize the category prediction error; the Gaussian kernel data includes: category features and corresponding category labels; The 32-bit category feature of each Gaussian kernel is input into the trained discriminator to obtain the category label of each Gaussian kernel; the Gaussian kernel of the specified category is filtered out according to the category label to generate the segmented 3D scene; According to the segmented 3D scene, the Gaussian kernel of the extracted road assets is spatially clustered to obtain a 3D Gaussian splash model of independent assets, and each asset is numbered; the extracted road assets include signs, gantries, and street lights; Performing asset status analysis includes: The Gaussian cluster of a single asset in the trained three-dimensional Gaussian splash model is projected to multiple camera perspectives, and the rendered image is classified using a classification model to obtain multiple classification results. The final asset status analysis result is determined based on a voting mechanism, and the asset number and asset status are recorded. The classification model adopts a fine-tuned convolutional neural network classification model.

9. A road asset inventory system based on three-dimensional Gaussian splashing, characterized in that: Used to implement the method according to any one of claims 1 to 7, the system comprises: a data acquisition module, a data preprocessing module, a three-dimensional Gaussian splash model training module, and a three-dimensional Gaussian splash model analysis module; The data acquisition module is used to collect multi-source data of road assets; The data preprocessing module is used to preprocess the multi-source data and obtain a mask grayscale mask image, camera intrinsic parameters, camera pose, and sparse point cloud; The three-dimensional Gaussian splash model training module is used to perform three-dimensional Gaussian splash model training using preprocessed data; it includes: Initialize the sparse point cloud as a 3D Gaussian distribution, project it onto the image plane for rendering, calculate the loss based on the grayscale mask map, and use the gradient descent method to optimize the Gaussian distribution parameters. Iterate the optimization until the loss converges or the maximum number of iterations is reached, and finally obtain the trained model. The three-dimensional Gaussian splash model analysis module is used to perform a road asset inventory based on the trained three-dimensional Gaussian splash model; including: Carrying out road asset inventory includes: road asset extraction, asset status analysis, GPS matching, and asset size measurement; Road asset extraction includes: Based on the trained three-dimensional Gaussian splash model, the encoder maps the Gaussian point features into high-dimensional feature vectors, and uses the grayscale mask image and neural network discriminator to determine the Gaussian kernel category label. The same type of Gaussian kernels are spatially clustered to obtain independent asset models and number them. Performing asset status analysis includes: The Gaussian cluster of a single asset in the trained 3D Gaussian splash model is projected to multiple camera viewpoints, the rendered images are classified, and the asset status is determined by combining a voting mechanism; GPS matching includes: Convert the coordinates of the center point of each Gaussian kernel in the trained three-dimensional Gaussian splash model to GPS coordinates, and calculate the average GPS coordinates of the asset, i.e., the real-world coordinates of the asset; Taking asset size measurements includes: The length, width and height information of the asset are obtained by fitting the minimum circumscribed cube of the world coordinate distribution of each asset in the trained three-dimensional Gaussian splash model.

10. The system for road asset inventory based on three-dimensional Gaussian splashing according to claim 9, characterized in that: The data acquisition module includes: a 360 panoramic camera, a professional-grade real-time motion measurement device with an inertial measurement unit, and a lidar sensor; each sensor is assembled and installed on the roof through a rigid structure; an industrial computer is installed in the car to be responsible for powering the equipment and controlling the data acquisition work. The industrial computer is equipped with GPU computing power and a 5G module to upload the collected data to the cloud server.

Citation Information

Patent Citations

  • One-map road asset detection method based on image and laser point cloud

    CN118967959A

  • Dynamic real-time rendering method for large assembly scene based on three-dimensional Gaussian splashing

    CN119229031A

  • Multi-modal three-dimensional instance segmentation method based on three-dimensional Gaussian splashing

    CN119296104A

  • Grid equipment three-dimensional reconstruction method based on Gaussian splashing

    CN119359955A

  • Road asset automatic identification and modeling system based on laser point cloud

    CN119538384A

Cited By

  • Internet of Things visual monitoring method and system based on three-dimensional Gaussian splashing and computer equipment

    CN121120935A