3DGS-based autonomous driving simulation scene construction system and method
By automatically constructing highly realistic 3D scenes using the 3DGS system, the problems of discrepancies between virtual scenes and real traffic environments and low efficiency in road test scene construction in existing technologies have been solved, achieving efficient and diverse simulation scene construction and cross-platform compatibility.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA AUTOMOTIVE ENG RES INST
- Filing Date
- 2026-01-20
- Publication Date
- 2026-04-28
AI Technical Summary
In existing autonomous driving simulation tests, preset virtual scenarios are difficult to simulate the complex interactive behaviors of real traffic environments. Scenario construction based on road test data is inefficient and cannot be iterated quickly, failing to meet the need for rapid construction of highly realistic simulation models.
An autonomous driving simulation scene construction system based on 3DGS is adopted. Through multi-source data fusion, scene mining, 3D scene generation, visualization editing and scene generalization modules, it automatically constructs a highly realistic 3D scene model and supports multi-format storage and organization.
It improves the automation and diversity of scene construction, enhances the coverage of the scene library, achieves high-fidelity 3D scene reconstruction and lighting effects, and supports multi-platform compatibility and resource reuse.
Smart Images

Figure CN121564239B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of automatic driving simulation test, in particular to an automatic driving simulation scene construction system and method based on 3DGS. BACKGROUND
[0002] In the development of automatic driving, simulation test is a key means to verify the safety and reliability of the system, which can efficiently and at low cost cover multiple working conditions and support system iteration. Its value is not only dependent on massive and diversified scene data, but also needs high-simulation three-dimensional scene models as a foundation. Currently, the construction of scene data and models supporting simulation test mainly comes from two paths, which together constitute the basic data system and model system of automatic driving simulation.
[0003] One of them is the construction mode of preset virtual scene, which defines the core elements such as road layout, obstacle distribution, and traffic signals by designers, and builds models based on these elements. This kind of scene has a significant efficiency advantage in testing the traffic rule compliance of the system, and can verify the system's performance in standard working conditions. However, due to the limitations of preset logic, on the one hand, it is difficult to integrate dynamic elements that change randomly in real traffic environments, resulting in a deviation between test scenes and actual road conditions, and failing to fully simulate interactive behaviors in complex traffic flows. On the other hand, the construction of three-dimensional scene models highly depends on personnel design and debugging, which requires a lot of time from element definition to model rendering, not only slow in generation, but also limited in detail restoration of manually built models, making it difficult to achieve high simulation matching with real roads.
[0004] The other type is a scene extraction and modeling mode based on vehicle road test logs, which extracts scene information from real road test data to build three-dimensional scene models. This kind of scene completely retains the traffic characteristics of real roads, and can accurately reflect various interactive details in complex road conditions, providing an actual environment basis for simulation test. Moreover, models based on real data have an advantage in scene authenticity. However, its scene content is relatively fixed and difficult to adjust and generate flexibly according to test requirements, and the modeling efficiency is low, which cannot meet the rapid iteration test requirements. It is also impossible to actively explore unknown "corner cases", which are the key hidden dangers that cause system safety risks.
[0005] Research shows that simulation test based on real road data is more likely to capture unknown scenes, but relying solely on road test log modeling cannot break through the data limitations to achieve rapid scene expansion, nor can it solve the problem of high-simulation model construction efficiency. This double contradiction makes simulation test fall into a dilemma, that is, relying on the preset mode cannot cover complex working conditions, and the model is slow and the simulation degree is low; relying on the road test mode is difficult to dig extreme risks, and the modeling efficiency cannot match the test requirements. SUMMARY
[0006] The application aims to provide a 3DGS-based automatic driving simulation scene construction system and method to solve the problem that the existing simulation test scene cannot meet the requirement of quickly constructing a high-imitation three-dimensional model.
[0007] To achieve the above-mentioned purpose, the application adopts the following technical scheme: a 3DGS-based automatic driving simulation scene construction system, comprising a multi-source data fusion module for time synchronization, filtering and registration of sensor data collected by a real vehicle, high-precision map data and other auxiliary data, and fusion into consistent environment perception information;
[0008] A scene mining module is used to automatically extract scene elements from the environment perception information; the scene elements include vehicle trajectories, traffic participant behaviors and triggering events; typical or extreme driving scenes are identified according to the scene elements and converted into simulation scene descriptions;
[0009] A 3D scene generation module is used to reconstruct a three-dimensional scene by jointly constructing the geometric shape and illumination characteristics of the scene based on the collected multi-view images and sensor data through a three-dimensional Gaussian point carrying a large number of parameters, and quickly generate a three-dimensional scene model;
[0010] A visual editing module is used to adjust and enrich the generated three-dimensional model through an interactive editing mode;
[0011] A scene generalization module is used to transform and parameterize a single scene to automatically generate scene variants;
[0012] A multi-format scene library construction module is used to store and organize the processed scenes into a scene library in multiple formats.
[0013] Meanwhile, the application also provides a 3DGS-based automatic driving simulation scene construction method applied to the above-mentioned 3DGS-based automatic driving simulation scene construction system, comprising the following steps:
[0014] S1, collecting original environment data through a plurality of sensors and high-precision map devices mounted on a test vehicle, and generating dense environment point clouds and image data and accurate self-vehicle positioning information after preprocessing;
[0015] S2, automatically searching and identifying key scenes through a scene mining module, extracting scene elements according to the required test target, and generating a structured scene description according to the scene requirement;
[0016] S3, initializing a three-dimensional Gaussian point array carrying a large number of parameters to represent static structures in the scene, and adjusting the parameters of the Gaussian points through back propagation optimization to render a three-dimensional scene model containing geometric and illumination information;
[0017] S4, according to the scene requirements, the generated scene model is edited to form a complete test environment;
[0018] S5, parameterizing a single scene, generating multiple scene variants by combining different scene element parameters; and saving the processed scene in multiple formats to the scene library.
[0019] The principle and advantages of the present scheme are:
[0020] First, the present scheme is based on multi-source data, replacing the traditional mode of relying on single data or manual construction, greatly reducing the need for manual intervention and significantly improving the automation level of scene construction; the fusion characteristics of multi-source data break the limitations of single data sources, effectively enriching the types and forms of scenes and improving the diversity of the scene library.
[0021] Second, the technical path of combining three-dimensional reconstruction and visual editing is adopted, the three-dimensional reconstruction technology ensures the accurate restoration of the geometric structure of the scene to the characteristics of the real environment, and the visual editing can optimize the details such as lighting and material, and the synergistic effect of the two makes the generated scene have higher fidelity in geometric shape and lighting effect, and is closer to the real road and traffic environment.
[0022] Third, the scene generalization function built in the present scheme can derive more new scene types based on existing data and scene templates, especially covering complex or extreme working conditions that traditional methods cannot reach, effectively breaking through the limitations of fixed scenes and further enhancing the coverage ability of the scene library, providing more comprehensive support for simulation testing.
[0023] Fourth, the multi-format scene library constructed by the present scheme supports the format requirements of multiple mainstream simulation platforms, solving the problem of incompatibility of scene data between different platforms, not only ensuring the good compatibility of the scene between different simulation platforms, but also realizing the efficient reuse of scene resources, reducing the cost of scene development and migration. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 is a structural schematic diagram of the automatic driving simulation scene construction system based on 3DGS of the present application;
[0025] Figure 2 is a flowchart of the automatic driving simulation scene construction method based on 3DGS of the present application. DETAILED DESCRIPTION
[0026] The following will be further described in detail through specific embodiments:
[0027] The 3DGS-based automatic driving simulation scene construction system and method in the embodiment can fuse multi-source information from real vehicle collection, maps and simulation data, automatically mine safety critical scenes in road test logs, and quickly generate high-fidelity virtual scenes by using advanced three-dimensional reconstruction and visualization editing technology, while parameterizing and generalizing a single scene to form diversified scenes.
[0028] Scheme one
[0029] A 3DGS-based automatic driving simulation scene construction system is provided, as shown in the accompanying drawings, Figure 1 as shown in the accompanying drawings, comprising,
[0030] The multi-source data fusion module is configured to perform time synchronization, filtering and registration on sensor data collected by a real vehicle, high-precision map data and other auxiliary data, and fuse the data into consistent environment perception information.
[0031] In the embodiment, the system first collects original environment data by various sensors (such as lidar, camera, millimeter wave radar, GPS / IMU, etc.) mounted on a test vehicle and high-precision map equipment. The multi-source data fusion module is responsible for collecting and preprocessing multi-source data from real vehicle sensors and maps, wherein the sensor data includes, for example, lidar point cloud, camera image, millimeter wave radar, GPS / IMU, etc., and other auxiliary data includes, for example, map data, etc. By performing time synchronization and fusion on the multi-source data, including calibrating the external parameters of each sensor, performing coordinate conversion and data filtering, etc., the close registration and consistency of each data source are ensured.
[0032] In the embodiment, the specific data fusion implementation process includes:
[0033] 1. Data preprocessing stage: The original data of each sensor is preprocessed, including laser radar point cloud denoising (using a combination of voxel filtering and statistical filtering to remove outliers), camera image distortion correction (based on Zhang's calibration method), millimeter wave radar target clustering (using the DBSCAN algorithm), and GPS / IMU data smoothing (Kalman filter preprocessing). For different sensor characteristics, an adaptive filtering strategy is adopted, such as using intensity threshold filtering for laser radar to filter ground points, and using Retinex algorithm for camera to enhance backlit scene images.
[0034] 2. Time synchronization phase: sensor time synchronization is achieved through hardware triggering (IEEE 1588 PTP protocol), and the timestamp accuracy is controlled within ±0.1 ms; spatial synchronization uses a coordinate conversion matrix to realize the coordination of multiple sensor coordinate systems, and the sensor external parameter calibration is completed through a combination of offline calibration and online calibration: in the offline stage, the initial external parameters are obtained by using a camera calibration method based on a checkerboard and a laser radar-camera joint calibration algorithm; in the online stage, the external parameter drift is optimized in real time through SLAM, and the ICP algorithm is used to realize the inter-frame matching of the laser radar and the high-precision map.
[0035] 3. Multi-modal fusion algorithm: a hierarchical fusion architecture is adopted, the original data splicing is realized through timestamp alignment in the data layer; the cross-modal feature is extracted by using a feature fusion network based on an attention mechanism (such as FusionNet) in the feature layer; the improved D-S evidence theory is used for processing the target detection results of heterogeneous sensors in the decision layer fusion. For dynamic targets, a multi-sensor tracking association algorithm (such as JPDAF) is used, and for static environment, the map fusion update is realized by using Bayesian estimation.
[0036] 4. Quality verification and optimization: a multi-dimensional fusion quality evaluation index system is established, including point cloud density uniformity (coefficient of variation <5%), target detection accuracy (mAP >95%) and positioning accuracy (RTK-GPS comparison error <0.5m). Through closed-loop detection and loop optimization (BA algorithm based on graph optimization), the cumulative error is eliminated, and finally dense environment point cloud and image data, as well as accurate self-positioning information are generated, ensuring the close registration and consistency of each data source.
[0037] Scene mining module: used for automatically extracting scene elements from environment perception information; scene elements include vehicle trajectory, traffic participant behavior and triggering event; typical or extreme driving scenes are identified according to scene elements, and are converted into simulation scene description.
[0038] In this embodiment, the scene mining module uses Log2World technology (log scene mining). After obtaining the processed real driving data, key scenes are automatically searched and identified from a large amount of real road log data after fusion. Key features are extracted for scene mining requirements, including vehicle motion features (speed, acceleration, yaw rate, etc.), traffic participant interaction features (relative distance, relative speed, TTC / THW safety indicators, etc.), and environment state features (weather code, light intensity, road type, etc.), scene elements are extracted, typical or extreme driving scenes are identified, and the real log data of the vehicle is converted into simulation scene description. The simulation scene description includes road geometry, vehicle state, obstacle configuration and other information.
[0039] To ensure the accuracy of the excavation, in this embodiment, the module is built-in quantifiable criteria, namely the scene mining module includes a typical scene defining unit and an extreme scene defining unit.
[0040] Among them, the typical scene defining unit identifies statistically representative common driving conditions such as "urban congestion following" and "highway stable cruise" by clustering analysis (such as K-means algorithm) on massive data according to traffic flow density, average speed, road type and other parameters.
[0041] Specifically, the typical scene defining is realized by multi-dimensional feature analysis and clustering algorithm. By constructing dynamic feature set (such as following distance standard deviation, lane change frequency) and static feature set (such as road curvature, speed limit information), the key features are screened through mutual information entropy (Top 15 features, cumulative contribution > 90%), the improved K-means++ algorithm is used, and the optimal clustering number (usually 8-12 classes) is determined by silhouette score and Davies-Bouldin index, combined with hierarchical clustering (Agglomerative Clustering) for subclass division.
[0042] The clustering results are manually annotated and semantically mapped to form a typical scene library such as "urban congestion following" (features: speed < 20km / h, following distance < 50m, traffic flow density > 30 vehicles / km) and "highway stable cruise" (features: speed 90-110km / h, lane keeping rate > 95%, no emergency acceleration and deceleration). And establish scene feature vector index (dimension 20).
[0043] The extreme scene defining unit determines by setting specific physical quantity thresholds. For example, when the time-to-collision (TTC) is less than 1.5 seconds, the post-encroachment time (PET) is less than 1 second, or the vehicle lateral / longitudinal acceleration / deceleration and impact exceed the preset safety threshold, the system automatically captures the scene segment and classifies it as an extreme or dangerous scene.
[0044] Specifically, the extreme scenario definition unit realizes accurate identification through a multi-parameter fusion threshold system and a dynamic event triggering mechanism. For example, based on industry safety standards (such as ISO 21448) and accident data analysis, a three-level threshold model is established: a first warning threshold (TTC = 2.5s, PET = 1.5s), a second danger threshold (TTC = 1.5s, PET = 1.0s), and a third emergency threshold (TTC < 1.0s, longitudinal deceleration > 8m / s²). The lateral parameters include steering wheel rotation rate > 120° / s, lateral acceleration > 0.8g. Fusion of laser radar point cloud (target distance accuracy ± 0.1m), millimeter wave radar (speed measurement error < 0.5km / h) and visual recognition results (target classification accuracy > 98%), using D-S evidence theory for data credibility weighting (laser radar weight 0.6, vision 0.3, millimeter wave 0.1).
[0045] After triggering the threshold condition, the complete data segment from 5 seconds before the event to 3 seconds after the event is automatically intercepted, including the state of the vehicle (position, speed, acceleration), the trajectory of the traffic participants (sampling frequency 100Hz) and the environmental parameters (weather, light intensity), and the scene segment is extracted. According to the risk level (L1-L5) and the scene rarity (based on the occurrence frequency per million kilometers), double classification is performed, such as "high-speed vehicle suddenly cuts in" (L4 level, frequency < 0.1 times / 10,000 kilometers) and "pedestrian ghost head" (L5 level, frequency < 0.01 times / 10,000 kilometers), and structured scene metadata including the triggering time, duration and impact range is generated.
[0046] 3D scene generation module: for reconstructing the geometry and lighting characteristics of the scene based on the collected multi-view images and sensor data, and quickly generating a three-dimensional scene model.
[0047] In this embodiment, the 3D scene generation module uses 3D Gaussian scene reconstruction technology (3DGS) for three-dimensional reconstruction. This module uses Gaussian point cloud to represent the geometry and lighting characteristics of static scenes, and quickly generates a three-dimensional model of the scene. To ensure high fidelity, this module uses a representation method centered on three-dimensional Gaussian "splat" points to jointly build the geometry and lighting characteristics of the scene, ultimately achieving high-fidelity three-dimensional model generation and photo-level rendering effects. Each three-dimensional Gaussian point is a basic unit that carries scene information, and its parameter system completely covers the two core dimensions of geometry and lighting, accurately restoring scene details through the coordinated action of parameters.
[0048] And the traditional 3D scene generation mostly adopts the way of "Mesh + Texture", which needs to construct the polygonal mesh structure of the object first, and then supplement the appearance details through pasting texture maps, and the lighting effect depends on the complex lighting model calculation in the rendering engine. This scheme completely abandons the dependence of the mesh structure, takes the discrete Gaussian point as the basic unit, encapsulates the geometric shape and lighting characteristics directly into the parameters of each Gaussian point, and realizes more detailed geometric fitting and more realistic lighting simulation without additional maps and complex lighting models, while having the advantage of fast scene generation.
[0049] In this embodiment, the parameters of the three-dimensional Gaussian point include geometric parameters and lighting characteristics, that is, each Gaussian point contains a set of parameters to jointly represent the geometry and lighting characteristics of the scene, and the two types of parameters jointly constitute the complete characteristics of the Gaussian point, and the core parameter relationship can be expressed as
[0050] ; (1)
[0051] In the formula, G is a single complete three-dimensional Gaussian point, which is a basic information unit for constructing a scene, integrating all geometric and lighting characteristics; is a set of geometric feature parameters of the Gaussian point, which determines the shape, position and size of the Gaussian point in space; is a set of lighting feature parameters of the Gaussian point, which determines the color, transparency and lighting response characteristics of the Gaussian point.
[0052] The geometric shape of the scene, i.e. the geometric shape of the Gaussian point, is determined by the spatial position (x, y, z) of the three-dimensional Gaussian point and the 3x3 covariance matrix. The position parameter determines the center of the point in space, and the covariance matrix defines the shape, size and orientation of the three-dimensional ellipsoid, which defines the spatial state of the three-dimensional ellipsoid through the cooperation of the two, and a large number of small ellipsoids jointly fit the fine geometric surface of complex objects such as roads, buildings and vegetation, and then fit the surface of the scene object, which can be expressed as
[0053] , ; (2)
[0054] In the formula, is the spatial position coordinate (x, y, z), which defines the center position of the ellipsoid corresponding to the Gaussian point in space; is a 3x3 covariance matrix, specifically , wherein is a rotation matrix, is a scale matrix, the elements on the diagonal correspond to the scaling of the three principal axes of the Gaussian ellipsoid, which is used to define the shape (such as spherical, ellipsoidal), size (volume scale) and spatial orientation of the ellipsoid, and a large number of ellipsoids are used to fit the fine surface of the scene object.
[0055] Based on the geometric parameters, the ellipsoid distribution formed by the Gaussian points in the three-dimensional space can be described by a Gaussian function, and the probability density distribution, i.e. the distribution of the Gaussian points in the three-dimensional space, can be represented as
[0056] ; (3)
[0057] In the formula, is the coordinate of any point in the three-dimensional space, used to describe the distribution range of the Gaussian point ellipsoid in the space.
[0058] The lighting characteristics of the Gaussian points are represented by color (R, G, B), opacity (Alpha) and spherical harmonic function (Spherical Harmonics) coefficients. Among them, the color and opacity directly determine the basic appearance and transparency of the point, and the spherical harmonic function coefficient is used to simulate the view-dependent lighting effect. Using the spherical harmonic function coefficient, each point can simulate the view-dependent lighting effect, such as material reflection, ambient light shading and shadow, so that photo-level rendering effect can be realized without complex lighting model in traditional rendering engine, and the lighting physical characteristics of real world are highly restored. Then the whole lighting parameter set and rendering effect relationship can be represented as
[0059] ; (4)
[0060] In the formula, is the color value (R, G, B) of the lighting parameter, which defines the basic color attribute of the Gaussian point; is the opacity parameter, with a value range of [0, 1], which determines the degree of transparency of the Gaussian point (0 is completely transparent, and 1 is completely opaque); is the spherical harmonic function coefficient, which is used to encode lighting information so that the Gaussian point can simulate the view-dependent lighting effect (such as reflection and shadow); is the observation angle parameter, reflecting the spatial angle of the observer relative to the Gaussian point, and the spherical harmonic function coefficient can dynamically adjust the lighting output according to V; is the lighting output value of a single Gaussian point at the angle V, which is the result of the action of color, opacity and spherical harmonic function.
[0061] Massive three-dimensional Gaussian points fit the scene geometric surface through spatial distribution, and the lighting effect of each point is superimposed to realize the rendering of the whole scene. The pixel color finally presented by the scene is the superposition of the lighting contribution of all Gaussian points in the field of view. That is, the lighting effect of massive three-dimensional Gaussian points is superimposed to realize the rendering of the whole scene, and the scene model is represented as
[0062] ; (5)
[0063] In the formula, The color value finally presented by a target pixel in the scene under a view angle V is a superposition of all relevant Gaussian point contributions; N is the total number of three-dimensional Gaussian points that contribute to the target pixel in the field of view, and is used to calculate the detail richness of the volumetric scene; The position coordinates of the target pixel in the three-dimensional space are used to calculate the contribution weight of each Gaussian point to the pixel.
[0064] Through the above generation method based on three-dimensional Gaussian points, in terms of geometric representation, an ellipsoid defined by a covariance matrix is used to replace the traditional grid, which can more finely fit the surfaces of complex objects such as roads and buildings; in terms of light simulation, the dependence on complex rendering engines is eliminated by means of spherical harmonic function coefficients, and the real light characteristics such as material reflection and environmental light shading are directly restored, so as to finally achieve the goal of quickly generating a high-fidelity 3D scene.
[0065] Meanwhile, in the reconstruction of the three-dimensional scene model, a loss function between the rendering output value and the original image pixel value is minimized , and the parameters of the three-dimensional Gaussian points are iteratively optimized by using a back propagation algorithm to update the set of three-dimensional Gaussian points. The loss function is represented as:
[0066] ; (6)
[0067] In the formula, is the mean absolute error between the rendered image and the original image; is a preset weight coefficient; is a structural similarity loss, which is used to constrain the geometric integrity of the static structure in the reconstructed scene and the consistency of the view angle and light.
[0068] The gradient is calculated and the set of Gaussian points is updated in the negative direction of the gradient , wherein includes the spatial position , the covariance matrix , the opacity parameter , and the spherical harmonic function coefficient .
[0069] A visual editing module is used to adjust and enrich the generated three-dimensional model through an interactive editing mode.
[0070] In the embodiment, the visual editing module provides an interactive editing interface and an automatic tool to adjust and enrich the generated three-dimensional scene. Users or algorithms can add or modify static elements (road signs, traffic signals, buildings, vegetation, etc.) and dynamic elements (traffic flow, pedestrians, etc.) in the scene, as well as configure environmental parameters such as light and weather, to improve the visual effect and diversity of the scene.
[0071] Scene generalization module: used to transform and parameterize a single scene, and automatically generate scene variants.
[0072] In this embodiment, parameterization operations include adjusting road layout, changing traffic density, randomly moving obstacles, and modifying traffic rules to obtain a wider range of scene distributions and higher test coverage.
[0073] Multi-format scene library building module: Used to store and organize the processed scenes into a scene library in multiple formats.
[0074] In this embodiment, the processed scenes are stored in multiple formats and organized into a scene library, supporting cross-platform use. The scene library includes scene data (such as 3DGS format, WorldSim format, OpenDRIVE format, etc.) and auxiliary resources (such as map data) for different simulation platforms. By combining various scene elements, almost unlimited scenes can be constructed, providing a rich variety of scene inputs for downstream simulation testing.
[0075] Option 2
[0076] A method for constructing autonomous driving simulation scenes based on 3DGS is provided, which is applied to the aforementioned 3DGS-based autonomous driving simulation scene construction system, as shown in the attached figure. Figure 2 As shown, it includes the following steps:
[0077] S1 collects raw environmental data through various sensors and high-precision map equipment installed on the test vehicle. After preprocessing, it generates dense environmental point clouds and image data, as well as accurate vehicle positioning information.
[0078] In this embodiment, the system first collects raw environmental data using various sensors (such as LiDAR, cameras, millimeter-wave radar, GPS / IMU, etc.) and high-precision map equipment mounted on the test vehicle. Then, the raw environmental data is preprocessed, including synchronizing and fusing the multi-sensor data according to timestamps, removing noise and redundant information, and fusing them into consistent environmental perception information.
[0079] In addition, this step may also include technical processing such as calibrating the extrinsic parameters of each sensor, performing coordinate transformation and data filtering, to ensure close registration and consistency of each data source.
[0080] S2 automatically searches for and identifies key scenarios through the scenario mining module, extracts scenario elements according to the required test objectives, and generates structured scenario descriptions based on scenario requirements.
[0081] In this embodiment, after obtaining processed real-world driving data, the scene mining module automatically searches for and identifies key scenes from a large amount of road logs. These scenes can be moments when the autonomous driving system takes over, collision warnings occur, or other safety-critical events take place. The module extracts scene parameters based on the required test objectives (such as traffic conditions, vehicle interaction patterns, etc.). For specific (typical and extreme) scene requirements, a structured scene description file (such as OpenSCENARIO format) is generated to meet simulation needs. In this embodiment, the scene description file records information including the static environment, dynamic objects, triggering conditions and events, and environmental parameters.
[0082] The key scenario detection and classification employs a dual-engine detection mechanism of "rules + learning." The rule engine filters potential risk scenarios based on preset safety thresholds (such as TTC < 2s, collision warning activation, system takeover request, etc.); the learning engine automatically identifies implicit dangerous scenarios without explicit triggering conditions (such as pedestrians suddenly appearing from behind obstacles, slow traffic in construction areas, etc.) by training a Transformer-based event classification model (input is a 10s time-series feature sequence, output is scenario category probability). Scenario classification uses a multi-level labeling system, with the first-level classification including 6 major categories such as "vehicle interaction," "vulnerable road users," and "road anomalies," and the second-level classification further refined into 32 typical scenario types.
[0083] For the detected key scenarios, standardized scenario parameters are generated. The static environment uses high-precision map data (such as OpenDRIVE files), automatically parsing the road topology, defining road geometry, lane lines, traffic signs, etc., and extracting geometric parameters such as the number of lanes, radius of curvature, and slope. The dynamic objects accurately describe the initial state (position, velocity, acceleration) and time-varying complete trajectory of all traffic participants (vehicles, pedestrians, etc.) in the scenario. Kalman filtering is used to smooth the traffic participant trajectories, and the RANSAC algorithm is used to remove trajectory noise, generating spatiotemporal sequence data (sampling frequency 10Hz) containing position (WGS84 coordinate system), velocity, and acceleration. Triggering conditions and events explicitly record key indicators for identifying the scenario (such as a cut-in event with TTC=1.2s) and define the start and end times of the scenario. Environmental parameters record information such as weather and lighting (e.g., noon, dusk) to provide a basis for high-fidelity reconstruction.
[0084] The scenario compiler converts parameterized data into multi-format scenario files. The core outputs include: OpenSCENARIO format (for dynamic behavior description), OpenDRIVE format (for road network definition), and JSON format metadata (containing scenario ID, risk level, data source, etc.).
[0085] Using the above methods, real driving situations can be accurately replayed in a closed-loop environment.
[0086] S3 initializes a massive number of 3D Gaussian point lattices carrying parameters to represent the static structure in the scene, and then optimizes and adjusts the parameters of the Gaussian points through backpropagation to render a 3D scene model containing geometric and lighting information.
[0087] In this embodiment, for complex scenes or environments requiring high realism, the 3DGS module uses camera and radar data for 3D reconstruction.
[0088] First, the camera pose is recovered and a sparse point cloud is generated using visual measurement technology (SfM). In this embodiment, SuperPoint is used to extract image feature points (≥2000 feature points per frame), and a FLANN matcher is used to achieve cross-frame feature association (matching accuracy >95%). The camera pose and 3D point coordinates are solved using Ceres Solver, with reprojection error controlled within 1.0 pixel. Triangulation calculations are performed on the matched feature points to generate a sparse point cloud containing ≥1 million points, with a point cloud density ≥50 points / m². 2 .
[0089] Then, based on the sparse point cloud distribution, a large number of 3D Gaussian point ellipsoids carrying position, covariance, color, opacity, and spherical harmonic coefficients are stacked to represent the static structure in the scene. An iterative optimization strategy can be used to adjust the Gaussian point parameters, and a virtual view can be generated using a differentiable renderer (such as 3DGS-Renderer). This involves projecting the 3D Gaussian ellipsoids onto a two-dimensional plane after affine transformation and calculating the photometric loss (MSE < 10) compared to the real image. -3 ) and structural loss (SSIM > 0.95).
[0090] The parameters of the Gaussian points are then continuously adjusted using optimization algorithms such as backpropagation. For example, the Adam optimizer (learning rate 1e-4) iteratively optimizes the position, covariance, color, and opacity parameters. Every 100 iterations, Gaussian points are split (when the eigenvalue ratio of the covariance matrix is >3) and pruned (removed when the opacity is <0.01), gradually improving the detail reproduction. This results in a realistic rendering of environmental details such as roads, buildings, and road surface textures. The output is a 3D scene model containing rich geometric and lighting information, which can be used for subsequent simulations and rendering.
[0091] S4 allows you to edit the generated scene model according to the scene requirements, forming a complete test environment.
[0092] In this embodiment, the visual editor (WorldSim editor) allows users to edit or add / remove static elements, dynamic elements, and environmental parameters of the generated scene model to enhance visual effects and diversity. Users can add static objects to the scene, set lighting and weather conditions, and arrange dynamic traffic flow.
[0093] Static elements include road signs, traffic signals, buildings, and vegetation. Dynamic elements include traffic flow and pedestrians. Environmental parameters include lighting and weather (such as sunshine, rain, snow, daytime, and nighttime).
[0094] In this embodiment, additional vehicles or pedestrian agents can be inserted according to testing requirements. This enriches the original reconstructed scene into a complete test environment usable in simulation. The editing process supports visual interaction and can automatically generate scene variations that comply with traffic rules using a rule engine, ensuring the effectiveness and safety of the scene.
[0095] S5 performs parameterization on a single scene, generating multiple scene variants by combining different scene element parameters; and saves the processed scenes to the scene library in multiple formats.
[0096] In this embodiment, the scenario is generalized to generate more variations. Parametric processing includes adjusting road layout, changing traffic density, randomly moving obstacles, and modifying traffic rules. The system can automatically process these variations, such as adjusting road connections, changing vehicle density, moving obstacle positions, or altering vehicle behavior patterns (e.g., acceleration, lane-changing strategies), to construct a series of "extreme" or "boundary" scenarios. This generalization further expands the scenario library's coverage, enhancing the robustness testing of autonomous driving algorithms. By combining different scenario element parameters, a near-infinite number of scenarios can be constructed, including more extreme cases and safety boundary scenarios, improving the reliability of simulation testing.
[0097] The system then saves the scenes, after being mined, reconstructed, edited, and generalized, into scene libraries in multiple formats to support the needs of different simulation platforms. For example, the 3DGS scene library retains high-fidelity raster representations; the WorldSim scene library stores executable scene files for direct use by the simulation editor; the Log2World scene library stores scene configurations based on actual logs; and a map library composed of high-definition map data is also maintained. Based on this, the scene library provides rich resources for downstream simulation testing modules and can output scene information to different simulation engines (such as Unreal Engine, Unity, CARLA, etc.) through a unified interface, ensuring cross-platform usability.
[0098] By constructing high-fidelity scenarios, it can be applied to test the robustness of perception algorithms such as cameras and LiDAR in complex lighting conditions (e.g., sunrise / sunset glare, tunnel entrances / exits), severe weather conditions (e.g., road surface glare in rainy weather), and complex texture environments, thus validating the perception algorithms. Furthermore, by comparing the output of virtual sensors in the simulation with real-world data at the pixel level or point cloud level, sensor model parameters can be precisely calibrated and optimized to achieve sensor simulation model calibration. Additionally, in highly realistic hazardous scenarios (e.g., sudden pedestrian appearances or emergency braking by the vehicle in front), it can test the extreme reaction capabilities of planning and control algorithms, or provide drivers with an immersive visual experience for research on autonomous driving system takeover, human-machine interaction, and other issues, realizing human-machine co-driving and driver-in-the-loop simulation.
[0099] In this embodiment, a closed-loop generation process from multi-source data to high-quality simulation scenes is achieved, effectively improving the automation and diversity of scene construction. By combining 3D reconstruction and visual editing, the generated scenes have higher geometric and lighting fidelity. The scene generalization function further enhances the coverage of the scene library. The multi-format library ensures the compatibility and reusability of scenes across different simulation platforms.
[0100] The implementation of this solution also overcame many challenging limiting factors and obstacles. First, the stringent requirements of 3DGS models for high-performance GPUs and large amounts of video memory constitute a computational power challenge from the outset. However, this solution did not rely on hardware upgrades. Instead, it explored the potential of computational power within a reasonable hardware configuration through model optimization, task splitting, and innovative parallel scheduling. After hundreds of algorithm iterations, it overcame the challenge of matching computational power with application requirements.
[0101] Secondly, the quality of the raw data directly determines the fidelity of the reconstruction, but the collected data often suffers from problems such as blurriness and missing perspectives. This solution constructs a multi-dimensional data preprocessing system, introducing image enhancement, perspective completion, and sensor fusion correction technologies. Combined with pre-control of data quality and optimization of the model, it achieves a stable transformation from non-ideal data to high-fidelity scenes.
[0102] The fusion of static and dynamic elements is the core barrier, as 3DGS excels at static reconstruction but struggles with dynamic objects. This solution creatively designs a "dynamic-static collaborative rendering" mechanism. After extracting trajectories using Log2World, it constructs a real-time lighting model for dynamic objects, allowing the lighting effects to blend naturally with the static background. Multiple scene tests have been conducted to avoid fusion defects.
[0103] Furthermore, the difficulty of geometric editing in 3DGS scenes once limited the practicality of the solution. This solution addresses this by developing and designing a "Gaussian point cloud and lightweight mesh hybrid editing interface," enabling indirect geometric adjustments based on mesh interaction. This enhances editing capabilities while retaining high fidelity, and its ease of use and reliability have been ensured through multi-tool compatibility testing. These technological breakthroughs and innovations have facilitated the implementation and realization of this solution.
[0104] The above descriptions are merely embodiments of the present invention, and common knowledge such as specific technical solutions and / or characteristics are not described in detail here. It should be noted that those skilled in the art can make various modifications and improvements without departing from the technical solutions of the present invention, and these should also be considered within the scope of protection of the present invention. These modifications and improvements will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.
Claims
1. A 3DGS-based autonomous driving simulation scene construction system, characterized in that, include, The multi-source data fusion module is used to synchronize, filter, and register sensor data, high-precision map data, and other auxiliary data collected from real vehicles, fusing them into consistent environmental perception information; it uses a coordinate transformation matrix to realize the multi-sensor coordinate system unification; it adopts a hierarchical fusion architecture, and in the data layer, it realizes the stitching of raw data through timestamp alignment; the feature layer uses a feature fusion network based on an attention mechanism to extract cross-modal features. The decision-level fusion uses an improved DS evidence theory to process target detection results from heterogeneous sensors; for dynamic targets, a multi-sensor tracking and association algorithm is used, and for static environments, Bayesian estimation is used to achieve map fusion and updating. Scene mining module: used to automatically extract scene elements from environmental perception information; the scene elements include vehicle trajectory, traffic participant behavior and triggering events; Identify typical or extreme driving scenarios based on scene elements and convert them into simulation scene descriptions; The scene mining module includes a typical scene definition unit and an extreme scene definition unit. The typical scene definition unit identifies representative common driving conditions based on traffic flow density, average vehicle speed, and road type, and achieves standardized scene classification through multi-dimensional feature analysis and clustering algorithms. The extreme scene definition unit makes judgments based on set physical quantity thresholds, and achieves accurate identification through a multi-parameter fusion threshold system and a dynamic event triggering mechanism. 3D Scene Generation Module: Based on acquired multi-view images and sensor data, it adopts a dynamic-static collaborative rendering mechanism to reconstruct the geometric shape and lighting characteristics of the scene through a massive set of 3D Gaussian points carrying parameters, and quickly generate 3D scene models. The 3D scene generation module uses three-dimensional Gaussian points as its core; it utilizes a massive set of three-dimensional Gaussian points carrying parameters. The geometry and lighting characteristics of the scene are jointly constructed; the geometry of the Gaussian points is determined by their spatial location in three dimensions and a 3×3 covariance matrix; when fitted to the surface of scene objects, the distribution of the Gaussian points in three-dimensional space is represented as follows: ; In the formula, Let be the coordinates of any point in three-dimensional space; Spatial location coordinates; It is a 3×3 covariance matrix. , For rotation matrix, The scale matrix has diagonal elements corresponding to the scaling scales of the three principal axes of the Gaussian ellipsoid, used to define the shape, size, and orientation of the 3D ellipsoid; it fits the fine surface of scene objects using a large number of ellipsoids. The lighting characteristics of a Gaussian point are represented by color, opacity, and spherical harmonic function coefficients; the spherical harmonic function coefficients are used to simulate view-dependent lighting effects, and the relationship between the overall lighting parameter set and the rendering effect is as follows: ; In the formula, The color value is a parameter of the lighting. This is the opacity parameter; These are the coefficients of the spherical harmonic function; These are the parameters for the observation viewpoint; This represents the illumination output value of a single Gaussian point at the viewing angle V. Each three-dimensional Gaussian point is a basic unit that carries scene information. Each Gaussian point contains a set of parameters that together represent the geometric and lighting characteristics of the scene. The two types of parameters together constitute the complete features of the Gaussian point. A massive number of tiny three-dimensional Gaussian points fit the geometric surface of the scene through spatial distribution. The lighting effects of each point are superimposed to achieve the rendering of the overall scene. The pixel color that the scene finally presents is the superposition of the lighting contributions of all Gaussian points within the field of view. By optimizing the parameters of the three-dimensional Gaussian points through backpropagation, a three-dimensional scene model containing geometric shape and viewpoint-related lighting characteristics is reconstructed. The scene model is then represented as ; In the formula, N is the total number of three-dimensional Gaussian points within the field of view that contribute to the target pixel; These are the position coordinates of the target pixel in three-dimensional space. Visual editing module: Used to adjust and enrich the generated 3D model through interactive editing; Scene generalization module: used to perform transformation and parameterization operations on a single scene, and automatically generate scene variants; Multi-format scene library building module: Used to store and organize the processed scenes into a scene library in multiple formats.
2. The autonomous driving simulation scene construction system based on 3DGS according to claim 1, characterized in that: When reconstructing a 3D scene model, the loss function between the rendered output value and the original image pixel value is minimized. The parameters of the three-dimensional Gaussian points are iteratively optimized using the backpropagation algorithm; the loss function is expressed as: ; In the formula, The mean absolute error between the rendered image and the original image; These are preset weighting coefficients; The structural similarity loss is used to constrain the geometric integrity and viewpoint lighting consistency of static structures in the reconstructed scene. By calculating the gradient And update the set of Gaussian points along the negative gradient direction. ,in Including spatial location Covariance matrix Opacity spherical harmonic coefficients .
3. A method for constructing autonomous driving simulation scenarios based on 3DGS, characterized in that, The 3DGS-based autonomous driving simulation scene construction system applied to any one of claims 1-2 includes the following steps: S1 collects raw environmental data through various sensors and high-precision map equipment installed on the test vehicle, and generates dense environmental point cloud and image data, as well as accurate vehicle positioning information after preprocessing. S2 automatically searches for and identifies key scenarios through the scenario mining module, extracts scenario elements according to the required test objectives, and generates structured scenario descriptions based on scenario requirements. S3 initializes a massive number of 3D Gaussian point lattices carrying parameters to represent the static structure in the scene, and then optimizes and adjusts the parameters of the Gaussian points through backpropagation to render a 3D scene model containing geometric and lighting information. S4: Edit the generated scene model according to the scene requirements to form a complete test environment; S5 performs parameterization on a single scene, generating multiple scene variants by combining different scene element parameters; and saves the processed scenes to the scene library in multiple formats.
4. The method for constructing an autonomous driving simulation scene based on 3DGS according to claim 3, characterized in that: In S2, scene descriptions include static environment, dynamic objects, triggering conditions and events, and environment parameters.
5. The method for constructing an autonomous driving simulation scene based on 3DGS according to claim 3, characterized in that: In S4, the generated scene model can be edited or have its static elements, dynamic elements, and environmental parameters added or removed; the static elements include road signs, traffic signals, buildings, and vegetation; the dynamic elements include traffic flow and pedestrians; and the environmental parameters include lighting and weather.
6. The method for constructing an autonomous driving simulation scene based on 3DGS according to claim 3, characterized in that: In S5, parametric processing includes adjusting road layout, changing traffic density, randomly moving obstacles, and modifying traffic rules.
Citation Information
Patent Citations
SCANeR-based automatic driving simulation test model construction method
CN111797001A
Rapid high-fidelity reconstruction method for automatic driving scene
CN121120895A