A four-dimensional gaussian-based automatic driving multi-modal driving world construction method

CN122597680APending Publication Date: 2026-08-18JILIN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611091622.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-22
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0004]为解决生成模型缺少稳定世界状态、重建模型缺少未来生成能力以及世界变化与传感器混淆的问题,本发明提供了一种基于四维高斯的自动驾驶多模态驾驶世界构建方法,建立包含相机多视角图像、激光雷达点云、语义、实例、运动及不确定性的四维高斯显式世界状态;将世界演化与传感器生成分层处理,采用确定性传播与概率残差生成未来状态,利用物理几何渲染与传感器残差生成形成激光雷达观测;并依据可观测区域和想象区域实施差异化循环验证

Benefits of technology

[0097]This invention provides a method for constructing a multimodal driving world for autonomous driving based on four-dimensional Gaussian geometry. It establishes a four-dimensional Gaussian explicit world state comprising multi-view images from cameras, LiDAR point clouds, semantics, instances, motion, and uncertainties. The method integrates world evolution with sensor generation in a hierarchical process, employing deterministic propagation and probabilistic residuals to generate future states, and utilizing physical geometry rendering and sensor residual generation to form LiDAR observations. Differential cyclic verification is implemented based on observable and imagined regions. This invention provides a more realistic multimodal driving world for the safety verification of end-to-end autonomous vehicles, supporting their safety verification and facilitating their industrialization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122597680A_ABST
    Figure CN122597680A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of automatic driving scene generation, multi-modal sensor simulation and world model, and particularly relates to a kind of automatic driving multi-modal driving world construction method based on four-dimensional Gauss. It comprises: S1, synchronously fusing multi-modal time series observation, constructing four-dimensional Gauss world state containing geometry, appearance, semantics, instance, motion and uncertainty;S2, dividing observable and imaginary Gauss, constructing hierarchical world label;S3, generating future world state through deterministic propagation and probability residual;S4, rendering camera image based on future state, and calculating laser radar hit, intensity, missing point and residual for each beam to generate final point cloud;S5, implementing multi-dimensional cyclic constraint and consistency verification. The application establishes explicit four-dimensional Gauss world state, fuses physical rendering and residual modeling, and can support end-to-end automatic driving safety verification and industrialization landing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of autonomous driving technology, specifically a method for constructing a multimodal driving world for autonomous driving based on four-dimensional Gaussian. Background Technology

[0002] Safety verification of autonomous driving systems not only requires reproducing real-world driving scenarios, but also the ability to generate time-synchronized camera images and LiDAR point clouds after changes in the vehicle's trajectory, the movement of traffic participants, and sensor configurations. Existing point cloud diffusion or image generation methods can learn the distribution of observational data, but their latent variables are often unstable and prone to geometric drift at adjacent time points and from different perspectives, making it difficult to generate editable and evolvable world states.

[0003] 3D Gaussian splashing technology utilizes explicit spatial primitives to represent scenes, offering advantages such as fast rendering speed and good multi-view consistency. However, existing methods primarily reconstruct scenes based on existing acquired information. When vehicle trajectories deviate from the original acquisition route, dynamic objects enter occluded areas, or future traffic layouts need to be generated, pure reconstruction models lack probabilistic generation capabilities, easily leading to issues such as localized scene holes, visual inconsistencies, and repetitive input of driving scenes. Furthermore, real-world changes and sensor observations belong to different levels. If the same generation network simultaneously alters object geometry and LiDAR echoes, the model may mistake sparse point clouds for geometric holes or rain and fog features for real-world entities. Summary of the Invention

[0004] To address the issues of the lack of stable world states in generative models, the lack of future generation capabilities in reconstructed models, and the confusion between world changes and sensor data, this invention provides a method for constructing a multimodal driving world for autonomous driving based on four-dimensional Gaussian. This method establishes a four-dimensional Gaussian explicit world state that includes multi-view images from cameras, LiDAR point clouds, semantics, instances, motion, and uncertainties. World evolution and sensor generation are processed in a hierarchical manner. Deterministic propagation and probabilistic residuals are used to generate future states, and physical geometry rendering and sensor residual generation are used to form LiDAR observations. Differential cyclic verification is implemented based on observable and imagined regions. The results of this invention can provide a more realistic multimodal driving world for the safety verification of end-to-end autonomous vehicles, supporting end-to-end autonomous vehicle safety verification and facilitating its industrialization.

[0005] The technical solution of this invention is described below in conjunction with the accompanying drawings:

[0006] This invention provides a method for constructing a multimodal driving world for autonomous driving based on four-dimensional Gaussian, comprising the following steps:

[0007] S1. Synchronization of multimodal temporal observations and construction of the state of a four-dimensional Gaussian world;

[0008] Time synchronization and feature fusion are performed on multi-view images, LiDAR point clouds, vehicle pose and sensor configuration to construct a four-dimensional Gaussian world state that includes geometric, appearance, echo, semantic, instance, motion and uncertainty attributes.

[0009] S2, Observable region division and hierarchical world marker construction;

[0010] Based on the coverage of camera and lidar rays, observable Gaussians and imagined Gaussians are classified and coded as global road and map markers, dynamic object markers, local surface markers, and incremental generation markers.

[0011] S3, Deterministic Propagation and Evolution of the Probabilistic Residual World;

[0012] Based on the current world state, the vehicle's actions, and the deterministic part of the map constraints, and by generating object responses, multiple candidate trajectories, the probability residuals corresponding to the appearance of new objects and the completion of occluded areas, the future four-dimensional Gaussian world state is obtained.

[0013] S4, camera and lidar observation generation;

[0014] The camera image is generated based on the future four-dimensional Gaussian world state, and the geometric hit, echo intensity, probability of missing points and sensor residual are calculated according to the beam emission time of the lidar to form a lidar point cloud.

[0015] S5, Cyclic Constraints and Automated Driving Safety Verification;

[0016] The generated observations are subjected to observation loops, latent space loops, new viewpoint loops, and cross-modal loops, and the generated scenarios that pass the consistency check are used for perception and failure verification of the autonomous driving system under test.

[0017] Furthermore, the specific method of S1 is as follows:

[0018] S11. Establish a multimodal time series observation set;

[0019] Time The input observation is represented by equation (1):

[0020] (1)

[0021] In the formula, For a moment A multimodal observation set; For a moment No. Images from a single camera perspective; For camera viewpoint numbers; For a moment LiDAR point cloud; For a moment Pose transformation from vehicle coordinate system to world coordinate system; For a moment The set of sensor configurations; For discrete time sequence numbers;

[0022] Unified time compensation is performed on the exposure times of each camera and the beam-by-beam scanning times of the LiDAR, interpolating the dynamic object state to the corresponding sampling time; camera configuration includes intrinsic parameters, extrinsic parameters, and image size; LiDAR configuration includes beam angle, scanning frequency, mounting pose, maximum range, and intensity calibration. Synchronized image pixels and LiDAR points can all be mapped to a unified world coordinate system.

[0023] S12, Construct the state of a four-dimensional Gaussian explicit world;

[0024] No. The properties of a four-dimensional Gaussian element are represented by equation (2):

[0025] (2)

[0026] In the formula, For a moment The A four-dimensional Gaussian element; The three-dimensional center position; It is a local rotation matrix; It is a three-axis scale vector; As to the degree of space occupation; Features of the camera's appearance; It exhibits near-infrared echo response characteristics; For the first The inherent tendency parameters of Gaussian elements to generate LiDAR point loss; These are motion state parameters, including the three-dimensional velocity vector and local deformation parameters; It is a semantic category vector; For instance category vectors; Let be the world state uncertainty parameter, used to characterize the . The geometric position, shape, and reliability of the generated results of each four-dimensional Gaussian element; The Gaussian element number;

[0027] The image branch extracts color, semantic, depth, and optical flow features, while the LiDAR branch extracts precise location, intensity, scan line, and dynamic point features. Based on the point cloud location, initial Gaussian anchors are constructed by combining sparse voxels and three-plane locations, and the remaining attributes of each initial Gaussian anchor are predicted through a multimodal fusion network. During training, image or point cloud modalities are randomly discarded to ensure that the single-modal coding results and the fusion coding results fall into the same normalized world coordinates.

[0028] The complete world state is divided into a static background, dynamic objects, and an incrementally generated Gaussian set, as shown in equation (3):

[0029] (3)

[0030] In the formula, For a moment The complete four-dimensional Gaussian world state; A static set of backgrounds fixed in the world coordinate system; For dynamic object serial numbers; Total number of dynamic objects; For the first An object at time Pose transformation from the normalized object coordinate system to the world coordinate system; For the first A Gaussian set of objects established in the canonical object coordinate system; For a moment Incremental generation of Gaussian sets; This is a set union operation;

[0031] Static backgrounds include roads, buildings, road edges, and fixed facilities; dynamic object Gaussian sets are used to maintain the cross-frame shape and instance category of vehicles or pedestrians; incrementally generated Gaussian sets are used to represent newly entering objects, surfaces that are first revealed after occlusion, and the birth, death, splitting, and merging of Gaussian elements.

[0032] Furthermore, the specific method of S2 is as follows:

[0033] S21. Divide the observable area and the imagined area according to the sensor rays;

[0034] The four-dimensional Gaussian world is divided according to the observed properties as shown in equation (4):

[0035] (4)

[0036] In the formula, For a moment The complete state of the world; An observable Gaussian set that is fully covered by the rays of a real camera or lidar; For the imagined Gaussian set that is obscured, outside the field of view, or inferred only from prior knowledge; This is a set union operation; For time sequence number;

[0037] S22. Construct a hierarchical world marker;

[0038] Encode the variable number of Gaussian sets into hierarchical latent states, as shown in equation (5):

[0039] (5)

[0040] In the formula, For a moment The hierarchical world potential state; It is a Gaussian set encoder; For global road and map marking; Mark a collection for dynamic objects; A set of local surface markers; Generate a set of tags for incremental generation; For time sequence number;

[0041] Global markers describe road topology, building layout, and overall traffic structure; object markers represent the category, pose, velocity, and shape of each object through set matching; local surface markers use sparse meshes to preserve curbs, fences, and vegetation; incrementally generated marker sets serve as the entry point for generating unobserved areas and new objects, while the Gaussian decoder simultaneously reconstructs geometry, camera appearance, LiDAR response, and instance category.

[0042] Furthermore, the specific method of S3 is as follows:

[0043] S31, The deterministic part of the propagation of the state of the world;

[0044] Based on the current world potential state, vehicle actions, and the deterministic part of map propagation, as shown in equation (6):

[0045] (6)

[0046] In the formula, The potential state of the deterministic world in the next moment; For coordinate and kinematic propagation operators; For a moment The hierarchical world potential state; For a moment The vehicle's self-control actions; For map constraint set; For time sequence number;

[0047] S32. Utilize conditional flow to generate probabilistic residuals of future world states;

[0048] In the flow matching training, an interpolation path is constructed between the noisy latent state and the true future residual, as shown in Equation (7):

[0049] (7)

[0050] In the formula, For continuous flow time The interpolation latent state at the location; The stream time is a value between 0 and 1; The initial noise latent state is sampled from a standard Gaussian distribution; The target residual latent state is obtained by encoding the real next-moment world state;

[0051] The conditional flow matching loss is shown in equation (8):

[0052] (8)

[0053] In the formula, Matching loss for world residual flow; To calculate the mathematical expectation of the training samples, noise, and streaming time; For parameters A defined conditional velocity field network; This refers to the current action of the vehicle. For map constraints; The square of the second norm of a vector;

[0054] The future world potential state is obtained by combining deterministic propagation and generation residuals, as shown in equation (9):

[0055] (9)

[0056] In the formula, The complete potential state of the world in the next moment; The probabilistic residuals generated for conditional flow describe object response, multimodal trajectory evolution, new object appearance, occlusion region completion, and local non-rigid changes. For time sequence number;

[0057] By changing the initial noise of the flow model under the same current state and vehicle actions, multiple physically plausible future traffic states are generated. When the vehicle actions are changed, the static world remains unchanged, while the viewpoint, occlusion, and other participant reactions change with the actions.

[0058] Furthermore, the specific method of S4 is as follows:

[0059] S41. Render camera observations based on the future state of the four-dimensional Gaussian world.

[0060] time The The observations from each camera are represented by equation (10):

[0061] (10)

[0062] In the formula, For a moment The generated first One camera image; Rendering operators for a four-dimensional Gaussian camera; For a moment The state of the four-dimensional Gaussian world; For the first Each camera at any time pose transformation; For the first Each camera at any time Configuration parameters;

[0063] The vehicle pose and dynamic object pose are updated according to the camera exposure time. The center position, rotation direction, and scale of the four-dimensional Gaussian units are transformed to the camera coordinate system and projected onto the image plane. According to the depth order of each Gaussian unit relative to the camera, the camera appearance features and opacity parameters of the Gaussian units are combined to synthesize the transparency to obtain the camera image. Based on the geometric, semantic, and instance attributes carried by the Gaussian units, depth maps, semantic segmentation maps, and instance segmentation maps are generated simultaneously. For cameras with non-zero exposure times, the vehicle pose and dynamic object pose are interpolated within the exposure time range, and the rendering results of multiple exposure sampling times are accumulated.

[0064] S42. Render the ideal lidar geometry according to the beam-by-beam scanning time;

[0065] No. The ideal hit distance of a laser beam is expressed by equation (11):

[0066] (11)

[0067] In the formula, For the first The ideal hit distance of a laser beam; Rendering operators for four-dimensional Gaussian geometry rays; For the scan time The corresponding four-dimensional Gaussian world state; For the first The emission time of the laser beam; Let this be the origin of the laser beam in the world coordinate system; Let be the unit direction vector of the laser beam; The laser beam number;

[0068] Each laser beam updates the vehicle pose and the dynamic object pose using an independent emission time. The cumulative occupancy and transmittance of Gaussian elements are calculated along the ray, and the foremost effective surface is selected as the main echo. For semi-transparent sparse structures, multiple candidate surfaces are retained to form multiple echoes.

[0069] The ideal echo intensity is calculated based on the material, incident angle, distance, and weather conditions, as shown in equation (12):

[0070] (12)

[0071] In the formula, For the first The physical intensity of a laser beam; For the first Near-infrared reflectance response of a Gaussian surface; The wavelength of the laser; The incident angle between the laser beam and the surface normal; It is a cosine function; The atmospheric attenuation function is the weather attenuation function. To prevent positive stability constants where the denominator is zero at close range; For the Gaussian sequence number; The laser beam number;

[0072] The probability of a missed point by the lidar is calculated by equation (13):

[0073] (13)

[0074] In the formula, For the first The probability of a laser beam producing a missing point; For the Sigmoid function; For predicting lost points; The Gaussian property of the hit surface; For ideal hitting distance; Angle of incidence; Configure for lidar equipment; This is a vector representing weather conditions.

[0075] The ideal point cloud and sensor residuals are combined into the final observation, as shown in equation (14):

[0076] (14)

[0077] In the formula, To ultimately generate a lidar point cloud; The sensor residual operator is used; the ideal hit distance of each laser beam is obtained according to equation (11), and the ideal hit point is calculated by combining the emission origin and unit direction vector of the laser beam in the world coordinate system. Then, the corresponding ideal echo intensity is calculated according to equation (12). The ideal point cloud is composed of the ideal hit points of all effective laser beams and their ideal echo intensities. ; This is a set of distance deviation, intensity noise, missing points, and weather scattering output from the sensor residual generator; Configure the equipment; This is a weather condition vector.

[0078] Furthermore, the specific method of S5 is as follows:

[0079] S51. Implement observability-aware cyclic constraints;

[0080] (15)

[0081] In the formula, Total cyclic loss; The deterministic reconstruction loss for the observable region includes image color, depth, point-to-surface distance, intensity, semantics, and instance error; Loss weights are assigned to the imagined regional distribution; Calibrate the loss for road constraints, object shape priors, multi-view conflicts, and uncertainties in the imagined region;

[0082] The overall training objective of the model is expressed as Equation (16):

[0083] (16)

[0084] In the formula, This represents the overall training loss. Reconstruct weights for the camera; Reconstruct loss for color, structure, and perceptual features; Reconstruct weights for the lidar; For distance, intensity, missing points, and point-to-surface error; Weights are assigned to the world flow; Matching loss for world residual flow; For recurrent weights; For dynamic constraint weights; For static background drift, object rigidity, trajectory smoothing and collision penetration loss; Weights for uncertainty; Calibrate the loss to cover low uncertainty in the observable area and the imagined area distribution;

[0085] The training adopts a phased approach. First, multimodal inversion and camera and LiDAR branches are trained. Second, the hierarchical encoder and decoder of Gaussian set are trained. Then, the future world residual flow is trained with the fixed reconstruction module. After the world generation is stable, new viewpoints and cross-modal loops are added. Finally, the frozen autonomous driving perception model is connected for task calibration.

[0086] S52. Evaluate the value of the generated data for autonomous driving verification;

[0087] The generated data is input into the sensing system under test and the task difference is calculated, as shown in equation (17):

[0088] (17)

[0089] In the formula, To perceive the overall differences in tasks; Weights for differences in 3D detection; Differences in target bounding box location, size, heading, and category; To occupy the difference weight; The difference between the drivable area and the 3D occupancy prediction; To track the difference weights; Differences in instance type, speed, and trajectory; For cross-sensor fusion stability weights; To measure the difference in sensing output before and after changing sensor configuration parameters;

[0090] The verification value is calculated by integrating global consistency, perceived differences, and hidden regional uncertainties, as shown in equation (18):

[0091] (18)

[0092] In the formula, The validation value of the generated scenario is scored; Weights for global consistency; It is a natural exponential function; For new perspectives and cross-modal cyclic errors; To perceive task weights; Weights to hide the uncertainty in the region; The hidden region's average uncertainty affects the decision-making of the system under test; only scenarios with cyclic errors below a set threshold are included in the autonomous driving test set to avoid mistaking world representation errors for defects in the system under test;

[0093] The effective cross-viewpoint coverage of the generated scene is evaluated by equation (19):

[0094] (19)

[0095] In the formula, Effective coverage across viewpoints; The number of viewpoints that pass geometric, cyclic, and perceptual consistency checks on a preset virtual sensor trajectory; This represents the total number of viewpoints to be evaluated. It is a natural exponential function; The larger the value, the better the apparent world state remains usable on sensor trajectories that deviate from the original log.

[0096] The beneficial effects of this invention are as follows:

[0097] This invention provides a method for constructing a multimodal driving world for autonomous driving based on four-dimensional Gaussian geometry. It establishes a four-dimensional Gaussian explicit world state comprising multi-view images from cameras, LiDAR point clouds, semantics, instances, motion, and uncertainties. The method integrates world evolution with sensor generation in a hierarchical process, employing deterministic propagation and probabilistic residuals to generate future states, and utilizing physical geometry rendering and sensor residual generation to form LiDAR observations. Differential cyclic verification is implemented based on observable and imagined regions. This invention provides a more realistic multimodal driving world for the safety verification of end-to-end autonomous vehicles, supporting their safety verification and facilitating their industrialization. Attached Figure Description

[0098] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0099] Figure 1 This is a flowchart of the present invention;

[0100] Figure 2 A schematic diagram illustrating the construction of the state of a four-dimensional Gaussian world and the hierarchical world labeling;

[0101] Figure 3 A flowchart for the generation and security verification of future world evolution and multimodal observation;

[0102] Figure 4 This is a schematic diagram of the multimodal driving world corresponding to Scene 1;

[0103] Figure 5 This is a schematic diagram of the multimodal driving world corresponding to Scenario 2. Detailed Implementation

[0104] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.

[0105] Example 1

[0106] See Figure 1 This invention provides a method for constructing a multimodal driving world for autonomous driving based on four-dimensional Gaussian, comprising the following steps:

[0107] S1. Synchronization of multimodal temporal observations and construction of the state of a four-dimensional Gaussian world;

[0108] We perform time synchronization and feature fusion on multi-view images, LiDAR point clouds, vehicle pose, and sensor configuration to construct a four-dimensional Gaussian world state that includes geometric, appearance, echo, semantic, instance, motion, and uncertainty attributes. The specific method is as follows:

[0109] S11. Establish a multimodal time series observation set;

[0110] Time The input observation is represented by equation (1):

[0111] (1)

[0112] In the formula, For a moment A multimodal observation set; For a moment No. Images from a single camera perspective; For camera viewpoint numbers; For a moment LiDAR point cloud; For a moment Pose transformation from vehicle coordinate system to world coordinate system; For a moment The set of sensor configurations; For discrete time sequence numbers;

[0113] Unified time compensation is performed on the exposure times of each camera and the beam-by-beam scanning times of the LiDAR, interpolating the dynamic object state to the corresponding sampling time; camera configuration includes intrinsic parameters, extrinsic parameters, and image size; LiDAR configuration includes beam angle, scanning frequency, mounting pose, maximum range, and intensity calibration. Synchronized image pixels and LiDAR points can all be mapped to a unified world coordinate system.

[0114] S12, Construct the state of a four-dimensional Gaussian explicit world;

[0115] No. The properties of a four-dimensional Gaussian element are represented by equation (2):

[0116] (2)

[0117] In the formula, For a moment The A four-dimensional Gaussian element; The three-dimensional center position; It is a local rotation matrix; It is a three-axis scale vector; As to the degree of space occupation; Features of the camera's appearance; It exhibits near-infrared echo response characteristics; For the first The inherent tendency parameters of Gaussian elements to generate LiDAR point loss; These are motion state parameters, including the three-dimensional velocity vector and local deformation parameters; It is a semantic category vector; For instance category vectors; Let be the world state uncertainty parameter, used to characterize the . The geometric position, shape, and reliability of the generated results of each four-dimensional Gaussian element; The Gaussian element number;

[0118] The image branch extracts color, semantic, depth, and optical flow features, while the LiDAR branch extracts precise location, intensity, scan line, and dynamic point features. Based on the point cloud location, initial Gaussian anchors are constructed by combining sparse voxels and three-plane locations, and the remaining attributes of each initial Gaussian anchor are predicted through a multimodal fusion network. During training, image or point cloud modalities are randomly discarded to ensure that the single-modal coding results and the fusion coding results fall into the same normalized world coordinates.

[0119] The complete world state is divided into a static background, dynamic objects, and an incrementally generated Gaussian set, as shown in equation (3):

[0120] (3)

[0121] In the formula, For a moment The complete four-dimensional Gaussian world state; A static set of backgrounds fixed in the world coordinate system; For dynamic object serial numbers; Total number of dynamic objects; For the first An object at time Pose transformation from the normalized object coordinate system to the world coordinate system; For the first A Gaussian set of objects established in the canonical object coordinate system; For a moment Incremental generation of Gaussian sets; This is a set union operation;

[0122] Static backgrounds include roads, buildings, road edges, and fixed facilities; dynamic object Gaussian sets are used to maintain the cross-frame shape and instance category of vehicles or pedestrians; incrementally generated Gaussian sets are used to represent newly entering objects, surfaces that are first revealed after occlusion, and the birth, death, splitting, and merging of Gaussian elements.

[0123] S2, Observable region division and hierarchical world marker construction;

[0124] Based on the camera and lidar ray coverage, observable Gaussians and imagined Gaussians are classified and encoded as global road and map markers, dynamic object markers, local surface markers, and incrementally generated markers, as follows:

[0125] S21. Divide the observable area and the imagined area according to the sensor rays;

[0126] The four-dimensional Gaussian world is divided according to the observed properties as shown in equation (4):

[0127] (4)

[0128] In the formula, For a moment The complete state of the world; An observable Gaussian set that is fully covered by the rays of a real camera or lidar; For the imagined Gaussian set that is obscured, outside the field of view, or inferred only from prior knowledge; This is a set union operation; For time sequence number;

[0129] S22. Construct a hierarchical world marker;

[0130] Encode the variable number of Gaussian sets into hierarchical latent states, as shown in equation (5):

[0131] (5)

[0132] In the formula, For a moment The hierarchical world potential state; It is a Gaussian set encoder; For global road and map marking; Mark a collection for dynamic objects; A set of local surface markers; Generate a set of tags for incremental generation; For time sequence number;

[0133] Global markers describe road topology, building layout, and overall traffic structure; object markers represent the category, pose, velocity, and shape of each object through set matching; local surface markers use sparse meshes to preserve curbs, fences, and vegetation; incrementally generated marker sets serve as the entry point for generating unobserved areas and new objects, while the Gaussian decoder simultaneously reconstructs geometry, camera appearance, LiDAR response, and instance category.

[0134] S3, Deterministic Propagation and Evolution of the Probabilistic Residual World;

[0135] Based on the current world state, vehicle actions, and the deterministic part of map constraint propagation, and generating object responses, multiple candidate trajectories, and probability residuals corresponding to the appearance of new objects and the completion of occluded regions, the future four-dimensional Gaussian world state is obtained. The specific method is as follows:

[0136] S31, The deterministic part of the propagation of the state of the world;

[0137] Based on the current world potential state, vehicle actions, and the deterministic part of map propagation, as shown in equation (6):

[0138] (6)

[0139] In the formula, The potential state of the deterministic world in the next moment; For coordinate and kinematic propagation operators; For a moment The hierarchical world potential state; For a moment The vehicle's self-control actions; For map constraint set; For time sequence number;

[0140] S32. Utilize conditional flow to generate probabilistic residuals of future world states;

[0141] In the flow matching training, an interpolation path is constructed between the noisy latent state and the true future residual, as shown in Equation (7):

[0142] (7)

[0143] In the formula, For continuous flow time The interpolation latent state at the location; The stream time is a value between 0 and 1; The initial noise latent state is sampled from a standard Gaussian distribution; The target residual latent state is obtained by encoding the real next-moment world state;

[0144] The conditional flow matching loss is shown in equation (8):

[0145] (8)

[0146] In the formula, Matching loss for world residual flow; To calculate the mathematical expectation of the training samples, noise, and streaming time; For parameters A defined conditional velocity field network; This refers to the current action of the vehicle. For map constraints; The square of the second norm of a vector;

[0147] The future world potential state is obtained by combining deterministic propagation and generation residuals, as shown in equation (9):

[0148] (9)

[0149] In the formula, The complete potential state of the world in the next moment; The probabilistic residuals generated for conditional flow describe object response, multimodal trajectory evolution, new object appearance, occlusion region completion, and local non-rigid changes. For time sequence number;

[0150] By changing the initial noise of the flow model under the same current state and vehicle actions, multiple physically plausible future traffic states are generated. When the vehicle actions are changed, the static world remains unchanged, while the viewpoint, occlusion, and other participant reactions change with the actions.

[0151] S4, camera and lidar observation generation;

[0152] Camera images are generated based on the future four-dimensional Gaussian world state, and geometric hit rate, echo intensity, missed point probability, and sensor residual are calculated according to the laser radar beam emission time. This forms a final laser radar point cloud that approximates the characteristics of the real sensor, thereby bridging the domain gap between simulation data and real data. The specific method is as follows:

[0153] S41. Render camera observations based on the future state of the four-dimensional Gaussian world.

[0154] time The The observations from each camera are represented by equation (10):

[0155] (10)

[0156] In the formula, For a moment The generated first One camera image; Rendering operators for a four-dimensional Gaussian camera; For a moment The state of the four-dimensional Gaussian world; For the first Each camera at any time pose transformation; For the first Each camera at any time Configuration parameters;

[0157] The vehicle pose and dynamic object pose are updated according to the camera exposure time. The center position, rotation direction, and scale of the four-dimensional Gaussian units are transformed to the camera coordinate system and projected onto the image plane. According to the depth order of each Gaussian unit relative to the camera, the camera appearance features and opacity parameters of the Gaussian units are combined to synthesize the transparency to obtain the camera image. Based on the geometric, semantic, and instance attributes carried by the Gaussian units, depth maps, semantic segmentation maps, and instance segmentation maps are generated simultaneously. For cameras with non-zero exposure times, the vehicle pose and dynamic object pose are interpolated within the exposure time range, and the rendering results of multiple exposure sampling times are accumulated.

[0158] S42. Render the ideal lidar geometry according to the beam-by-beam scanning time;

[0159] No. The ideal hit distance of a laser beam is expressed by equation (11):

[0160] (11)

[0161] In the formula, For the first The ideal hit distance of a laser beam; Rendering operators for four-dimensional Gaussian geometry rays; For the scan time The corresponding four-dimensional Gaussian world state; For the first The emission time of the laser beam; Let this be the origin of the laser beam in the world coordinate system; Let be the unit direction vector of the laser beam; The laser beam number;

[0162] Each laser beam updates the vehicle pose and the dynamic object pose using an independent emission time. The cumulative occupancy and transmittance of Gaussian elements are calculated along the ray, and the foremost effective surface is selected as the main echo. For semi-transparent sparse structures, multiple candidate surfaces are retained to form multiple echoes.

[0163] The ideal echo intensity is calculated based on the material, incident angle, distance, and weather conditions, as shown in equation (12):

[0164] (12)

[0165] In the formula, For the first The physical intensity of a laser beam; For the first Near-infrared reflectance response of a Gaussian surface; The wavelength of the laser; The incident angle between the laser beam and the surface normal; It is a cosine function; The atmospheric attenuation function is the weather attenuation function. To prevent positive stability constants where the denominator is zero at close range; For the Gaussian sequence number; The laser beam number;

[0166] The probability of a missed point by the lidar is calculated by equation (13):

[0167] (13)

[0168] In the formula, For the first The probability of a laser beam producing a missing point; For the Sigmoid function; For predicting lost points; The Gaussian property of the hit surface; For ideal hitting distance; Angle of incidence; Configure for lidar equipment; This is a vector representing weather conditions.

[0169] The ideal point cloud and sensor residuals are combined into the final observation, as shown in equation (14):

[0170] (14)

[0171] In the formula, To ultimately generate a lidar point cloud; The sensor residual operator is used; the ideal hit distance of each laser beam is obtained according to equation (11), and the ideal hit point is calculated by combining the emission origin and unit direction vector of the laser beam in the world coordinate system. Then, the corresponding ideal echo intensity is calculated according to equation (12). The ideal point cloud is composed of the ideal hit points of all effective laser beams and their ideal echo intensities. ; This is a set of distance deviation, intensity noise, missing points, and weather scattering output from the sensor residual generator; Configure the equipment; This is a vector representing weather conditions.

[0172] The sensor residual generator only modifies the observations and does not change the principal surface of the four-dimensional Gaussian world. Therefore, when the number of lines, vertical field of view, scanning frequency and installation height of the lidar are changed, the world geometry remains unchanged, but the point density and noise domain change accordingly. The camera observations are differentiable rendered based on the Gaussian appearance, transparency, camera pose and exposure parameters, and share the world state at the same moment with the lidar.

[0173] S5, Cyclic Constraints and Automated Driving Safety Verification;

[0174] The generated observations undergo observation loops, latent space loops, new viewpoint loops, and cross-modal loops. The generated scenarios that pass consistency checks are then used for perception and failure verification of the autonomous driving system under test. The specific methods are as follows:

[0175] S51. Implement observability-aware cyclic constraints;

[0176] (15)

[0177] In the formula, Total cyclic loss; The deterministic reconstruction loss for the observable region includes image color, depth, point-to-surface distance, intensity, semantics, and instance error; Loss weights are assigned to the imagined regional distribution; Calibrate the loss for road constraints, object shape priors, multi-view conflicts, and uncertainties in the imagined region;

[0178] The overall training objective of the model is expressed as Equation (16):

[0179] (16)

[0180] In the formula, This represents the overall training loss. Reconstruct weights for the camera; Reconstruct loss for color, structure, and perceptual features; Reconstruct weights for the lidar; For distance, intensity, missing points, and point-to-surface error; Weights are assigned to the world flow; Matching loss for world residual flow; For recurrent weights; For dynamic constraint weights; For static background drift, object rigidity, trajectory smoothing and collision penetration loss; Weights for uncertainty; Calibrate the loss to cover low uncertainty in the observable area and the imagined area distribution;

[0181] The training adopts a phased approach. First, multimodal inversion and camera and LiDAR branches are trained. Second, the hierarchical encoder and decoder of Gaussian set are trained. Then, the future world residual flow is trained with the fixed reconstruction module. After the world generation is stable, new viewpoints and cross-modal loops are added. Finally, the frozen autonomous driving perception model is connected for task calibration.

[0182] S52. Evaluate the value of the generated data for autonomous driving verification;

[0183] The generated data is input into the sensing system under test and the task difference is calculated, as shown in equation (17):

[0184] (17)

[0185] In the formula, To perceive the overall differences in tasks; Weights for differences in 3D detection; Differences in target bounding box location, size, heading, and category; To occupy the difference weight; The difference between the drivable area and the 3D occupancy prediction; To track the difference weights; Differences in instance type, speed, and trajectory; For cross-sensor fusion stability weights; To measure the difference in sensing output before and after changing sensor configuration parameters;

[0186] The verification value is calculated by integrating global consistency, perceived differences, and hidden regional uncertainties, as shown in equation (18):

[0187] (18)

[0188] In the formula, The validation value of the generated scenario is scored; Weights for global consistency; It is a natural exponential function; For new perspectives and cross-modal cyclic errors; To perceive task weights; Weights to hide the uncertainty in the region; The hidden region's average uncertainty affects the decision-making of the system under test; only scenarios with cyclic errors below a set threshold are included in the autonomous driving test set to avoid mistaking world representation errors for defects in the system under test;

[0189] The effective cross-viewpoint coverage of the generated scene is evaluated by equation (19):

[0190] (19)

[0191] In the formula, Effective coverage across viewpoints; The number of viewpoints that pass geometric, cyclic, and perceptual consistency checks on a preset virtual sensor trajectory; This represents the total number of viewpoints to be evaluated. It is a natural exponential function; The larger the value, the better the apparent world state remains usable on sensor trajectories that deviate from the original log.

[0192] In summary, this invention establishes a four-dimensional Gaussian explicit world state encompassing multi-view camera images, LiDAR point clouds, semantics, instances, motion, and uncertainties. It layers the world evolution with sensor generation, employs deterministic propagation and probabilistic residuals to generate future states, and utilizes physical geometry rendering and sensor residual generation to form LiDAR observations. Furthermore, it implements differentiated cyclical verification based on observable and imagined regions. The results of this invention can provide a more realistic multimodal driving world for the safety verification of end-to-end autonomous vehicles, supporting end-to-end autonomous vehicle safety verification and facilitating its industrialization.

[0193] Example 2

[0194] In this embodiment, 700 temporal scenes are selected from the nuScenes dataset to form a training set. The method of this invention is trained using continuous multi-view images, LiDAR point clouds, vehicle pose, and sensor calibration parameters from each temporal scene. A four-dimensional Gaussian explicit world state is constructed based on the input historical multimodal temporal observations, and the future world state evolves under the constraints of vehicle actions and map conditions. This generates time-synchronized camera images and LiDAR point clouds. The generated multimodal driving worlds of the two scenes are as follows: Figure 4 and Figure 5 As shown, where, Figure 4 and Figure 5 In the text, (a) represents a multi-view image branch; Figure 4 and Figure 5 (b) in the diagram represents the LiDAR point cloud branch.

[0195] Depend on Figure 4 and Figure 5 It can be seen that the generated camera images maintain the continuity of the appearance of road structures, traffic participants, and road facilities, and the generated LiDAR point clouds reflect the spatial location and geometric contours of the corresponding objects. The road boundaries, dynamic objects, and fixed facilities observed in the camera and LiDAR observations correspond to each other in spatial location. This demonstrates that the present invention can generate multimodal driving scenarios with spatiotemporal consistency based on a unified four-dimensional Gaussian world state, which can be used for safety verification of end-to-end autonomous driving systems.

[0196] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for constructing a multimodal driving world for autonomous driving based on four-dimensional Gaussian, characterized in that, Includes the following steps: S1. Synchronization of multimodal temporal observations and construction of the state of a four-dimensional Gaussian world; Time synchronization and feature fusion are performed on multi-view images, LiDAR point clouds, vehicle pose and sensor configuration to construct a four-dimensional Gaussian world state that includes geometric, appearance, echo, semantic, instance, motion and uncertainty attributes. S2, Observable region division and hierarchical world marker construction; Based on the coverage of camera and lidar rays, observable Gaussians and imagined Gaussians are classified and coded as global road and map markers, dynamic object markers, local surface markers, and incremental generation markers. S3, Deterministic Propagation and Evolution of the Probabilistic Residual World; Based on the current world state, the vehicle's actions, and the deterministic part of the map constraints, and by generating object responses, multiple candidate trajectories, the probability residuals corresponding to the appearance of new objects and the completion of occluded areas, the future four-dimensional Gaussian world state is obtained. S4, camera and lidar observation generation; The camera image is generated based on the future four-dimensional Gaussian world state, and the geometric hit, echo intensity, probability of missing points and sensor residual are calculated according to the beam emission time of the lidar to form a lidar point cloud. S5, Cyclic Constraints and Automated Driving Safety Verification; The generated observations are subjected to observation loops, latent space loops, new viewpoint loops, and cross-modal loops, and the generated scenarios that pass the consistency check are used for perception and failure verification of the autonomous driving system under test.

2. The method for constructing a multimodal driving world for autonomous driving based on four-dimensional Gaussian as described in claim 1, characterized in that, The specific method of S1 is as follows: S11. Establish a multimodal time series observation set; Time The input observation is represented by equation (1): (1) In the formula, For a moment A multimodal observation set; For a moment No. Images from a single camera perspective; For camera viewpoint numbers; For a moment LiDAR point cloud; For a moment Pose transformation from vehicle coordinate system to world coordinate system; For a moment The set of sensor configurations; For discrete time sequence numbers; Unified time compensation is performed on the exposure time of each camera and the beam-by-beam scanning time of the LiDAR, and the dynamic object state is interpolated to the corresponding sampling time; the camera configuration includes intrinsic parameters, extrinsic parameters and image size; the LiDAR configuration includes beam angle, scanning frequency, installation pose, maximum range and intensity calibration; the synchronized image pixels and LiDAR points can be mapped to a unified world coordinate system. S12, Construct the state of a four-dimensional Gaussian explicit world; No. The properties of a four-dimensional Gaussian element are represented by equation (2): (2) In the formula, For a moment The A four-dimensional Gaussian element; The three-dimensional center position; It is a local rotation matrix; It is a three-axis scale vector; As to the degree of space occupation; Features of the camera's appearance; It exhibits near-infrared echo response characteristics; For the first The inherent tendency parameters of Gaussian elements to generate LiDAR point loss; These are motion state parameters, including the three-dimensional velocity vector and local deformation parameters; It is a semantic category vector; For instance category vectors; Let be the world state uncertainty parameter, used to characterize the . The geometric position, shape, and reliability of the generated results of each four-dimensional Gaussian element; The Gaussian element number; The image branch extracts color, semantic, depth, and optical flow features, while the LiDAR branch extracts precise location, intensity, scan line, and dynamic point features. Based on the point cloud location, initial Gaussian anchors are constructed by combining sparse voxels and three-plane locations, and the remaining attributes of each initial Gaussian anchor are predicted through a multimodal fusion network. During training, image or point cloud modalities are randomly discarded to ensure that the single-modal coding results and the fusion coding results fall into the same normalized world coordinates. The complete world state is divided into a static background, dynamic objects, and an incrementally generated Gaussian set, as shown in equation (3): (3) In the formula, For a moment The complete four-dimensional Gaussian world state; A static set of backgrounds fixed in the world coordinate system; For dynamic object serial numbers; Total number of dynamic objects; For the first An object at time Pose transformation from the normalized object coordinate system to the world coordinate system; For the first A Gaussian set of objects established in the canonical object coordinate system; For a moment Incremental generation of Gaussian sets; This is a set union operation; Static backgrounds include roads, buildings, road edges, and fixed facilities; dynamic object Gaussian sets are used to maintain the cross-frame shape and instance category of vehicles or pedestrians; incrementally generated Gaussian sets are used to represent newly entering objects, surfaces that are first revealed after occlusion, and the birth, death, splitting, and merging of Gaussian elements.

3. The method for constructing a multimodal driving world for autonomous driving based on four-dimensional Gaussian as described in claim 1, characterized in that, The specific method of S2 is as follows: S21. Divide the observable area and the imagined area according to the sensor rays; The four-dimensional Gaussian world is divided according to the observed properties as shown in equation (4): (4) In the formula, For a moment The complete state of the world; An observable Gaussian set that is fully covered by the rays of a real camera or lidar; For the imagined Gaussian set that is obscured, outside the field of view, or inferred only from prior knowledge; This is a set union operation; For time sequence number; S22. Construct a hierarchical world marker; Encode the variable number of Gaussian sets into hierarchical latent states, as shown in equation (5): (5) In the formula, For a moment The hierarchical world potential state; It is a Gaussian set encoder; For global road and map marking; Mark a collection for dynamic objects; A set of local surface markers; Generate a set of tags for incremental generation; For time sequence number; Global markers describe road topology, building layout, and overall traffic structure; object markers represent the category, pose, velocity, and shape of each object through set matching; local surface markers use sparse meshes to preserve curbs, fences, and vegetation; incrementally generated marker sets serve as the entry point for generating unobserved areas and new objects, while the Gaussian decoder simultaneously reconstructs geometry, camera appearance, LiDAR response, and instance category.

4. The method for constructing a multimodal driving world for autonomous driving based on four-dimensional Gaussian as described in claim 1, characterized in that, The specific method of S3 is as follows: S31, The deterministic part of the propagation of the state of the world; Based on the current world potential state, vehicle actions, and the deterministic part of map propagation, as shown in equation (6): (6) In the formula, The potential state of the deterministic world in the next moment; For coordinate and kinematic propagation operators; For a moment The hierarchical world potential state; For a moment The vehicle's self-control actions; For map constraint set; For time sequence number; S32. Utilize conditional flow to generate probabilistic residuals of future world states; In the flow matching training, an interpolation path is constructed between the noisy latent state and the true future residual, as shown in Equation (7): (7) In the formula, For continuous flow time The interpolation latent state at the location; The stream time is a value between 0 and 1; The initial noise latent state is sampled from a standard Gaussian distribution; The target residual latent state is obtained by encoding the real next-moment world state; The conditional flow matching loss is shown in equation (8): (8) In the formula, Matching loss for world residual flow; To calculate the mathematical expectation of the training samples, noise, and streaming time; For parameters A defined conditional velocity field network; This refers to the current action of the vehicle. For map constraints; The square of the second norm of a vector; The future world potential state is obtained by combining deterministic propagation and generation residuals, as shown in equation (9): (9) In the formula, The complete potential state of the world in the next moment; The probabilistic residuals generated for conditional flow describe object response, multimodal trajectory evolution, new object appearance, occlusion region completion, and local non-rigid changes. For time sequence number; By changing the initial noise of the flow model under the same current state and vehicle actions, multiple physically plausible future traffic states are generated. When the vehicle actions are changed, the static world remains unchanged, while the viewpoint, occlusion, and other participant reactions change with the actions.

5. The method for constructing a multimodal driving world for autonomous driving based on four-dimensional Gaussian as described in claim 1, characterized in that, The specific method of S4 is as follows: S41. Render camera observations based on the future state of the four-dimensional Gaussian world. time The The observations from each camera are represented by equation (10): (10) In the formula, For a moment The generated first One camera image; Rendering operators for a four-dimensional Gaussian camera; For a moment The state of the four-dimensional Gaussian world; For the first Each camera at any time pose transformation; For the first Each camera at any time Configuration parameters; The vehicle pose and dynamic object pose are updated according to the camera exposure time. The center position, rotation direction, and scale of the four-dimensional Gaussian units are transformed to the camera coordinate system and projected onto the image plane. According to the depth order of each Gaussian unit relative to the camera, the camera appearance features and opacity parameters of the Gaussian units are combined to synthesize the transparency to obtain the camera image. Based on the geometric, semantic, and instance attributes carried by the Gaussian units, depth maps, semantic segmentation maps, and instance segmentation maps are generated simultaneously. For cameras with non-zero exposure times, the vehicle pose and dynamic object pose are interpolated within the exposure time range, and the rendering results of multiple exposure sampling times are accumulated. S42. Render the ideal lidar geometry according to the beam-by-beam scanning time; No. The ideal hit distance of a laser beam is expressed by equation (11): (11) In the formula, For the first The ideal hit distance of a laser beam; Rendering operators for four-dimensional Gaussian geometry rays; For the scan time The corresponding four-dimensional Gaussian world state; For the first The emission time of the laser beam; Let this be the origin of the laser beam in the world coordinate system; Let be the unit direction vector of the laser beam; The laser beam number; Each laser beam updates the vehicle pose and the dynamic object pose using an independent emission time. The cumulative occupancy and transmittance of Gaussian elements are calculated along the ray, and the foremost effective surface is selected as the main echo. For semi-transparent sparse structures, multiple candidate surfaces are retained to form multiple echoes. The ideal echo intensity is calculated based on the material, incident angle, distance, and weather conditions, as shown in equation (12): (12) In the formula, For the first The physical intensity of a laser beam; For the first Near-infrared reflectance response of a Gaussian surface; The wavelength of the laser; The incident angle between the laser beam and the surface normal; It is a cosine function; The atmospheric attenuation function is the weather attenuation function. To prevent positive stability constants where the denominator is zero at close range; For the Gaussian sequence number; The laser beam number; The probability of a missed point by the lidar is calculated by equation (13): (13) In the formula, For the first The probability of a laser beam producing a missing point; For the Sigmoid function; For predicting lost points; The Gaussian property of the hit surface; For ideal hitting distance; Angle of incidence; Configure for lidar equipment; This is a vector representing weather conditions. The ideal point cloud and sensor residuals are combined into the final observation, as shown in equation (14): (14) In the formula, To ultimately generate a lidar point cloud; The sensor residual operator is used; the ideal hit distance of each laser beam is obtained according to equation (11), and the ideal hit point is calculated by combining the emission origin and unit direction vector of the laser beam in the world coordinate system. Then, the corresponding ideal echo intensity is calculated according to equation (12). The ideal point cloud is composed of the ideal hit points of all effective laser beams and their ideal echo intensities. ; This is a set of distance deviation, intensity noise, missing points, and weather scattering output from the sensor residual generator; Configure the equipment; This is a weather condition vector.

6. The method for constructing a multimodal driving world for autonomous driving based on four-dimensional Gaussian as described in claim 1, characterized in that, The specific method of S5 is as follows: S51. Implement observability-aware cyclic constraints; (15) In the formula, Total cyclic loss; The deterministic reconstruction loss for the observable region includes image color, depth, point-to-surface distance, intensity, semantics, and instance error; Loss weights are assigned to the imagined regional distribution; Calibrate the loss for road constraints, object shape priors, multi-view conflicts, and uncertainties in the imagined region; The overall training objective of the model is expressed as Equation (16): (16) In the formula, This represents the overall training loss. Reconstruct weights for the camera; Reconstruct loss for color, structure, and perceptual features; Reconstruct weights for the lidar; For distance, intensity, missing points, and point-to-surface error; Weights are assigned to the world flow; Matching loss for world residual flow; For recurrent weights; For dynamic constraint weights; For static background drift, object rigidity, trajectory smoothing and collision penetration loss; Weights for uncertainty; Calibrate the loss to cover low uncertainty in the observable area and the imagined area distribution; The training adopts a phased approach. First, multimodal inversion and camera and LiDAR branches are trained. Second, the hierarchical encoder and decoder of Gaussian set are trained. Then, the future world residual flow is trained with the fixed reconstruction module. After the world generation is stable, new viewpoints and cross-modal loops are added. Finally, the frozen autonomous driving perception model is connected for task calibration. S52. Evaluate the value of the generated data for autonomous driving verification; The generated data is input into the sensing system under test and the task difference is calculated, as shown in equation (17): (17) In the formula, To perceive the overall differences in tasks; Weights for differences in 3D detection; Differences in target bounding box location, size, heading, and category; To occupy the difference weight; The difference between the drivable area and the 3D occupancy prediction; To track the difference weights; Differences in instance type, speed, and trajectory; For cross-sensor fusion stability weights; To measure the difference in sensing output before and after changing sensor configuration parameters; The verification value is calculated by integrating global consistency, perceived differences, and hidden regional uncertainties, as shown in equation (18): (18) In the formula, The validation value of the generated scenario is scored; Weights for global consistency; It is a natural exponential function; For new perspectives and cross-modal cyclic errors; To perceive task weights; Weights to hide the uncertainty in the region; The hidden region's average uncertainty affects the decision-making of the system under test; only scenarios with cyclic errors below a set threshold are included in the autonomous driving test set to avoid mistaking world representation errors for defects in the system under test; The effective cross-viewpoint coverage of the generated scene is evaluated by equation (19): (19) In the formula, Effective coverage across viewpoints; The number of viewpoints that pass geometric, cyclic, and perceptual consistency checks on a preset virtual sensor trajectory; This represents the total number of viewpoints to be evaluated. It is a natural exponential function; The larger the value, the better the apparent world state remains usable on sensor trajectories that deviate from the original log.