Automatic driving simulation mode switching method and device, equipment and medium
By acquiring target scene information at the switching moment to construct a virtual test scenario with a closed-loop simulation mode, the problem of environmental deviation during mode switching in autonomous driving simulation testing is solved, and seamless connection and smooth switching between the simulation environment and playback data are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD
- Filing Date
- 2026-02-26
- Publication Date
- 2026-06-02
AI Technical Summary
In current autonomous driving simulation testing, when switching from open-loop playback mode to closed-loop simulation mode, the simulation environment after the switch deviates significantly from the original real scene, making it difficult to maintain consistency and authenticity.
By acquiring the target scene information at the switching moment of the open-loop playback mode, a virtual test scene of the closed-loop simulation mode is constructed. The scene is dynamically updated using the feedback control commands of the autonomous driving system to ensure the spatiotemporal continuity between the simulation environment and the playback data.
Seamless switching between the simulation environment and playback data was achieved, ensuring the consistency and realism of the test scenario and guaranteeing efficient and smooth switching of simulation tests.
Smart Images

Figure CN122133334A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving simulation testing technology, and in particular to autonomous driving simulation mode switching methods, devices, equipment and media. Background Technology
[0002] The development and testing of autonomous driving systems heavily rely on simulation technology. Currently, the mainstream testing process typically includes two phases: open-loop playback mode and closed-loop simulation mode. Open-loop playback mode replays real-world sensor data (images, point clouds, etc.) to the autonomous driving system at specific times, verifying the accuracy of the perception algorithm under known truth conditions. Closed-loop simulation, on the other hand, runs the complete autonomous driving system in a virtual environment, testing the system's comprehensive decision-making and control capabilities in an interactive virtual environment.
[0003] However, these two modes are currently disconnected. When testers want to switch from a potentially risky scenario discovered during playback to a closed-loop simulation for in-depth interactive testing, the initial state of the closed-loop simulation is difficult to keep in line with the state of the real scenario at the moment of playback interruption, resulting in a significant deviation between the switched simulation environment and the original real scenario. Summary of the Invention
[0004] Based on this, a method, device, equipment, and medium for switching autonomous driving simulation modes are provided to solve the problem that the simulation environment after switching from open-loop playback mode to closed-loop simulation mode deviates significantly from the original real scene.
[0005] Firstly, this application provides a method for switching autonomous driving simulation modes, the method comprising: In response to the command to switch from open-loop playback mode to closed-loop simulation mode, determine the switching time; Obtain the target scene information corresponding to the open-loop playback mode at the switching time; the target scene information includes at least the description information of the static environment and the state information of at least one dynamic target at the switching time; Based on the target scenario information, a virtual test scenario for the closed-loop simulation mode is constructed; Based on the virtual test scenario, simulation data is generated and provided to the autonomous driving system. In response to the control commands fed back by the autonomous driving system based on the simulation data, the virtual test scenario is updated to complete the switch to the closed-loop simulation mode.
[0006] Using the above method, a closed-loop virtual test scenario is seamlessly constructed based on the static environment and dynamic target state information corresponding to the switching moment. This achieves continuous spatiotemporal connection between the simulation environment and playback data. The scenario is dynamically updated by relying on the control commands fed back by the autonomous driving system. Thus, while ensuring the consistency and authenticity of the test scenario, the test mode is switched efficiently and smoothly.
[0007] In one embodiment, obtaining the target scene information corresponding to the open-loop playback mode at the switching time includes: Obtain the vehicle pose corresponding to the switching time; Based on the switching time and the corresponding vehicle pose, the scene reconstruction information of the switching time is queried from the preset scene model, wherein the preset scene model is constructed based on the log data of the open-loop playback mode; The scene reconstruction information is analyzed to separate the first type of elements representing the static environment and the second type of elements representing the dynamic target; The description information of the static environment is reconstructed based on the first type of elements, and the state information of at least one dynamic target at the switching moment is reconstructed based on the second type of elements.
[0008] By combining the vehicle pose at the switching moment with the preset scene model, the static environment elements and dynamic target elements corresponding to the switching moment can be efficiently and accurately retrieved from the complex scene reconstruction information constructed based on log data. Based on this, the target scene information used to initialize the closed-loop simulation can be reconstructed, thereby ensuring that the starting point of the simulation test and the historical data of the open-loop playback are strictly aligned in geometric position and logical state.
[0009] In one embodiment, reconstructing the state information of at least one dynamic target at the switching moment based on the second type of elements includes: Based on the spatiotemporal correlation of the second type of elements in a preset neighborhood at the switching time, the second type of elements are clustered to obtain at least one set of elements, wherein each set of elements corresponds to a dynamic target; Based on each set of elements, the geometric information, appearance information, and motion state information of the corresponding dynamic target are determined to generate the state information.
[0010] Using the above method, intelligent clustering of dynamic target elements based on spatiotemporal correlation can effectively identify and separate different dynamic targets from discrete raw data, thereby accurately reconstructing the geometric shape, visual features and instantaneous motion state of each target. This provides accurate, complete and individual-discriminative dynamic target state information for closed-loop simulation, ensuring the realism and accuracy of dynamic interactions in the simulation environment.
[0011] In one embodiment, the method for constructing the preset scene model includes: Obtain log data for the open-loop playback mode; wherein the log data includes at least a time-synchronized sequence of real-scene images, a 3D point cloud sequence, and a corresponding vehicle pose sequence; Based on the log data, a three-dimensional scene representation model is constructed; the three-dimensional scene representation model is used to perform unified spatiotemporal coding of static environment and dynamic targets in real scene. Select the target time and corresponding vehicle pose from the log data, and render the corresponding reconstructed image from the three-dimensional scene representation model based on the target time and corresponding vehicle pose. The parameters of the 3D scene representation model are optimized with the goal of minimizing the difference between the reconstructed image and the real scene image corresponding to the target time. The optimized 3D scene representation model is the preset scene model.
[0012] Using the above method, a unified spatiotemporal encoded 3D scene representation model is constructed using time-synchronized multimodal log data. A rendering optimization method based on the difference between real and reconstructed images is used to iteratively correct the model parameters. Finally, a preset scene model that can faithfully reproduce static environments and dynamic targets in the real world is generated, providing an accurate and efficient scene data foundation for subsequent mode switching.
[0013] In one embodiment, based on the target scenario information, constructing a virtual test scenario for the closed-loop simulation mode includes: Based on the description information of the static environment in the target scene information, a virtual static environment is constructed; Based on the state information of the dynamic target in the target scene information, the target three-dimensional model of the dynamic target and the initial state of the dynamic target are determined, wherein the initial state includes at least the initial pose, initial velocity and initial acceleration; In the virtual static environment, an entity corresponding to the dynamic target is created; The entity is bound to the target 3D model, and the initial state is configured on the entity to construct the virtual test scene.
[0014] Using the above method, the parsed target scene information is converted into elements that can be directly processed by the simulation, and these elements are accurately associated and instantiated, thereby efficiently and accurately constructing a virtual test scene that is strictly aligned with the historical playback data of the open-loop playback mode in terms of time, space and state, and can be directly used for closed-loop interactive testing.
[0015] In one embodiment, configuring the initial state after the entity further includes: The entity is bound to a preset behavior model, and the behavior model is controlled to enter a first stage; in the first stage, the behavior model is configured to control the entity to maintain the initial state; In response to the fulfillment of preset triggering conditions, the behavior model is controlled to enter the second stage from the first stage; in the second stage, the behavior model is configured to perform autonomous decision-making and interactive behaviors based on the generated simulation data.
[0016] By using the above method, the behavior model of the entity is bound and controlled in two stages after the entity is created. This ensures that the dynamic target can maintain the initial state of the open-loop playback in the early stage of mode switching to ensure scene consistency. After meeting specific conditions, it smoothly transitions to the autonomous decision-making and interaction stage, thereby realizing the seamless and controllable connection from the real state of historical playback data to the dynamic interactive behavior in closed-loop simulation.
[0017] In one embodiment, generating simulation data based on the virtual test scenario and providing it to the autonomous driving system includes: Obtain the simulation data generated at the current moment, and obtain the historical playback data of the open-loop playback mode at the switching moment; Determine the fusion weight corresponding to the current moment, and fuse the historical playback data and the simulation data generated at the current moment based on the fusion weight to obtain fused data; The fused data is provided to the autonomous driving system.
[0018] Using the above method, after the closed-loop simulation is started, the fusion weight is dynamically adjusted according to time, and the real-time generated simulation data is fused with the historical playback data at the switching moment. This generates a hybrid perception input that retains the characteristics of the real historical scene and includes the current simulation evolution, thereby providing a test environment for the autonomous driving system with continuous and smooth state transitions, effectively alleviating the problem of perception abrupt changes or data inconsistency that may be caused by mode switching.
[0019] Secondly, this application provides an autonomous driving simulation mode switching device, the device comprising: The determination module is used to determine the switching time in response to the instruction to switch from open-loop playback mode to closed-loop simulation mode; The acquisition module is used to acquire the target scene information corresponding to the open-loop playback mode at the switching time; the target scene information includes at least the description information of the static environment and the state information of at least one dynamic target at the switching time; A construction module is used to construct a virtual test scenario for the closed-loop simulation mode based on the target scenario information; The switching module is used to generate simulation data based on the virtual test scenario and provide it to the autonomous driving system, and to update the virtual test scenario in response to the control command fed back by the autonomous driving system based on the simulation data, thereby completing the switch to the closed-loop simulation mode.
[0020] Thirdly, this application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the autonomous driving simulation mode switching method described in the first aspect above.
[0021] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the autonomous driving simulation mode switching method described in the first aspect above.
[0022] The technical effects of each of the second to fourth aspects mentioned above, as well as the technical effects that each aspect may achieve, are described above with reference to the technical effects that can be achieved for the first aspect or the various possible solutions in the first aspect, and will not be repeated here. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the system architecture of an autonomous driving simulation mode switching method in one embodiment; Figure 2 This is a flowchart illustrating an autonomous driving simulation mode switching method in one embodiment; Figure 3 This is a flowchart illustrating a preset scene model construction method in one embodiment; Figure 4 This is a schematic diagram of the structure of an autonomous driving simulation mode switching device in one embodiment; Figure 5 This is a schematic diagram of the internal structure of a computer device in one embodiment. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The specific operational methods in the method embodiments can also be applied to the device embodiments or system embodiments. It should be noted that in the description of this application, "multiple" is understood as "at least two". "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing together, or B existing alone. A connected to B can represent: A and B directly connected, or A and B connected through C. Furthermore, in the description of this application, terms such as "first" and "second" are used only for distinguishing the purpose of description and should not be construed as indicating or implying relative importance or order.
[0025] Before introducing the autonomous driving simulation mode switching method provided in this application, the technical background of this application will be described in detail below for ease of understanding.
[0026] The development and testing of autonomous driving systems heavily rely on simulation technology. Currently, the mainstream testing process typically includes two relatively independent phases: open-loop playback verification and closed-loop simulation testing. Open-loop playback directly replays the collected sensor data back to the autonomous driving system to verify the accuracy of the perception module under known truth conditions. However, for testing the decision-making and control systems, the collected data needs to be reconstructed into a dynamic virtual environment understandable by the simulation platform, within which interactive testing is performed.
[0027] In related technologies, a simulation testing method based on scene reconstruction and editing is disclosed. This method first performs offline analysis on log data collected from real roads, extracting the historical motion trajectories of dynamic targets. Based on these trajectories, a parameterized simplified behavioral model is established, generating a structured, editable scene description file. Testers can manually modify this file, such as adding or deleting dynamic targets, adjusting their speed or path, to construct new test cases. During the simulation execution phase, the simulation engine loads and runs this edited scene description file, using it as a basis to uniformly render and generate a complete virtual sensor data stream, which is then output to the autonomous driving system for testing. This method can retain the basic framework of the real scene while introducing a certain degree of interactive variability through manual editing, thus enabling the testing of autonomous driving systems in specific scenarios.
[0028] However, the above methods still have significant limitations. First, model simplification leads to insufficient state fidelity, failing to accurately reproduce the instantaneous motion state (such as precise velocity and acceleration) and fine-grained appearance attributes of a target at a given moment in a real-world scenario. Second, the generation of interactive behaviors relies on predefined simplified models, making it difficult to simulate the complex, diverse, and even seemingly irrational decisions of real human drivers, thus limiting the depth of testing. Finally, its essence is to construct an independent and complete simulation test case. When testers want to immediately switch from an observed real-world scene segment in open-loop playback to closed-loop simulation, the initial state of the closed-loop simulation is difficult to maintain consistency with the real-world scene at the moment of interruption in open-loop playback, resulting in a significant deviation between the switched simulation environment and the original real-world scene.
[0029] Therefore, how to achieve highly realistic scene reconstruction, seamless and smooth switching between different simulation modes, and generate rich and diverse interactive behaviors have become urgent technical problems to be solved.
[0030] In view of this, this application provides an autonomous driving simulation mode switching method, device, computer equipment and storage medium to solve the problems of low scene realism, difficulty in keeping the initial state of closed-loop simulation consistent with the real scene at the moment of open-loop playback interruption and rigid interactive behavior in related technologies.
[0031] The following is a brief introduction to the system architecture to which the technical solution of this application is applicable. It should be noted that the system architecture described below is for illustrative purposes only and not for limitation. In specific implementations, the technical solution provided in this application can be flexibly applied according to actual needs.
[0032] This application provides a method for switching autonomous driving simulation modes, which can be applied to, for example... Figure 1 The system architecture shown mainly includes: an autonomous driving system, a data playback unit, a scene reconstruction unit, a scene parsing unit, a simulation unit, and a simulation scheduling and management unit. The connections and data interaction relationships between these units are as follows: The output of the data playback unit is connected to the input of the autonomous driving system to provide playback data to the autonomous driving system; The data interface of the data playback unit is connected to the data interface of the scene reconstruction unit to transmit log data collected by the vehicle in real-world scenarios. The output of the scene reconstruction unit is connected to the input of the scene parsing unit, and is used to transmit the preset scene model built based on log data; The output of the scene parsing unit is connected to the configuration interface of the simulation unit, and is used to provide the simulation unit with scene information queried from the preset scene model; The sensor data output terminal of the simulation unit is connected to the input terminal of the simulation scheduling and management unit to provide feedback on the generated simulation data (including at least virtual sensor data and simulated vehicle status signals). The output control terminal of the simulation scheduling management unit is connected to the start / stop control terminal of the data playback unit and the start / stop control terminal of the simulation unit, thereby realizing the overall coordination and control of the operating status of the two units; the sensor data output terminal of the simulation scheduling management unit is connected to the input terminal of the autonomous driving system to provide simulation data to the autonomous driving system. The vehicle control command input terminal of the simulation unit is connected to the output terminal of the autonomous driving system to receive control commands generated by the autonomous driving system and update the virtual test scenario accordingly, thereby forming a closed loop.
[0033] For example, the data playback unit is the core execution unit of the open-loop playback mode. Its main functions are: to receive and parse the log data collected by the autonomous vehicle in real-world scenarios, and to play back the sensor data and vehicle status signals (such as vehicle speed, wheel speed, steering angle, etc.) in the log data to the autonomous driving system strictly according to the original time sequence through the data distribution service and the controller area network bus.
[0034] The log data includes time-synchronized real-scene image sequences, 3D point cloud sequences, and corresponding vehicle pose sequences.
[0035] For example, multiple vehicle-mounted cameras acquire continuous image frames (which are the basis for reconstructing color, texture, and appearance); vehicle-mounted LiDAR acquires continuous 3D point cloud frames (which provide precise geometric spatial constraints and initialization for scene reconstruction); and a combination of a Global Positioning System (GPS) and an Inertial Measurement Unit (IMU), or a Global Navigation Satellite System (GNSS) and an IMU, acquires the six-degree-of-freedom pose corresponding to each frame of image and point cloud (described by three translational degrees of freedom and three rotational degrees of freedom, including X-axis translation: horizontal forward and backward movement, Y-axis translation: horizontal left and right movement, Z-axis translation: vertical up and down movement; roll: rotation around the X-axis, pitch: rotation around the Y-axis, and yaw: rotation around the Z-axis).
[0036] The collected raw data is preprocessed to obtain a time-synchronized real-scene image sequence, a 3D point cloud sequence, and a corresponding vehicle pose sequence. The preprocessing includes, but is not limited to: time synchronization processing (aligning data streams from different sensors based on a unified time reference), sensor calibration processing (calibrating the internal parameters of each camera and the external parameters between the camera and the LiDAR to establish a unified sensor coordinate system), and data cleaning processing (removing invalid image frames, point cloud noise points, and abnormal positioning data) to ensure data quality.
[0037] The scene reconstruction unit is the foundation for building the virtual environment. Its main function is to receive log data from the data playback unit and, based on preset scene reconstruction technology, construct a high-fidelity, queryable preset scene model that includes static backgrounds and dynamic targets (such as vehicles and pedestrians). This model serves as the information source for subsequently extracting scene information at any given time.
[0038] The scene parsing unit's main function is to automatically retrieve the corresponding target scene information from the preset scene model in response to the command to switch from open-loop playback mode to closed-loop simulation mode, based on the switching time specified by the simulation scheduling and management unit. This target scene information includes at least: Description of the static environment: This is high-precision map data described in standard formats (such as OpenDRIVE, OpenSCENARIO, etc.), which defines in detail the location, geometry, and semantic type of all static elements such as lane lines, curbs, and traffic signs.
[0039] Status information of at least one dynamic target at the switching moment: The status information of at least one dynamic target at the switching moment is listed in the form of a list (the list contains the dynamic target identifier).
[0040] The simulation unit is the core execution unit of the closed-loop simulation mode, and it has a built-in virtual world engine. Its main function is to construct a virtual test scenario for the closed-loop simulation mode based on the target scene information received from the scene parsing unit.
[0041] In addition, the simulation unit also has a built-in general vehicle dynamics model to simulate the basic physical behavior of vehicles in virtual test scenarios, including dynamic responses such as acceleration, braking, and steering, providing basic vehicle motion simulation capabilities for closed-loop simulation.
[0042] Optionally, to improve the realism and accuracy of the simulation, the simulation scheduling and management unit can also maintain a high-precision model database (containing various detailed vehicle and human models). Based on the appearance information in the state information of the dynamic target, the corresponding high-precision model can be matched from the model database to replace the default general model, thereby more accurately reproducing the unique dynamic characteristics of a specific dynamic target.
[0043] The simulation scheduling and management unit is the control center of the entire system. It receives mode switching commands from users and coordinates and controls all the aforementioned units to achieve seamless, high-fidelity switching. Its main functions include: The system controls the start and stop of the data playback unit to freeze the open-loop state; the scheduling scene parsing unit extracts scene information at a specified time; the instruction simulation unit loads the scene information and constructs a virtual test scene; at the mode switching time, it manages the switching of the data link (that is, switches the data source and control loop of the autonomous driving system to be driven by the simulation unit) and ensures data synchronization between units, thereby controlling the entire system to transition from open-loop playback mode to closed-loop simulation mode.
[0044] The autonomous driving system, as the core driver of the test object and the closed-loop formation, has the following main functions: In open-loop playback mode, it passively receives sensor data and vehicle status signals from the data playback unit, and performs perception, localization, planning, and decision-making based on this data to output vehicle control commands (such as throttle, brake, steering angle, etc.). At this time, the output is mainly used for algorithm verification and historical behavior reproduction, and does not interact with the environment.
[0045] When switching to closed-loop simulation mode, the system receives virtual sensor data and simulated vehicle state signals rendered in real time from the simulation unit. Based on these new perceptual inputs, its algorithm modules (perception, prediction, planning, and control) make driving decisions in real time and output new vehicle control commands. These control commands are fed back to the vehicle dynamics model in the simulation unit in real time, driving its movement. The movement of the vehicle dynamics model updates the virtual test scenario, which in turn affects the simulation data input to the autonomous driving system in the next moment. This forms a complete "perception-decision-control-scenario update" test closed loop, allowing the autonomous driving system to undergo testing in complex scenarios involving dynamic interactions in a high-fidelity, safe, controllable, and reproducible simulation environment.
[0046] The deployment of the above systems is highly flexible and can be adapted to different computing resources, network conditions, and test scales. This mainly includes: Integrated local deployment mode: All functional units in the system architecture are integrated and deployed on the same high-performance local workstation or server. All data interaction and process control are completed through local inter-process communication or shared memory. This mode is simple to deploy, has the lowest latency, and is suitable for single-machine development, debugging, and small-scale testing scenarios.
[0047] Cloud-edge collaborative distributed deployment mode: Adopting a cloud-edge collaborative architecture, it separates computationally intensive offline tasks from real-time interactive tasks.
[0048] For example, the scene reconstruction unit can be deployed as a high-performance computing service in the cloud. This service receives raw data packets from the local machine, utilizes the massive computing power of the cloud's graphics processing units to efficiently construct a preset scene model, and stores the preset scene model in cloud object storage; other architectures besides the scene reconstruction unit are deployed on the local test terminal. The terminal pulls the preset scene model from the cloud as needed and performs real-time data playback, model parsing, scene construction, and closed-loop simulation testing locally.
[0049] This mode effectively solves the bottleneck of local hardware computing power and realizes elastic resource scheduling, making it particularly suitable for teams that need to process massive amounts of road data or build large-scale scenario libraries.
[0050] Furthermore, all core functionalities of the system can be encapsulated as independent microservices, each with a clear application programming interface (API). This architecture achieves complete decoupling of business logic and resource deployment, greatly improving the system's scalability and flexibility in continuous integration / continuous deployment environments.
[0051] The above deployment methods can be used individually or in combination. Users can choose the most suitable architecture based on their actual testing needs, infrastructure conditions, and development and operation strategies.
[0052] The technical solution provided in this application will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0053] Figure 2 The following is a flowchart illustrating an autonomous driving simulation mode switching method in one embodiment. The process includes the following steps: S201, in response to the command to switch from open-loop playback mode to closed-loop simulation mode, determines the switching time; S202, Obtain the target scene information corresponding to the switching time of the open-loop playback mode; the target scene information includes at least the description information of the static environment and the state information of at least one dynamic target at the switching time; S203, based on target scene information, constructs a virtual test scene with a closed-loop simulation mode; S204 generates simulation data based on the virtual test scenario and provides it to the autonomous driving system. In response to the control commands fed back by the autonomous driving system based on the simulation data, it updates the virtual test scenario and completes the switch to closed-loop simulation mode.
[0054] Among them, the open-loop playback mode is a mode that replays the sensor data stream from the log data to the autonomous driving system based on log data.
[0055] For example, when the simulation scheduling management unit receives a user instruction or an instruction issued by an automatic triggering rule to switch from open-loop playback mode to closed-loop simulation mode, it determines the switching time and immediately commands the data playback unit to pause after completing the playback of the frame data corresponding to the switching time. The data playback unit then pauses playback, freezes the current state, and records the context, for example, recording the precise timestamp corresponding to the current playback data packet (i.e., the switching time), and recording all output states of the autonomous driving system at the switching time, such as vehicle speed and vehicle pose.
[0056] The simulation scheduling and management unit initiates the scene reconstruction unit and the scene parsing unit. The scene reconstruction unit queries the scene reconstruction information at the switching time from the preset scene model based on the switching time and the corresponding vehicle pose. Subsequently, the scene parsing unit parses the scene reconstruction information to obtain a structured target scene information in a format readable by the simulation unit, such as a lightweight data exchange format (JavaScript Object Notation, JSON) or an Extensible Markup Language (XML). This target scene information includes at least: a description of the static environment and the state information of at least one dynamic target at the switching time.
[0057] The simulation scheduling and management unit sends the target scenario information to the simulation unit. Based on the target scenario information, the simulation unit constructs a virtual test scenario in a closed-loop simulation mode.
[0058] The simulation unit includes, but is not limited to: an open-source simulation platform for autonomous driving research (CarLearning to Act, CARLA), the LG Silicon Valley Lab simulator (LGSVL), a joint simulation scheme of vehicle dynamics simulation software (CarSim) and dynamic system modeling and simulation platform (Simulink), or any other simulation environment that supports custom maps and external control interfaces, depending on the circumstances and not limited here.
[0059] The simulation unit generates simulation data (including at least virtual sensor data and simulated vehicle state signals) based on a virtual test scenario. At this point, the simulation scheduling and management unit disconnects the data transmission link from the data playback unit to the autonomous driving system and establishes a new data transmission link between the simulation unit and the autonomous driving system. The virtual sensor states (such as initial pose and viewpoint) in the simulation unit are aligned with the real sensor states in the log data at the switching moment to ensure the spatiotemporal continuity of the perception input. Simultaneously, the vehicle state signals calculated by the vehicle dynamics model in the simulation unit are simulated and sent to the autonomous driving system via the controller area network bus, replacing the vehicle state signals replayed by the data playback unit.
[0060] The autonomous driving system makes decisions based on simulation data sent by the simulation unit and outputs new vehicle control commands. These new control commands are fed back to the vehicle dynamics model in the simulation unit in real time, driving the vehicle dynamics model to move and update the virtual test scenario. This completes the switch to closed-loop simulation mode.
[0061] Using the above methods, the switching moment is precisely locked and the structured target scene information is reconstructed using a preset scene model, thereby constructing a virtual test environment with consistent spatiotemporal state. At the same time, by switching data links, aligning the initial state of sensors, and introducing real-time interaction between vehicle dynamics models and autonomous driving control commands, the continuity of perception input and the dynamic response of the simulation environment are ensured. Finally, while maintaining the authenticity of historical data, a smooth transition to closed-loop interactive testing is achieved.
[0062] Alternatively, the virtual test scenario construction in the above scheme can be replaced with an offline preprocessing-online index matching mode to adapt to scenarios with stringent runtime performance requirements, including but not limited to: A series of potential high-value switching moments are pre-selected, and an interactive simulation scenario file is generated for each switching moment; during the actual switching, the closest file is called through a matching mechanism.
[0063] For example, log data is scanned and analyzed, and a series of candidate moments are manually or automatically marked. For each candidate moment, a corresponding high-fidelity static environment is generated based on 3D reconstruction technology (such as 3D Gaussian Splatting (3DGS) or Neural Radiance Fields (NeRF) technology), and an artificial intelligence behavior model is used to simulate multiple possible behavioral branches from the real log data. Finally, the complete scene state (static environment + initial state of dynamic target) of each candidate moment is packaged with the simulated behavioral branches to generate an independent, interactive simulation scene file, and an index relationship is established between each candidate moment and this simulation scene file.
[0064] In response to the command to switch from open-loop playback mode to closed-loop simulation mode, the switching time is determined. Based on the switching time, the target simulation scene file corresponding to the closest target candidate time is queried from a pre-established index relationship. The simulation unit generates simulation data based on the target simulation scene file and provides it to the autonomous driving system, completing the switch to closed-loop simulation mode.
[0065] The above method sacrifices the arbitrariness of the switching time and the absolute precision of the state in exchange for extremely low runtime computation latency and stable performance. It is suitable for batch testing needs with relatively fixed test scenarios that can be pre-enumerated.
[0066] Figure 3 This is a flowchart illustrating a preset scene model construction method in one embodiment, illustrating the application of this method to... Figure 1 Taking the scene reconstruction unit as an example, the process includes the following steps: S301, acquire log data in open-loop playback mode; the log data includes at least a time-synchronized sequence of real-scene images, a 3D point cloud sequence, and the corresponding vehicle pose sequence; S302, based on log data, constructs a 3D scene representation model; the 3D scene representation model is used to perform unified spatiotemporal coding of static environment and dynamic targets in real scene; S303: Select the target time and corresponding vehicle pose from the log data, and render the corresponding reconstructed image from the 3D scene representation model based on the target time and corresponding vehicle pose. S304 optimizes the parameters of the 3D scene representation model with the goal of minimizing the difference between the reconstructed image and the real scene image corresponding to the target time; the optimized 3D scene representation model is the preset scene model.
[0067] In this embodiment, the scene reconstruction unit uses 3DGS technology to construct the scene model. Under this technology, static environmental elements and dynamic targets in the real scene are explicitly represented by a series of 3D Gaussian ellipsoids with optimizable properties. The properties of each 3D Gaussian ellipsoid include, but are not limited to: Location: The center coordinates of the 3D Gaussian ellipsoid in three-dimensional space; Covariance: Defines the shape and spatial extension of the 3D Gaussian ellipsoid; Opacity: Controls the contribution weight of the 3D Gaussian ellipsoid to the final rendered color; Spherical Harmonic Function Coefficients: Used to model the view-dependent color appearance to accurately reproduce the gloss and texture details of materials.
[0068] The preset scene model is constructed as follows: First, an initial 3D Gaussian ellipsoidal distribution is generated based on the sparse point cloud recovered from the real scene image sequence of the log data using motion reconstruction structure technology, or the 3D point cloud sequence of the log data.
[0069] To provide a unified representation of static environmental elements and dynamic targets in real-world scenes, a temporal modeling capability is introduced into a 3D Gaussian ellipsoidal distribution, thereby constructing a 3D scene representation model with spatiotemporal encoding capabilities. The temporal modeling capability is achieved through at least one of the following methods: Train a time-conditioned neural network deformation field to predict the position and shape changes of each 3D Gaussian ellipsoid at a specific time. The 3D Gaussian ellipsoids are divided into static and dynamic categories, and the property parameters of the dynamic 3D Gaussian ellipsoids as they evolve over time are explicitly modeled.
[0070] Then, the target time and corresponding vehicle pose are selected from the log data, and the 3D Gaussian ellipsoid corresponding to the 3D scene representation model at the target time is projected onto the 2D image plane through a differentiable rasterizer to render the corresponding reconstructed image.
[0071] Finally, based on a preset loss function, and with the objective of minimizing the difference between the reconstructed image and the real scene image corresponding to the target time, the attribute parameters of the 3D Gaussian ellipsoid are optimized through backpropagation until the preset loss function value is less than a preset threshold (the specific threshold depends on the situation and is not limited here). The resulting 3D scene representation model is the preset scene model. This model contains high-fidelity geometry and appearance of the static environment and inherently encodes the motion trajectory and state changes of dynamic targets in the time dimension. The preset loss function can be constructed as follows: Based on the difference in pixel color between the reconstructed image and the real scene image, a luminosity and loss function L is constructed. P , where L P To optimize appearance and texture, L1 loss (absolute error) or L2 loss (mean squared error) are commonly used; a depth consistency loss L is constructed based on the difference between the rendered depth of the reconstructed image and the true depth of the real scene image. D , where L D To constrain the accuracy of geometry, L1 loss or L2 loss is commonly used; based on L... P With L D Construct a pre-defined loss function Loss=λ1×L P +λ2×L D λ1 and λ2 are used to balance the contributions of different loss terms in the optimization process. Their specific values depend on the situation and are not limited here.
[0072] Using the methods described above, an explicit and differentiable spatiotemporal unified 3D scene representation model based on 3DGS dynamic scene modeling technology was constructed. This model introduces a time dimension to unify the encoding of static environment and dynamic targets, and optimizes it using a loss function that includes photometric and depth constraints. Finally, a preset scene model that can reproduce the geometric appearance and dynamic evolution of real scene with high fidelity is obtained, providing an accurate, coherent and renderable scene reconstruction foundation for seamless switching of simulation modes.
[0073] In one embodiment, exemplarily described, the acquisition of target scene information corresponding to the switching time of open-loop playback mode includes, but is not limited to: The vehicle pose corresponding to the switching time is obtained. That is, the simulation scheduling management unit sends the vehicle pose corresponding to the switching time recorded by the data playback unit to the scene reconstruction unit.
[0074] The scene reconstruction unit queries the scene reconstruction information at the switching time from a preset scene model based on the switching time and the corresponding vehicle pose. The preset scene model is constructed based on log data from the open-loop playback mode. The scene reconstruction information is then sent to the scene parsing unit through the simulation scheduling management unit.
[0075] For example, the scene reconstruction unit extracts the Gaussian points (a 3D Gaussian ellipsoid and its attribute parameters) corresponding to the switching time from the preset scene model based on the switching time and the corresponding vehicle pose, and renders the corresponding switching reconstruction image (which is consistent with the above rendering method and will not be described again here). The internal scene representation and the switching reconstruction image corresponding to the switching time are used as scene reconstruction information.
[0076] The scene parsing unit analyzes the scene reconstruction information to separate the first type of elements representing the static environment and the second type of elements representing the dynamic target.
[0077] For example, through motion consistency analysis, a set of Gaussian points with time-varying poses (such as moving objects like vehicles, pedestrians, and cyclists) is identified from the Gaussian points in the scene reconstruction information as the second type of elements representing dynamic targets, and the remaining set of Gaussian points (such as background elements like roads, buildings, and vegetation) is identified as the first type of elements representing the static environment.
[0078] Thus, the description information of the static environment can be reconstructed based on the first type of elements, and the state information of at least one dynamic target at the switching moment can be reconstructed based on the second type of elements.
[0079] Using the above method, the corresponding Gaussian point set and reconstructed image are located and rendered based on the switching time and vehicle pose. Motion consistency analysis is used to intelligently distinguish between static and dynamic elements, thereby reconstructing a structured description of the static environment and the state of the dynamic target, providing a key and reliable data foundation for seamlessly constructing the initial scene of the closed-loop simulation.
[0080] In one embodiment, exemplarily illustrated, the description information of the static environment is reconstructed based on the first type of elements, including but not limited to: The scene parsing unit performs target detection and semantic segmentation on the switched reconstructed image in the scene reconstruction information, and identifies the 2D semantic region and semantic category of each target in the image (such as car, truck, pedestrian, bicycle, traffic cone, etc.).
[0081] Based on the semantic segmentation results, the semantic category corresponding to the first type of element is determined.
[0082] For example, the first type of element is projected onto the switched reconstructed image to form a corresponding first projection region. The overlap (e.g., intersection-union ratio) of each 2D semantic region in the semantic segmentation result of the first projection region and the switched reconstructed image is calculated, and the semantic category of the semantic region with the highest overlap is taken as the semantic category corresponding to the first type of element.
[0083] Therefore, based on the first type of elements and their corresponding semantic categories, a high-precision map conforming to the open and standardized road network description format (OpenDRIVE) is generated. This map accurately describes the geometric and semantic information of static environmental elements such as lane lines, curbs, traffic signs, and traffic lights.
[0084] By combining the semantic recognition of 2D images with the projection matching of 3D Gaussian point clouds, static elements are assigned precise semantic categories, thereby generating high-precision maps and realizing the accurate reconstruction of static environments from discrete geometric point clouds to structured and semantic road network models.
[0085] In one embodiment, it is exemplarily illustrated that the state information of at least one dynamic target at the switching moment is reconstructed based on the second type of elements, including but not limited to: Based on the spatiotemporal correlation of the second type of elements in the preset neighborhood at the switching time, the second type of elements are clustered to obtain at least one set of elements, where each set of elements corresponds to a dynamic target.
[0086] For example, the cooperative motion pattern of each second type element in the spatiotemporal dimension within the preset neighborhood at the switching moment is analyzed; second type elements with consistent motion trajectories and spatial clustering are aggregated into an element set, and each element set represents the instantaneous, dense 3D representation of a dynamic target at the switching moment, with clear spatiotemporal boundaries.
[0087] Based on each set of elements, the geometric information, appearance information, and motion state information of the corresponding dynamic target are determined to generate state information. The geometric information includes at least the position, orientation, and size of the dynamic target; the appearance information is used to characterize the visual features of the dynamic target; and the motion state information includes at least the instantaneous velocity and acceleration of the dynamic target.
[0088] For example, the method for determining geometric information is as follows: perform spatial distribution analysis on all second-class elements in each element set, and calculate the center position of the corresponding dynamic target by weighted average; based on the spatial covariance matrix of the point cloud, determine the three-dimensional orientation (yaw angle, pitch angle, roll angle) of the corresponding dynamic target by principal component analysis; fit a 3D bounding box along the principal component direction to obtain the length, width, and height dimensions of the corresponding dynamic target.
[0089] The method for determining appearance information is as follows: aggregate the spherical harmonic function coefficients of all second-class elements in each element set to reconstruct the color and texture features of the corresponding dynamic target; and statistically analyze the opacity distribution of all second-class elements to characterize the visual density and internal structural features of the corresponding dynamic target.
[0090] The method for determining motion state information is as follows: analyze the change of the center position of each element set in the target neighborhood (greater than the preset neighborhood) at the switching moment, and calculate the instantaneous velocity vector and acceleration vector of the corresponding dynamic target based on the change of the center position using the numerical differentiation method.
[0091] Using the methods described above, the spatiotemporal correlation-based clustering algorithm accurately aggregates dynamic Gaussian point clouds into independent dynamic targets. By comprehensively utilizing spatial distribution analysis, spherical harmonic coefficient aggregation, and motion trajectory differentiation, the precise geometric shape, visual appearance, and instantaneous kinematic state of each dynamic target are fully reconstructed. This provides high-fidelity initial state information of dynamic targets with complete physical properties for closed-loop simulation.
[0092] In one embodiment, the state information further includes, exemplarily, the semantic category of the dynamic target, and the method for determining the semantic category of the dynamic target includes, but is not limited to: The semantic category of a dynamic target is determined based on the semantic segmentation results of the switched-reconstructed image. For example, for each set of dynamic elements, its Gaussian points are projected onto the plane of the switched-reconstructed image to form a corresponding projection region. The overlap between this projection region and each 2D semantic region in the semantic segmentation result is calculated, and the semantic category of the semantic region with the highest overlap is taken as the semantic category of the dynamic target corresponding to that set of dynamic elements.
[0093] By projecting the 3D point cloud of a dynamic target onto a 2D image and matching it with the semantic segmentation results, we can efficiently and accurately assign a specific semantic category (such as a car, a pedestrian, etc.) to each dynamic target, thereby enhancing the semantic integrity of the dynamic target's state information and supporting accurate behavior modeling and interaction based on categories in subsequent simulations.
[0094] In one embodiment, exemplarily illustrated, after reconstructing the state information of at least one dynamic target at the switching moment based on the second type of elements, the method further includes: The simulation analysis unit creates a corresponding digital twin data structure for each dynamic target based on its semantic category, geometric information, appearance information, and motion state information. This structure serves as the state information for each dynamic target at each transition point. The digital twin data structure includes a dynamic target identifier: a globally unique ID to ensure consistent tracking throughout the system; a semantic category; geometric information; appearance information; and motion state information. Thus, the description information of the static environment and the digital twin data structure of at least one dynamic target are used as the target scene information.
[0095] Using the above method, a digital twin data structure containing a globally unique ID, semantics, geometry, appearance, and motion state information is created for each dynamic target. Discrete attributes are integrated into a logically unified and independently traceable simulation entity, thereby providing a standardized data foundation for closed-loop simulation to perform accurate identity management, continuous state updates, and complex interactive simulation of dynamic targets.
[0096] Optionally, the preset scene model can also be built based on NeRF technology to provide a high-quality static environment base for the simulation unit with photorealistic quality.
[0097] The method for constructing the preset scene model is as follows: The 3D scene representation and reconstructed image rendering in 3DGS technology are configured as a neural rendering unit based on NeRF and its efficient variants. This unit takes log data used in open-loop playback mode as input, trains a multilayer perceptron neural network to implicitly model the continuous volumetric radiation field of the scene, and uses the trained multilayer perceptron neural network as a preset scene model. During simulation mode switching, when static environmental information at the switching moment is needed, a high-fidelity image observed from any specified viewpoint at the switching moment can be rendered in real time by querying the preset scene model.
[0098] To meet the simulation unit's requirement for importing explicit geometry, this solution can also include a geometry extraction process: using a preset isosurface extraction algorithm (such as the moving cube method or the traveling tetrahedron method), the explicit 3D mesh of the scene is extracted from a preset scene model, and a corresponding texture map is generated. This explicit 3D mesh and texture map are then imported into the simulation unit as a directly interactive static 3D model. This allows for the construction of a virtual test scene for a closed-loop simulation mode based on the retrieved high-fidelity image and the static 3D model.
[0099] Using the methods described above, an implicit scene model based on NeRF is constructed, achieving high-fidelity, photorealistic reconstruction of real-world static environments. It supports both real-time generation of realistic images from any perspective through neural rendering to directly drive simulations, and generation of explicit 3D meshes and textures through a geometry extraction process for import and use by simulation units. This provides another technologically advanced and flexible complete path to eliminate visual discrepancies in simulations.
[0100] In one embodiment, an exemplary illustration is provided, which constructs a virtual test scenario with a closed-loop simulation mode based on target scenario information, including but not limited to: The scene analysis unit sends the target scene information to the simulation unit through the simulation scheduling and management unit. Based on the description information of the static environment in the target scene information, the simulation unit constructs a virtual static environment, including lane lines, curbs, traffic signs, building outlines, etc., forming a static base for high-fidelity simulation.
[0101] Subsequently, for each dynamic target in the target scene information, the following refined initialization process is executed to replace the use of the default general model of the simulation unit: First, based on the state information of the dynamic target in the target scene information, the target 3D model of the dynamic target is determined. The method for determining the target 3D model can be as follows: Using surface reconstruction algorithms (such as Poisson reconstruction or neural surface reconstruction), a basic 3D mesh is generated based on the geometric information in the state information. Then, texture mapping is applied to the basic 3D mesh based on the appearance information in the state information, ultimately forming a target 3D model with realistic textures. Alternatively, Based on the appearance information in the state information, a similarity match is performed with a high-precision model database, and the model with the highest appearance similarity is selected as the base 3D model. Subsequently, a non-rigid registration algorithm is used to fine-tune the base 3D model according to the geometric information in the state information to obtain the target 3D model, which can align its geometry with the real target while preserving the model's own high-detail features.
[0102] Simultaneously, based on the state information of the dynamic target in the target scene information, the initial state of the dynamic target is determined. The initial state includes at least the initial pose (position and orientation in geometric information), the initial velocity (instantaneous velocity in motion state information), and the initial acceleration (acceleration in motion state information).
[0103] Then, in the virtual static environment, the simulation unit creates an entity corresponding to the dynamic target (the entity's ID is the same as the corresponding dynamic target's ID). Subsequently, the entity is bound to the target's 3D model as its visual representation in the virtual static environment. An initial state is then configured on the entity, precisely placing it in the position and orientation defined by the initial pose, and assigning it initial velocity and initial acceleration.
[0104] After completing all the aforementioned steps, the simulation unit successfully initialized a high-fidelity virtual test scenario. This virtual test scenario not only possesses a static road network consistent with the real world, but each dynamic target entity within it also exhibits geometric consistency, appearance consistency, and motion state consistency.
[0105] Thus, the initial conditions of the closed-loop simulation and the physical and visual states of the real scene at the switching moment are mapped and aligned with a high degree of losslessness, laying a solid and reliable simulation environment foundation for subsequent testing of autonomous driving algorithms based on real scenes and featuring both high-fidelity vision and precise physical interaction.
[0106] In one embodiment, exemplarily illustrated, configuring the initial state after the entity further includes: The simulation unit binds entities to preset behavioral models, which are responsible for driving the behavior of the entities during simulation. It also controls the behavioral models to enter the first phase; in the first phase, the behavioral models are configured to control the entities to maintain their initial state.
[0107] Among them, the behavior model of the dynamic target can be the artificial intelligence behavior model built into the simulation unit, which can quickly build basic interactions.
[0108] Optionally, for tests requiring more complex, more human-like, or scenario-specific behaviors, trained reinforcement learning models or more advanced cognitive behavioral models can be deployed for key dynamic targets to generate richer and more challenging interactive behaviors.
[0109] Optionally, the behavioral model can be designed as a hybrid system. This system first uses an intent classifier to analyze the entity's historical trajectory and state in the log data of the open-loop playback mode, identifying its initial behavioral intent (such as cutting in, maintaining a stable following position, or merging into a lane). Subsequently, based on the identified initial behavioral intent, the system activates a set of predefined, finely tuned rule controllers (such as a more aggressive cutting-in model or a more conservative following model).
[0110] In response to the fulfillment of preset triggering conditions, the control behavior model moves from the first stage to the second stage. In the second stage, the behavior model is configured to perform autonomous decision-making and interactive behaviors based on the acquired simulation data (such as the positions of surrounding vehicles, traffic signal status, etc.). For example, a vehicle that is moving at a constant speed in the log data may decelerate and avoid collisions with other vehicles in the closed-loop simulation.
[0111] The preset trigger conditions can be flexibly configured according to testing needs, including but not limited to: Time trigger: When the simulation duration is detected to be longer than the preset duration threshold (e.g., 2 seconds, the specific threshold depends on the situation and is not limited here); Event Trigger: A deviation is detected between the behavior of the tested vehicle and its historical behavior in the log data. For example, the lateral control offset of the tested vehicle exceeds a preset offset threshold (the specific threshold depends on the situation and is not limited here), or there is a significant discrepancy between the planned trajectory and the historical trajectory; Command Trigger: Testers can manually trigger the switch via commands to achieve precise control over the testing process.
[0112] By using the above method, the evolution of the behavior model is configured and flexibly triggered in stages, so that the dynamic target can strictly maintain its historical state in the initial stage of simulation to ensure the consistency of the scene. Then, after the preset triggering conditions are met, it smoothly transitions to the autonomous decision-making and interaction stage based on real-time simulation data. This realizes the controlled and smooth switching of behavior mode from historical playback reproduction to closed-loop intelligent interaction, thereby effectively enhancing the authenticity, complexity and controllability of the test scene.
[0113] In one embodiment, exemplarily illustrated, before generating simulation data based on a virtual test scenario and providing it to the autonomous driving system, the method further includes: The scheduling and management unit drives the simulation unit to render a simulation image based on the virtual test scenario and the same sensor states at the switching moment in open-loop simulation mode. This simulation image aims to completely reproduce the scenario state at the instant of switching in open-loop mode.
[0114] The scheduling and management unit compares the generated simulation image with the original real image corresponding to the open-loop playback mode at the switching time, such as pixel-level comparison or semantic feature-level comparison, and calculates the difference value.
[0115] If the difference value is less than the preset difference threshold (the specific threshold depends on the situation and is not limited here), the driving simulation unit generates simulation data according to the virtual test scenario and provides it to the autonomous driving system.
[0116] If the difference value is greater than or equal to a preset difference threshold, it indicates that the initial state of the virtual test scene deviates from the real scene at the level of detail (such as lighting, material texture, precise outline or position of objects). At this time, the system will automatically start an optimization iteration process: based on the difference value, according to preset adjustment rules (including the adjusted parameters and the corresponding adjustment amounts), the virtual test scene parameters are adjusted (such as adjusting the ambient light intensity and angle, object surface material properties, pose of dynamic targets, etc.). Subsequently, rendering and comparison are performed again until the difference value is less than the preset difference threshold or the number of iterations is greater than the preset number of iterations threshold (the specific threshold depends on the situation and is not limited here).
[0117] By introducing an automatic verification and iterative optimization mechanism based on image comparison, the rendering effect of the virtual test scene at the moment of switching is highly consistent with the real scene visually. This effectively eliminates the problem of mismatch in perception input caused by model reconstruction or environmental rendering deviation, and provides more realistic and reliable simulation data input for the initial state of the autonomous driving system.
[0118] Optionally, to significantly reduce computational resource consumption while ensuring the accuracy of the simulation images, the image sequences in the log data used in the open-loop playback mode can be processed. Each frame of image is extracted as a high-fidelity background library, and its corresponding precise vehicle pose is calculated to establish an image frame-vehicle pose index relationship.
[0119] When running in closed-loop simulation mode, the simulation unit renders a simulation image based on the virtual test scenario and the same sensor states at the switching time in open-loop simulation mode. Then, based on the initial pose of the vehicle under test in the simulation image, the unit quickly retrieves and obtains the most matching real background image from the background library using a preset image frame-vehicle pose index.
[0120] Image processing techniques (such as semantic segmentation) are used to extract dynamic targets from the simulated image as foreground elements and seamlessly integrate them into the real background image obtained in the previous step to generate the final synthetic image. This synthetic image, along with simulation data from other sensors (such as LiDAR and vehicle bus), is then output to the autonomous driving system.
[0121] This solution achieves a high degree of scene realism and dynamic interactivity at the visual level by directly fusing "real background" and "simulated foreground driven by high-fidelity virtual scene" at the image pixel level. It is suitable for test scenarios that are sensitive to runtime performance, have high requirements for visual fidelity, and can accept consistency at the image level (rather than absolute three-dimensional geometric space).
[0122] In one embodiment, simulation data is generated based on a virtual test scenario and provided to an autonomous driving system, including but not limited to: First, obtain the simulation data generated at the current moment, and obtain the historical playback data of the open-loop playback mode at the switching moment.
[0123] For example, after receiving the mode switching instruction, the simulation scheduling management unit commands the data playback unit to pause after completing the transmission of the data frame corresponding to the switching time. This data frame is the historical playback data. At the same time, the simulation scheduling management unit commands the simulation unit to render and generate the sensor data at the current time in real time based on the high-fidelity virtual test scene initialized at the switching time. The sensor data includes at least image data and point cloud data, which is the simulation data generated at the current time.
[0124] Then, the fusion weight corresponding to the current moment is determined, and the historical replay data and the simulation data generated at the current moment are fused based on the fusion weight to generate fused data. The fusion weight consists of a first weight w1 corresponding to the historical replay data and a second weight w2 corresponding to the simulation data, and w1 + w2 = 1.
[0125] The fusion weights are dynamically determined based on the simulation duration from the current time to the switching time. Specifically, at the switching time, w1=1 and w2=0; as the simulation duration increases, w1 decreases gradually and w2 increases gradually; when the simulation duration reaches the preset target duration threshold (the specific threshold depends on the situation and is not limited here), w1=0 and w2=1; thereafter, fusion is no longer performed, and the simulation data is directly output to the autonomous driving system.
[0126] The fused data is provided to the perception module of the autonomous driving system as input, completing a smooth transition from open-loop playback to closed-loop simulation of the data source.
[0127] By dynamically adjusting the fusion weights using the above method, the system initially relies entirely on historical playback data during mode switching, and then gradually transitions to relying entirely on simulation data. This provides the autonomous driving perception system with a continuous and seamless fusion data stream of visual and perceptual inputs, effectively avoiding perception jitter or decision instability caused by instantaneous switching of data sources.
[0128] In one embodiment, exemplarily illustrating the generation of simulation data based on a virtual test scenario and its provision to the autonomous driving system, the method further includes: Within a continuous N-frame transition window starting from the switching moment, the fusion weight is determined based on the frame order of the simulation data frames generated at the current moment within the transition window. Here, N is an integer greater than 1, typically ranging from [2, 5], corresponding to a physical duration usually of tens to over one hundred milliseconds. The fusion weight consists of a first weight w1 corresponding to the historical playback data and a second weight w2 corresponding to the simulation data, where w1 + w2 = 1. Based on the fusion weight, the historical playback data and the simulation data generated at the current moment are fused to obtain fused data, which is then provided to the autonomous driving system.
[0129] The fusion weights vary according to the frame order of the simulation data frame within the transition window, allowing the fused data to smoothly transition from being entirely derived from historical playback data to being entirely derived from simulation data.
[0130] For example, N=3, meaning the transition window frame order includes 0, 1, and 2. When the simulation data frame's frame order within the transition window is 0 (the start frame of the window), w1=1 and w2=0, meaning the fused data is 100% historical playback data. When the simulation data frame's frame order within the transition window is 1 (the middle frame of the window), w1=0.5 and w2=0.5. When the simulation data frame's frame order within the transition window is 2 (the end frame of the window), w1=0 and w2=1, meaning the fused data is 100% simulation image. After the transition window ends, fusion stops, and each subsequent frame of simulation data is directly output to the autonomous driving system.
[0131] By using the above method, a very short but precise transition window is set up to quickly and smoothly transition from 100% historical playback data to 100% simulation data within several consecutive frames after the switch. This effectively ensures the visual continuity and temporal coherence of the perceptual input at the switching critical point, avoids screen jumps, and minimizes computational overhead and system latency.
[0132] In one embodiment, exemplarily illustrating the fusion of historical playback data and simulation data based on fusion weights, including but not limited to: If the simulation data is image data, then the historical playback data and the simulation data are weighted and fused pixel by pixel based on the fusion weight to obtain the fused image.
[0133] For example, in a historical playback image, a pixel's value in a specific color channel (such as red, green, or blue) is Pixel_Log, while in the simulation image, the value in the same channel at the corresponding position is Pixel_Sim. Then, based on the first weight w1 and the second weight w2 determined at the current moment, the value of the corresponding pixel in that channel in the fused image is calculated as Pixel_Fused = w1 × Pixel_Log + w2 × Pixel_Sim. By iterating through all color channels of all pixels in both the historical playback image and the simulation image, a visually smooth fused image is generated.
[0134] If the simulation data is point cloud data, then spatial and attribute interpolation fusion operations are performed on the historical playback data and the simulation data based on the fusion weight.
[0135] For example, interpolation and fusion of coordinates and attributes can be performed on point cloud pairs that have a corresponding relationship between historical playback point clouds and simulated point clouds, or on point clouds that are spatially adjacent: Coordinate fusion: The 3D coordinates of a point cloud in the historical playback point cloud are (x1, y1, z1), and the 3D coordinates of the corresponding point cloud in the simulated point cloud are (x2, y2, z2). Then the 3D coordinates of the fused point cloud are (w1×x1+w2×x2, w1×y1+w2×y2, w1×z1+w2×z2). Attribute fusion: The reflection intensity or other attribute values of the point cloud are weighted and fused according to the fusion weights.
[0136] In addition, to address the sparsity of point cloud data and the robustness of the algorithm, the system also provides a progressive switching strategy as an alternative: at the beginning frame of the transition window, 100% of the historical playback point cloud is output; at the end frame of the transition window, 100% of the simulated point cloud is output; in the middle stage of the transition window, the above interpolation fusion can be performed according to actual needs, or the point cloud data source can be switched directly by using a single-frame hard switching method according to the tolerance of the perception algorithm (the perception algorithm of the autonomous driving system has a higher tolerance for point cloud mutations than that of images).
[0137] By employing the methods described above, differentiated fusion methods are designed for different data types, and flexible gradual or hard switching strategies are provided. This enables a smooth and natural transition of perception data from historical replay to simulation generation in both visual appearance and physical space dimensions. This effectively reduces the impact of mode switching on autonomous driving perception algorithms and ensures the continuity and stability of perception results.
[0138] In one embodiment, as exemplarily illustrated, when generating and providing simulation data to an autonomous driving system based on a virtual test scenario, to ensure strict timing consistency of various simulation data (such as images, point clouds, and controller area network bus signals) in closed-loop simulation, this solution deploys a high-precision synchronization system that combines software and hardware, including but not limited to: Global hardware clock synchronization: All physical computing nodes that generate simulation data in the simulation unit (such as image rendering server, point cloud simulation unit, vehicle dynamics solution module, etc.) are synchronized at the hardware level through a high-precision time synchronization protocol to establish a unified global clock reference and achieve microsecond-level time alignment.
[0139] Simulation Clock Stepping and Scheduling: Based on a global clock reference, the simulation scheduling management unit employs a step-based clock driving mechanism. This unit broadcasts synchronized stepping instructions to each physical computing node at a fixed simulation step size. Each physical computing node independently completes its data calculations within the specified logical simulation time according to the instructions and reports its completion status to the simulation scheduling management unit. Once the simulation scheduling management unit confirms that all physical computing nodes are ready, it advances the global simulation clock to the next time step. This mechanism logically ensures the causal consistency of calculations from all data sources at the same simulation time.
[0140] Time-aligned data publishing: The system has a data publishing scheduler. This scheduler continuously receives data frames (such as image frames, point cloud frames, and bus signal frames) from each physical computing node, each marked with a precise target publishing timestamp. For each target publishing time point, the scheduler waits until all types of data frames belonging to that time have arrived in its buffer before publishing this complete, time-aligned multimodal simulation data packet to the autonomous driving system.
[0141] Through the multi-layered collaborative mechanism from physical clock and simulation logic to data release, the high-precision timing consistency of the simulation data stream is systematically guaranteed, providing a high-fidelity and reliable closed-loop testing environment for autonomous driving algorithms.
[0142] It should be understood that, although Figure 2 and 3 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 2 and 3 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0143] In one embodiment, such as Figure 4As shown, an autonomous driving simulation mode switching device is provided, including: a determining module 401, an acquiring module 402, a constructing module 403, and a switching module 404, wherein: The determination module 401 is used to determine the switching time in response to the instruction to switch from open-loop playback mode to closed-loop simulation mode; The acquisition module 402 is used to acquire the target scene information corresponding to the switching time of the open-loop playback mode; the target scene information includes at least the description information of the static environment and the state information of at least one dynamic target at the switching time; Module 403 is used to construct a virtual test scenario with a closed-loop simulation mode based on target scene information; The switching module 404 is used to generate simulation data based on the virtual test scenario and provide it to the autonomous driving system, and to update the virtual test scenario in response to the control commands fed back by the autonomous driving system based on the simulation data, thereby completing the switch to the closed-loop simulation mode.
[0144] For specific limitations regarding an autonomous driving simulation mode switching device, please refer to the limitations of an autonomous driving simulation mode switching method described above, which will not be repeated here. Each module in the aforementioned autonomous driving simulation mode switching device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0145] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores data for switching between autonomous driving simulation modes. The network interface communicates with external terminals via a network connection. When the processor executes the computer program, it implements a method for switching between autonomous driving simulation modes. The display screen can be an LCD screen or an e-ink display screen. The input device can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device casing, or an external keyboard, touchpad, or mouse.
[0146] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0147] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: In response to the command to switch from open-loop playback mode to closed-loop simulation mode, determine the switching time; Obtain the target scene information corresponding to the switching time in open-loop playback mode; the target scene information includes at least the description information of the static environment and the state information of at least one dynamic target at the switching time; Based on the target scenario information, a virtual test scenario with a closed-loop simulation mode is constructed; Based on the virtual test scenario, simulation data is generated and provided to the autonomous driving system. In response to the control commands fed back by the autonomous driving system based on the simulation data, the virtual test scenario is updated to complete the switch to closed-loop simulation mode.
[0148] In one embodiment, the processor, when executing a computer program, also performs the following steps: Obtain the vehicle pose at the time of the switch; Based on the switching time and the corresponding vehicle pose, the scene reconstruction information at the switching time is queried from the preset scene model. The preset scene model is constructed based on the log data of the open-loop playback mode. The scene reconstruction information is analyzed to separate the first type of elements representing the static environment and the second type of elements representing the dynamic target; Based on the first type of elements, the description information of the static environment is reconstructed, and based on the second type of elements, the state information of at least one dynamic target at the switching moment is reconstructed.
[0149] In one embodiment, the processor, when executing a computer program, also performs the following steps: Based on the spatiotemporal correlation of the second type of elements in the preset neighborhood at the switching time, the second type of elements are clustered to obtain at least one set of elements, where each set of elements corresponds to a dynamic target. Based on each set of elements, the geometric information, appearance information, and motion state information of the corresponding dynamic target are determined to generate state information.
[0150] In one embodiment, the processor, when executing a computer program, also performs the following steps: Obtain log data in open-loop playback mode; the log data should include at least time-synchronized real-scene image sequences, 3D point cloud sequences, and corresponding vehicle pose sequences; Based on log data, a 3D scene representation model is constructed; the 3D scene representation model is used to perform unified spatiotemporal coding of static environment and dynamic targets in real scene. Select the target time and corresponding vehicle pose from the log data, and render the corresponding reconstructed image from the 3D scene representation model based on the target time and corresponding vehicle pose. The parameters of the 3D scene representation model are optimized with the goal of minimizing the difference between the reconstructed image and the real scene image at the target time. The optimized 3D scene representation model is the preset scene model.
[0151] In one embodiment, the processor, when executing a computer program, also performs the following steps: A virtual static environment is constructed based on the description information of the static environment in the target scene information; Based on the state information of the dynamic target in the target scene information, determine the target 3D model of the dynamic target and the initial state of the dynamic target, wherein the initial state includes at least the initial pose, initial velocity and initial acceleration. In a virtual static environment, create entities that correspond to dynamic targets; Bind entities to the target 3D model and configure the initial state on the entities to construct a virtual test scenario.
[0152] In one embodiment, the processor, when executing a computer program, also performs the following steps: Bind the entity to a preset behavior model and control the behavior model to enter the first stage; in the first stage, the behavior model is configured to control the entity to maintain its initial state; In response to the fulfillment of preset triggering conditions, the control behavior model moves from the first stage to the second stage; in the second stage, the behavior model is configured to perform autonomous decision-making and interactive behaviors based on simulation data.
[0153] In one embodiment, the processor, when executing a computer program, also performs the following steps: Obtain the simulation data generated at the current moment, and obtain the historical playback data at the switching moment of the open-loop playback mode; Determine the fusion weight corresponding to the current moment, and fuse the historical replay data and the simulation data generated at the current moment based on the fusion weight to obtain fused data; The fused data is provided to the autonomous driving system.
[0154] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: In response to the command to switch from open-loop playback mode to closed-loop simulation mode, determine the switching time; Obtain the target scene information corresponding to the switching time in open-loop playback mode; the target scene information includes at least the description information of the static environment and the state information of at least one dynamic target at the switching time; Based on the target scenario information, a virtual test scenario with a closed-loop simulation mode is constructed; Based on the virtual test scenario, simulation data is generated and provided to the autonomous driving system. In response to the control commands fed back by the autonomous driving system based on the simulation data, the virtual test scenario is updated to complete the switch to closed-loop simulation mode.
[0155] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: Obtain the vehicle pose at the time of the switch; Based on the switching time and the corresponding vehicle pose, the scene reconstruction information at the switching time is queried from the preset scene model. The preset scene model is constructed based on the log data of the open-loop playback mode. The scene reconstruction information is analyzed to separate the first type of elements representing the static environment and the second type of elements representing the dynamic target; Based on the first type of elements, the description information of the static environment is reconstructed, and based on the second type of elements, the state information of at least one dynamic target at the switching moment is reconstructed.
[0156] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: Based on the spatiotemporal correlation of the second type of elements in the preset neighborhood at the switching time, the second type of elements are clustered to obtain at least one set of elements, where each set of elements corresponds to a dynamic target. Based on each set of elements, the geometric information, appearance information, and motion state information of the corresponding dynamic target are determined to generate state information.
[0157] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: Obtain log data in open-loop playback mode; the log data should include at least time-synchronized real-scene image sequences, 3D point cloud sequences, and corresponding vehicle pose sequences; Based on log data, a 3D scene representation model is constructed; the 3D scene representation model is used to perform unified spatiotemporal coding of static environment and dynamic targets in real scene. Select the target time and corresponding vehicle pose from the log data, and render the corresponding reconstructed image from the 3D scene representation model based on the target time and corresponding vehicle pose. The parameters of the 3D scene representation model are optimized with the goal of minimizing the difference between the reconstructed image and the real scene image at the target time. The optimized 3D scene representation model is the preset scene model.
[0158] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: A virtual static environment is constructed based on the description information of the static environment in the target scene information; Based on the state information of the dynamic target in the target scene information, determine the target 3D model of the dynamic target and the initial state of the dynamic target, wherein the initial state includes at least the initial pose, initial velocity and initial acceleration. In a virtual static environment, create entities that correspond to dynamic targets; Bind entities to the target 3D model and configure the initial state on the entities to construct a virtual test scenario.
[0159] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: Bind the entity to a preset behavior model and control the behavior model to enter the first stage; in the first stage, the behavior model is configured to control the entity to maintain its initial state; In response to the fulfillment of preset triggering conditions, the control behavior model moves from the first stage to the second stage; in the second stage, the behavior model is configured to perform autonomous decision-making and interactive behaviors based on simulation data.
[0160] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: Obtain the simulation data generated at the current moment, and obtain the historical playback data at the switching moment of the open-loop playback mode; Determine the fusion weight corresponding to the current moment, and fuse the historical replay data and the simulation data generated at the current moment based on the fusion weight to obtain fused data; The fused data is provided to the autonomous driving system.
[0161] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0162] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0163] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for switching autonomous driving simulation modes, characterized in that, The method includes: In response to the command to switch from open-loop playback mode to closed-loop simulation mode, determine the switching time; Obtain the target scene information corresponding to the open-loop playback mode at the switching time; the target scene information includes at least the description information of the static environment and the state information of at least one dynamic target at the switching time; Based on the target scenario information, a virtual test scenario for the closed-loop simulation mode is constructed; Based on the virtual test scenario, simulation data is generated and provided to the autonomous driving system. In response to the control commands fed back by the autonomous driving system based on the simulation data, the virtual test scenario is updated to complete the switch to the closed-loop simulation mode.
2. The method according to claim 1, characterized in that, The step of obtaining the target scene information corresponding to the open-loop playback mode at the switching time includes: Obtain the vehicle pose corresponding to the switching time; Based on the switching time and the corresponding vehicle pose, the scene reconstruction information of the switching time is queried from the preset scene model, wherein the preset scene model is constructed based on the log data of the open-loop playback mode; The scene reconstruction information is analyzed to separate the first type of elements representing the static environment and the second type of elements representing the dynamic target; The description information of the static environment is reconstructed based on the first type of elements, and the state information of at least one dynamic target at the switching moment is reconstructed based on the second type of elements.
3. The method according to claim 2, characterized in that, The process of reconstructing the state information of at least one dynamic target at the switching moment based on the second type of elements includes: Based on the spatiotemporal correlation of the second type of elements in a preset neighborhood at the switching time, the second type of elements are clustered to obtain at least one set of elements, wherein each set of elements corresponds to a dynamic target; Based on each set of elements, the geometric information, appearance information, and motion state information of the corresponding dynamic target are determined to generate the state information.
4. The method according to claim 2, characterized in that, The method for constructing the preset scene model includes: Obtain the log data of the open-loop playback mode; the log data includes at least a time-synchronized real scene image sequence, a 3D point cloud sequence, and a corresponding vehicle pose sequence; Based on the log data, a three-dimensional scene representation model is constructed; the three-dimensional scene representation model is used to perform unified spatiotemporal coding of static environment and dynamic targets in real scene. Select the target time and corresponding vehicle pose from the log data, and render the corresponding reconstructed image from the three-dimensional scene representation model based on the target time and corresponding vehicle pose. The parameters of the 3D scene representation model are optimized with the goal of minimizing the difference between the reconstructed image and the real scene image corresponding to the target time. The optimized 3D scene representation model is the preset scene model.
5. The method according to claim 1, characterized in that, The construction of the virtual test scenario for the closed-loop simulation mode based on the target scenario information includes: Based on the description information of the static environment in the target scene information, a virtual static environment is constructed; Based on the state information of the dynamic target in the target scene information, the target three-dimensional model of the dynamic target and the initial state of the dynamic target are determined, wherein the initial state includes at least the initial pose, initial velocity and initial acceleration; In the virtual static environment, an entity corresponding to the dynamic target is created; The entity is bound to the target 3D model, and the initial state is configured on the entity to construct the virtual test scene.
6. The method according to claim 5, characterized in that, The step of configuring the initial state after the entity is defined further includes: The entity is bound to a preset behavior model, and the behavior model is controlled to enter a first stage; in the first stage, the behavior model is configured to control the entity to maintain the initial state; In response to the fulfillment of preset triggering conditions, the behavior model is controlled to enter the second stage from the first stage; in the second stage, the behavior model is configured to perform autonomous decision-making and interactive behaviors based on simulation data.
7. The method according to claim 1, characterized in that, The process of generating simulation data based on the virtual test scenario and providing it to the autonomous driving system includes: Obtain the simulation data generated at the current moment, and obtain the historical playback data of the open-loop playback mode at the switching moment; Determine the fusion weight corresponding to the current moment, and fuse the historical playback data and the simulation data generated at the current moment based on the fusion weight to obtain fused data; The fused data is provided to the autonomous driving system.
8. An autonomous driving simulation mode switching device, characterized in that, The device includes: The determination module is used to determine the switching time in response to the command to switch from open-loop playback mode to closed-loop simulation mode; The acquisition module is used to acquire the target scene information corresponding to the open-loop playback mode at the switching time; the target scene information includes at least the description information of the static environment and the state information of at least one dynamic target at the switching time; A construction module is used to construct a virtual test scenario for the closed-loop simulation mode based on the target scenario information; The switching module is used to generate simulation data based on the virtual test scenario and provide it to the autonomous driving system, and to update the virtual test scenario in response to the control command fed back by the autonomous driving system based on the simulation data, thereby completing the switch to the closed-loop simulation mode.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.