Simulation scene automatic generation method and device, electronic equipment and storage medium

By using video generation models and 3D Gaussian splashing algorithms to automatically generate simulation scenes in the autonomous driving simulation platform, the problem of low efficiency in manually creating scenes is solved, and rapid and efficient simulation scene generation and diversity support are achieved.

CN121785928APending Publication Date: 2026-04-03CONTINENTAL HOLDING CHINA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing autonomous driving simulation platforms rely on manually creating scenarios, resulting in low generation efficiency, especially when dynamically generating and adjusting complex scenarios in real time.

Method used

By inputting the intent commands of the simulation requirements into the video generation model, and using object detection algorithms and 3D Gaussian splashing algorithms, 3D models of dynamic and static objects are generated, and the simulation scene is automatically generated in conjunction with the simulation platform.

Benefits of technology

It enables the rapid generation of high-quality simulation scenarios based on user intent commands, reducing the requirements and costs of manual skills, improving the efficiency and diversity of scenario generation, and supporting the training and verification of autonomous driving algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121785928A_ABST
    Figure CN121785928A_ABST
Patent Text Reader

Abstract

The invention provides an automatic generation method and device of a simulation scene, electronic equipment and a storage medium. The specific implementation scheme is as follows: inputting an intention instruction corresponding to a simulation demand into a video generation model to obtain a target video; determining a dynamic object and a static object in the target video based on a target detection algorithm; determining multiple pieces of object pose information of different frames of the target video corresponding to the dynamic object; based on a 3D Gaussian splashing algorithm, determining a first 3D Gaussian sputtering model of the dynamic object and a second 3D Gaussian sputtering model of the static object; and based on a simulation platform, automatically generating a simulation scene corresponding to the target video according to the first 3D Gaussian sputtering model, the second 3D Gaussian sputtering model and the pose information of the plurality of objects. According to the scheme, the target video generated based on the intention instruction of the user can be converted into the corresponding simulation scene for automatic driving simulation, and the generation efficiency of the simulation scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of simulation technology, and in particular to methods, apparatus, electronic devices and storage media for automatically generating simulation scenes. Background Technology

[0002] With the rapid development of autonomous driving technology, the industry's demand for high-quality simulation scenarios has increased significantly. These scenarios are not only used to test and optimize autonomous driving algorithms, but also play a crucial role in verifying safety and reliability. Currently, many mainstream autonomous driving simulation platforms rely primarily on manually creating and configuring scenarios. While this approach is intuitive, it requires a significant investment of time, effort, and skill in modeling toolchains. Furthermore, it proves inefficient when faced with complex scenarios that require dynamic generation and real-time adjustments. Summary of the Invention

[0003] This disclosure provides a method, apparatus, electronic device, and storage medium for automatically generating simulation scenes to solve or alleviate one or more technical problems in the prior art.

[0004] On the one hand, this disclosure provides a method for automatically generating simulation scenes, including:

[0005] Input the intent command corresponding to the simulation requirements into the video generation model to obtain the target video;

[0006] Based on object detection algorithms, dynamic and static objects in target videos are identified.

[0007] Determine the pose information of multiple objects in different frames of the target video corresponding to the dynamic object;

[0008] Based on the 3D Gaussian splashing algorithm, the first 3D Gaussian splashing model of dynamic objects and the second 3D Gaussian splashing model of static objects are determined.

[0009] Based on the simulation platform, a simulation scene corresponding to the target video is automatically generated according to the first 3D Gaussian sputtering model, the second 3D Gaussian sputtering model, and the pose information of multiple objects.

[0010] On the other hand, this disclosure provides an automatic simulation scene generation device, including:

[0011] The video generation module is used to input the intent commands corresponding to the simulation requirements into the video generation model to obtain the target video;

[0012] The first determination module is used to determine dynamic and static objects in the target video based on the target detection algorithm;

[0013] The second determining module is used to determine the pose information of multiple objects in different frames of the target video corresponding to the dynamic object;

[0014] The third determination module is used to determine the first 3D Gaussian sputtering model of dynamic objects and the second 3D Gaussian sputtering model of static objects based on the 3D Gaussian sputtering algorithm.

[0015] The simulation module is used to automatically generate a simulation scene corresponding to the target video based on the simulation platform, the first 3D Gaussian sputtering model, the second 3D Gaussian sputtering model, and the pose information of multiple objects.

[0016] On the other hand, this disclosure provides an electronic device, including:

[0017] At least one processor; and

[0018] The memory is communicatively connected to the at least one processor; wherein,

[0019] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods of any embodiment of the present disclosure.

[0020] On the other hand, this disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to perform a method according to any embodiment of this disclosure.

[0021] On the other hand, this disclosure provides a computer program product including a computer program that, when executed by a processor, implements a method according to any embodiment of this disclosure.

[0022] According to the scheme disclosed herein, target videos generated based on user intent commands can be converted into corresponding simulation scenarios for autonomous driving simulation, thereby improving the efficiency of simulation scenario generation.

[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0024] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments provided according to this disclosure and should not be construed as limiting the scope of this disclosure.

[0025] Figure 1 This is a flowchart illustrating a method for automatically generating simulation scenes according to an embodiment of the present disclosure.

[0026] Figure 2 This is a schematic diagram of an automatic simulation scene generation device according to an embodiment of the present disclosure.

[0027] Figure 3 This is a block diagram of an electronic device used to implement the automatic generation method of simulation scenes according to embodiments of the present disclosure. Detailed Implementation

[0028] The present disclosure will now be described in further detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0029] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0030] like Figure 1 As shown, this disclosure provides a method for automatically generating simulation scenes, including:

[0031] Step S101: Input the intent command corresponding to the simulation requirements into the video generation model to obtain the target video.

[0032] Step S102: Based on the object detection algorithm, determine the dynamic and static objects in the target video.

[0033] Step S103: Determine the pose information of multiple objects in different frames of the target video corresponding to the dynamic object.

[0034] Step S104: Based on the 3D Gaussian sputtering algorithm, determine the first 3D Gaussian sputtering model of the dynamic object and the second 3D Gaussian sputtering model of the static object.

[0035] Step S105: Based on the simulation platform, automatically generate a simulation scene corresponding to the target video according to the first 3D Gaussian sputtering model, the second 3D Gaussian sputtering model, and the pose information of multiple objects.

[0036] According to the embodiments of this disclosure, it should be noted that:

[0037] Intent instructions can be generated based on at least one of text, voice, and images. Based on the intent instructions, the user can determine the object categories of dynamic objects, the object category information of static objects, the object position information of dynamic objects, the object motion information of dynamic objects, and the object position information of static objects that the target video needs to be generated. For example, an intent instruction could be: "On a busy street, the roadside is full of parked cars. Suddenly, a pedestrian crosses the parked cars from the crosswalk and walks onto the road. At this moment, a car is speeding across the road. Due to obstructed view, the pedestrian and the car collide."

[0038] The video generation model can employ existing technologies such as text-to-video, image-to-video, or text-and-image-to-video models; no specific limitations are imposed here. Examples include video generation models like Sora and Runway.

[0039] Dynamic objects can include motor vehicles, non-motor vehicles, and VRUs (Vulnerable Road Users).

[0040] Static objects can include roads, buildings, transportation facilities, etc.

[0041] Object pose information includes the object's position and rotation orientation. For example, when the dynamic object is a vehicle, its position and orientation change in each frame of the target video because the vehicle is moving. Similarly, when the dynamic object is a pedestrian, its position and orientation change in each frame of the target video because the pedestrian is moving.

[0042] The object detection algorithm can be any existing technology, as long as it can identify dynamic and static objects in the target video, and can identify the 3D bounding boxes of dynamic objects in different frames of the target video.

[0043] The first and second 3D Gaussian sputtering models can include information such as the object's geometry, texture features, and motion trajectory.

[0044] The simulation platform can be any existing autonomous driving simulation platform, and no specific limitation is made here.

[0045] According to the technology of this disclosure, target videos of desired driving scenarios are directly generated using intent commands corresponding to simulation requirements. These target videos are then used to generate simulation scenarios for vehicle simulation, enabling users to quickly create the required simulation scenarios based solely on text, voice, or images. This not only reduces the skill requirements for using the simulation platform and the workload of manual design and modeling, but also saves manpower and time costs, improving the efficiency of simulation scenario generation. Furthermore, since the target videos are generated based on user intent commands (voice, text, and images), different target videos with different video content can be generated freely without restriction. Because the generated target videos are unrestricted, the diversity of simulation scenarios obtained based on the target videos can be increased. This allows for the generation of various types of traffic simulation scenarios that meet the needs of autonomous driving simulation based on different intent commands. According to the method of this disclosure, the scenario described by the user based on intent commands can be accurately reproduced. The vehicles and pedestrians in this scenario are in the same positions as in the target video, possessing physical simulation characteristics. Simultaneously, vehicle and pedestrian state controllers are also generated, enabling the movement described in the intent commands. The vehicle can be connected to autonomous driving algorithms for training and validating various scenarios in autonomous driving.

[0046] In one implementation, step S101: inputting the intent command corresponding to the simulation requirement into the video generation model to obtain the target video, including:

[0047] Based on the simulation requirements, generate intent instructions containing text information, which includes at least the object category information, object location information, and object motion information of the object to be generated in the video.

[0048] Input the intent command into the video generation model to obtain the target video. Alternatively, input the intent command and an image associated with the simulation requirements into the video generation model to obtain the target video.

[0049] According to the technology of this disclosure, since the target video is generated based on the user's intention commands via voice, text, and images, target videos with different scenes and different video content can be generated freely and without restriction. Because the generated target videos are unrestricted, the scene diversity of simulation scenarios obtained based on the target videos can be improved. This enables the generation of various types of traffic simulation scenarios that meet the needs of autonomous driving simulation based on different intention commands.

[0050] In one implementation, step S104: determining a first 3D Gaussian sputtering model for a dynamic object and a second 3D Gaussian sputtering model for a static object based on a 3D Gaussian sputtering algorithm, including:

[0051] Step S1041: Use computer vision algorithms to determine the first camera pose information of the dynamic object and the second camera pose information of the static object.

[0052] Step S1042: Based on the 3D Gaussian splashing algorithm and the pose information of the first camera, determine the first 3D Gaussian splashing model of the dynamic object.

[0053] Step S1043: Based on the 3D Gaussian splashing algorithm and the pose information of the second camera, determine the second 3D Gaussian splashing model of the static object.

[0054] According to the technology of the embodiments of this disclosure, by using computer vision algorithms to determine the camera pose information corresponding to each dynamic object and static object in the target video, it can help to accurately simulate dynamic objects and static objects in the simulation platform.

[0055] In one implementation, step S1042: Based on the 3D Gaussian sputtering algorithm and the pose information of the first camera, determine the first 3D Gaussian sputtering model of the dynamic object, including:

[0056] When the dynamic object is a vulnerable road user, the human body parameter model of the vulnerable road user is determined by using a human pose recognition algorithm based on the pose information of multiple objects of the vulnerable road user.

[0057] Using a linear hybrid skinning algorithm, Gaussian binding is applied to the human body parameter model to obtain the binding results of vulnerable road users in different frames of the target video.

[0058] Based on the 3D Gaussian splashing algorithm, the first 3D Gaussian splashing model of vulnerable road users is determined according to the binding results and the pose information of the first camera.

[0059] According to the embodiments of this disclosure, it should be noted that:

[0060] Based on human body parameter models, joint information of vulnerable road users in videos can be obtained.

[0061] Gaussian binding can be understood as using a linear blending skinning algorithm to bind Gaussian spheres at different locations on the body of a vulnerable road user to the corresponding joints.

[0062] According to the technology of this disclosure, the dynamic deformation and motion of vulnerable road users in world space can be obtained. This enables the vulnerable road users to be controlled at the joint level when simulating them based on this first 3D Gaussian sputtering model and simulation platform, accurately simulating pedestrian walking, running, and other actions.

[0063] In one implementation, step S1042: Based on the 3D Gaussian sputtering algorithm and the pose information of the first camera, determine the first 3D Gaussian sputtering model of the dynamic object, including:

[0064] When the dynamic object is a vehicle, the first 3D Gaussian sputtering model of the vehicle is determined based on the 3D Gaussian sputtering algorithm and the pose information of multiple objects of the vehicle and the pose information of the first camera.

[0065] According to the embodiments of this disclosure, it should be noted that:

[0066] Since the vehicle is a rigid body, the object pose information and 3D bounding box of the vehicle in different frames of the target video can be obtained through rigid body transformation algorithms.

[0067] According to the technology of this disclosure, the global pose (object pose information) of a vehicle can be obtained over time, while the Gaussian sputtering within the vehicle remains unchanged in the local space. By applying a corresponding rigid body transformation to the Gaussian sputtering in the local space, a dynamic representation of the vehicle in world space and an accurate first 3D Gaussian sputtering model of the vehicle can be obtained.

[0068] In one embodiment, the automatic generation method for simulation scenes provided in this disclosure, after step S104: determining the first 3D Gaussian sputtering model of the dynamic object and the second 3D Gaussian sputtering model of the static object based on the 3D Gaussian sputtering algorithm, further includes:

[0069] Determine the first standard metric for dynamic objects based on their category.

[0070] Based on the category corresponding to static objects, determine the second standard scale for static objects.

[0071] Based on the first standard scale, the first 3D Gaussian sputtering model is preprocessed to obtain the scaled first 3D Gaussian sputtering model.

[0072] Based on the second standard scale, the second 3D Gaussian sputtering model is preprocessed to obtain the scaled second 3D Gaussian sputtering model.

[0073] According to the embodiments of this disclosure, it should be noted that:

[0074] The first standard scale can be understood as the typical scale of a dynamic object in the real world. For example, the width of a vehicle, the distance between its wheels, etc.

[0075] The second standard scale can be understood as the typical scale of a static object in the real world. For example, the standard width of a lane line is 3.5 meters.

[0076] Preprocessing the first 3D Gaussian sputtering model can be understood as adjusting the coordinates of the Gaussian sphere in the point cloud world of the dynamic object to obtain a reasonable first 3D Gaussian sputtering model of the dynamic object that conforms to the scale in the real world.

[0077] Preprocessing the second 3D Gaussian sputtering model can be understood as adjusting the coordinates of the Gaussian sphere in the point cloud world of the static object to obtain a reasonable second 3D Gaussian sputtering model of the dynamic object that conforms to the scale in the real world.

[0078] According to the technology of the present disclosure embodiments, by scaling processing, static and dynamic objects that are close to the real size can be obtained, which makes up for the problem that objects in the target video do not have absolute scale, thereby making the first 3D Gaussian sputtering model and the second 3D Gaussian sputtering model input to the simulation platform more accurate, so as to simulate static and dynamic objects that are more in line with the simulation requirements.

[0079] In one implementation, step S105: Based on the simulation platform, according to the first 3D Gaussian sputtering model, the second 3D Gaussian sputtering model, and multiple object pose information, automatically generate a simulation scene corresponding to the target video, including:

[0080] Step S1051: Use a mesh generation algorithm to convert the first 3D Gaussian sputtering model into first mesh data.

[0081] Step S1052: Use a mesh generation algorithm to convert the second 3D Gaussian sputtering model into second mesh data.

[0082] Step S1053: Based on the simulation platform, configure the corresponding simulation attributes for the first grid data and the second grid data.

[0083] Step S1054: Based on the first grid data, the second grid data, simulation attributes, and multiple object pose information, automatically generate a simulation scene corresponding to the target video using the simulation platform.

[0084] According to the embodiments of this disclosure, it should be noted that:

[0085] The first and second grid data can be understood as mesh format data used by the simulation platform.

[0086] Simulation attributes can include collision attributes (whether an object can be collided with), dynamic interaction attributes (whether the object possesses traffic element information, such as traffic lights), color attributes, speed attributes, driving control attributes, joint control attributes, behavioral strategy attributes, and path planning attributes. Specific simulation attributes are set according to the different usage requirements of the object in the simulation scenario.

[0087] When the simulation platform's database includes object models corresponding to static and / or dynamic objects, the relevant object models can be directly called, along with their pre-configured simulation attributes. For example, if the database contains object models of people with similar height and weight, then that person's object model can be called. Similarly, if the database contains object models of vehicles with similar appearances, then that vehicle's object model can be called, and simulation attributes such as color, character, and speed can be obtained, along with simulation attributes for automatically configuring related control behavior scripts, such as simulation attributes for driving control, path planning, and behavioral strategies.

[0088] According to the technology of the embodiments of this disclosure, through scene reconstruction technology and mesh generation algorithm, the static and dynamic objects in the generated simulation scene have high-precision geometric shapes and texture details, which can provide a high-quality simulation environment for the simulation scene.

[0089] In one implementation, step S1052: converting the second 3D Gaussian sputtering model into second mesh data using a mesh generation algorithm, including:

[0090] When the static object is a road, the corresponding bird's-eye view is rendered based on the road's second 3D Gaussian sputtering model.

[0091] Lane information is extracted from bird's-eye view based on lane line detection algorithm.

[0092] Based on lane information, the vectorized OpenDrive map data used to describe the road network is determined.

[0093] Using a grid generation algorithm and map data, the second 3D Gaussian sputtering model is converted into second grid data.

[0094] In one implementation, step S1054: Based on the first grid data, the second grid data, simulation attributes, and multiple object pose information, automatically generate a simulation scene corresponding to the target video using a simulation platform, including:

[0095] The motion trajectory of a dynamic object is determined based on the pose information of multiple objects.

[0096] The motion trajectory is smoothed to obtain the repaired motion trajectory.

[0097] Based on the first grid data, the second grid data, simulation attributes, pose information of multiple objects, and the repaired motion trajectory, the simulation platform automatically generates a simulation scene corresponding to the target video.

[0098] According to the embodiments of this disclosure, it should be noted that:

[0099] The method for smoothing the motion trajectory is not specifically limited here; Kalman filtering or optimization algorithms based on kinematic models can be used.

[0100] According to the technology of the embodiments of this disclosure, the problem of "clipping" or "teleportation" of dynamic objects such as vehicles and pedestrians in the generated target video can be effectively solved, ensuring that the movement of dynamic objects input into the simulation platform is logical, thereby making the movement of dynamic objects in the final generated simulation platform logical.

[0101] like Figure 2 As shown, this disclosure provides an automatic simulation scene generation device, including:

[0102] The video generation module 210 is used to input the intent command corresponding to the simulation requirements into the video generation model to obtain the target video.

[0103] The first determining module 220 is used to determine dynamic and static objects in the target video based on the target detection algorithm.

[0104] The second determining module 230 is used to determine the pose information of multiple objects in different frames of the target video corresponding to the dynamic object.

[0105] The third determining module 240 is used to determine the first 3D Gaussian sputtering model of a dynamic object and the second 3D Gaussian sputtering model of a static object based on the 3D Gaussian sputtering algorithm.

[0106] The simulation module 250 is used to automatically generate a simulation scene corresponding to the target video based on the simulation platform, according to the first 3D Gaussian sputtering model, the second 3D Gaussian sputtering model, and the pose information of multiple objects.

[0107] In one implementation, the video generation module 210 is used to:

[0108] Based on the simulation requirements, generate intent instructions containing text information, which includes at least the object category information, object location information, and object motion information of the object to be generated in the video.

[0109] Input the intent command into the video generation model to obtain the target video. Alternatively, input the intent command and an image associated with the simulation requirements into the video generation model to obtain the target video.

[0110] In one implementation, the third determining module 240 is used to:

[0111] Using computer vision algorithms, the first camera pose information of dynamic objects and the second camera pose information of static objects are determined.

[0112] Based on the 3D Gaussian sputtering algorithm and the pose information of the first camera, the first 3D Gaussian sputtering model of the dynamic object is determined.

[0113] Based on the 3D Gaussian sputtering algorithm and the pose information of the second camera, a second 3D Gaussian sputtering model of a static object is determined.

[0114] In one implementation, a first 3D Gaussian sputtering model of a dynamic object is determined based on a 3D Gaussian sputtering algorithm and the pose information of a first camera, including:

[0115] When the dynamic object is a vulnerable road user, the human body parameter model of the vulnerable road user is determined by using a human pose recognition algorithm based on the pose information of multiple objects of the vulnerable road user.

[0116] Using a linear hybrid skinning algorithm, Gaussian binding is applied to the human body parameter model to obtain the binding results of vulnerable road users in different frames of the target video.

[0117] Based on the 3D Gaussian splashing algorithm, the first 3D Gaussian splashing model of vulnerable road users is determined according to the binding results and the pose information of the first camera.

[0118] In one implementation, a first 3D Gaussian sputtering model of a dynamic object is determined based on a 3D Gaussian sputtering algorithm and the pose information of a first camera, including:

[0119] When the dynamic object is a vehicle, the first 3D Gaussian sputtering model of the vehicle is determined based on the 3D Gaussian sputtering algorithm and the pose information of multiple objects of the vehicle and the pose information of the first camera.

[0120] In one embodiment, the automatic simulation scene generation device provided in this disclosure further includes:

[0121] The fourth determination module is used to determine the first standard metric of a dynamic object based on its category.

[0122] The fifth determination module is used to determine the second standard scale of static objects based on the category corresponding to static objects.

[0123] The first preprocessing module is used to preprocess the first 3D Gaussian sputtering model according to the first standard scale to obtain the scaled first 3D Gaussian sputtering model.

[0124] The second preprocessing module is used to preprocess the second 3D Gaussian sputtering model according to the second standard scale to obtain the scaled second 3D Gaussian sputtering model.

[0125] In one implementation, the simulation module 250 is used for:

[0126] The first 3D Gaussian sputtering model is converted into first grid data using a mesh generation algorithm.

[0127] The second 3D Gaussian sputtering model is converted into second grid data using a mesh generation algorithm.

[0128] Based on the simulation platform, corresponding simulation attributes are configured for the first grid data and the second grid data.

[0129] Based on the first grid data, the second grid data, simulation attributes, and the pose information of multiple objects, the simulation platform automatically generates a simulation scene corresponding to the target video.

[0130] In one implementation, a mesh generation algorithm is used to convert the second 3D Gaussian sputtering model into second mesh data, including:

[0131] When the static object is a road, the corresponding bird's-eye view is rendered based on the road's second 3D Gaussian sputtering model.

[0132] Lane information is extracted from bird's-eye view based on lane line detection algorithm.

[0133] Based on lane information, lanes are determined to be vectorized map data used to describe the road network.

[0134] Using a grid generation algorithm and map data, the second 3D Gaussian sputtering model is converted into second grid data.

[0135] In one implementation, based on first grid data, second grid data, simulation attributes, and multiple object pose information, a simulation platform automatically generates a simulation scene corresponding to the target video, including:

[0136] The motion trajectory of a dynamic object is determined based on the pose information of multiple objects.

[0137] The motion trajectory is smoothed to obtain the repaired motion trajectory.

[0138] Based on the first grid data, the second grid data, simulation attributes, pose information of multiple objects, and the repaired motion trajectory, the simulation platform automatically generates a simulation scene corresponding to the target video.

[0139] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.

[0140] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0141] Figure 3This is a structural block diagram of an electronic device according to an embodiment of the present disclosure. Figure 3 As shown, the electronic device includes a memory 310 and a processor 320. The memory 310 stores a computer program that can run on the processor 320. The number of memories 310 and processors 320 can be one or more. The memory 310 can store one or more computer programs, which, when executed by the electronic device, cause the electronic device to perform the methods provided in the above-described method embodiments. The electronic device may also include a communication interface 330 for communicating with external devices and performing data exchange and transmission.

[0142] If the memory 310, processor 320, and communication interface 330 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0143] Optionally, in a specific implementation, if the memory 310, processor 320 and communication interface 330 are integrated on a single chip, the memory 310, processor 320 and communication interface 330 can communicate with each other through an internal interface.

[0144] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Machines (ARM) architecture.

[0145] Further, optionally, the aforementioned memory may include read-only memory and random access memory, and may also include non-volatile random access memory. The memory may be volatile or non-volatile, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. Many forms of RAM are available by way of example, but not limitation. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate Synchronous DRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct RAMBUS RAM (DR RAM).

[0146] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this disclosure is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line, DSL) or wireless (e.g., infrared, Bluetooth, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., Solid State Disks (SSDs)). It is worth noting that the computer-readable storage media mentioned in this disclosure can be non-volatile storage media; in other words, they can be non-transient storage media.

[0147] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0148] In the description of the embodiments of this disclosure, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0149] In the description of the embodiments disclosed herein, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone.

[0150] In the description of embodiments of this disclosure, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more.

[0151] The above description is merely an exemplary embodiment of this disclosure and is not intended to limit this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the protection scope of this disclosure.

Claims

1. A method for automatically generating simulation scenes, comprising: Input the intent command corresponding to the simulation requirements into the video generation model to obtain the target video; Based on the object detection algorithm, dynamic and static objects in the target video are determined; Determine the pose information of multiple objects corresponding to different frames of the target video for the dynamic object; Based on the 3D Gaussian sputtering algorithm, a first 3D Gaussian sputtering model of the dynamic object and a second 3D Gaussian sputtering model of the static object are determined. Based on the simulation platform, a simulation scene corresponding to the target video is automatically generated according to the first 3D Gaussian sputtering model, the second 3D Gaussian sputtering model, and the pose information of the multiple objects.

2. The method according to claim 1, wherein, The intent command corresponding to the simulation requirements is input into the video generation model to obtain the target video, including: Based on the simulation requirements, generate intent instructions containing text information, wherein the text information includes at least the object category information, object location information, and object motion information of the object to be generated in the video; The intent command is input into the video generation model to obtain the target video; or, the intent command and the image associated with the simulation requirement are input into the video generation model to obtain the target video.

3. The method according to claim 1, wherein, Based on the 3D Gaussian sputtering algorithm, the first 3D Gaussian sputtering model of the dynamic object and the second 3D Gaussian sputtering model of the static object are determined, including: Using computer vision algorithms, the first camera pose information of the dynamic object and the second camera pose information of the static object are determined; Based on the 3D Gaussian splashing algorithm and the pose information of the first camera, the first 3D Gaussian splashing model of the dynamic object is determined; Based on the 3D Gaussian splashing algorithm and the pose information of the second camera, a second 3D Gaussian splashing model of the static object is determined.

4. The method according to claim 3, wherein, Based on the 3D Gaussian sputtering algorithm and the pose information of the first camera, a first 3D Gaussian sputtering model of the dynamic object is determined, including: When the dynamic object is a vulnerable road user, the human body parameter model of the vulnerable road user is determined by using a human body pose recognition algorithm based on the multiple object pose information of the vulnerable road user. Using a linear hybrid skinning algorithm, Gaussian binding is performed on the human body parameter model to obtain the binding results of the vulnerable road user in different frames of the target video; Based on the 3D Gaussian splashing algorithm, the first 3D Gaussian splashing model of the vulnerable road user is determined according to the binding result and the pose information of the first camera.

5. The method according to claim 3, wherein, Based on the 3D Gaussian sputtering algorithm and the pose information of the first camera, a first 3D Gaussian sputtering model of the dynamic object is determined, including: In the case where the dynamic object is a vehicle, a first 3D Gaussian sputtering model of the vehicle is determined based on the 3D Gaussian sputtering algorithm, according to the pose information of multiple objects of the vehicle and the pose information of the first camera.

6. The method according to claim 1, wherein, After determining the first 3D Gaussian sputtering model of the dynamic object and the second 3D Gaussian sputtering model of the static object based on the 3D Gaussian sputtering algorithm, the method further includes: Based on the category of the dynamic object, determine the first standard scale of the dynamic object; Based on the category corresponding to the static object, determine the second standard scale of the static object; Based on the first standard scale, the first 3D Gaussian sputtering model is preprocessed to obtain the scaled first 3D Gaussian sputtering model. Based on the second standard scale, the second 3D Gaussian sputtering model is preprocessed to obtain the scaled second 3D Gaussian sputtering model.

7. The method according to any one of claims 1 to 6, wherein, Based on the simulation platform, and according to the first 3D Gaussian sputtering model, the second 3D Gaussian sputtering model, and the pose information of the multiple objects, a simulation scene corresponding to the target video is automatically generated, including: The first 3D Gaussian sputtering model is converted into first grid data using a mesh generation algorithm; Using the aforementioned mesh generation algorithm, the second 3D Gaussian sputtering model is converted into second mesh data; Based on the simulation platform, configure corresponding simulation attributes for the first grid data and the second grid data; Based on the first grid data, the second grid data, the simulation attributes, and the pose information of the multiple objects, the simulation platform automatically generates a simulation scene corresponding to the target video.

8. The method according to claim 7, wherein, Using the aforementioned mesh generation algorithm, the second 3D Gaussian sputtering model is converted into second mesh data, including: When the static object is a road, a corresponding bird's-eye view is rendered based on the second 3D Gaussian sputtering model of the road. Based on the lane line detection algorithm, lane information is extracted from the bird's-eye view; Based on the lane information, determine the vectorized map data used to describe the road network for the lanes; Using the grid generation algorithm and the map data, the second 3D Gaussian sputtering model is converted into second grid data.

9. The method according to claim 7, wherein, Based on the first grid data, the second grid data, the simulation attributes, and the pose information of the multiple objects, the simulation platform automatically generates a simulation scene corresponding to the target video, including: Based on the pose information of the multiple objects, determine the motion trajectory of the dynamic object; The motion trajectory is smoothed to obtain the repaired motion trajectory; Based on the first grid data, the second grid data, the simulation attributes, the pose information of the multiple objects, and the repaired motion trajectory, the simulation platform automatically generates a simulation scene corresponding to the target video.

10. An automatic simulation scene generation device, comprising: The video generation module is used to input the intent commands corresponding to the simulation requirements into the video generation model to obtain the target video; The first determining module is used to determine dynamic and static objects in the target video based on the target detection algorithm; The second determining module is used to determine the pose information of multiple objects corresponding to different frames of the target video of the dynamic object; The third determining module is used to determine the first 3D Gaussian sputtering model of the dynamic object and the second 3D Gaussian sputtering model of the static object based on the 3D Gaussian sputtering algorithm. The simulation module is used to automatically generate a simulation scene corresponding to the target video based on the simulation platform, according to the first 3D Gaussian sputtering model, the second 3D Gaussian sputtering model, and the pose information of the multiple objects.

11. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 9.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 9.

13. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 9.