Video generation methods, apparatus, electronic devices and storage media
By acquiring single-frame environmental data and related information collected by the vehicle, and using a pre-trained model to generate future environment prediction videos, the problem of high training cost and insufficient realism of traditional autonomous driving models is solved, and low-cost, high-fidelity and diversified simulation video generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU XIAOPENG MOTORS TECH CO LTD
- Filing Date
- 2026-03-20
- Publication Date
- 2026-05-26
AI Technical Summary
Traditional autonomous driving model training relies on simulation-generated videos, which are costly and lack realism. Relying on real-collected videos is even more expensive.
By acquiring single-frame environmental data collected by the vehicle, and combining the driving data of the vehicle and other vehicles, road topology information and scene description information, feature extraction is performed, and a pre-trained video generation model is used to generate a video predicting the future environment.
It achieves low-cost, high-fidelity, and diverse autonomous driving simulation video generation, reduces data collection requirements, and the generated videos conform to physical laws.
Smart Images

Figure CN122093641A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video generation technology, specifically to a video generation method, apparatus, electronic device, and storage medium. Background Technology
[0002] Traditional autonomous driving model training relies either on training videos generated from autonomous driving simulations or on real-world driving videos. However, the approach of using simulation-generated training videos suffers from high costs and insufficient simulation realism, while the approach of using massive amounts of real-world driving videos, although providing sufficiently realistic footage, is even more expensive. Summary of the Invention
[0003] This application provides a video generation method, apparatus, electronic device, and storage medium that can generate diverse autonomous driving simulation videos in a controllable, high-fidelity manner that conforms to physical laws.
[0004] On one hand, embodiments of this application provide a video generation method, the method comprising: Acquire input data at the target time; wherein the input data includes at least one of the following: single-frame environmental data collected by any acquisition device of the first vehicle; the input data also includes at least one of the following: first driving data of the first vehicle; second driving data of the second vehicle contained in the single-frame environmental data; road topology information associated with the single-frame environmental data; and scene description information of the single-frame environmental data. Feature extraction is performed on the input data to obtain a feature vector; The feature vector is input into a pre-trained video generation model, which outputs a future environment prediction video after the target time. The video generation model is trained based on sample single-frame environment data and sample future environment prediction videos.
[0005] On the other hand, embodiments of this application provide a video generation apparatus, including: An acquisition module is used to acquire input data at a target time; wherein the input data includes at least one of the following: single-frame environmental data acquired by any acquisition device of the first vehicle; the input data also includes at least one of the following: first driving data of the first vehicle; second driving data of the second vehicle contained in the single-frame environmental data; road topology information associated with the single-frame environmental data; and scene description information of the single-frame environmental data. The feature extraction module is used to extract features from the input data to obtain feature vectors; The generation module is used to input the feature vector into a pre-trained video generation model and output a future environment prediction video after the target time. The video generation model is trained based on sample single-frame environment data and sample future environment prediction videos.
[0006] On the other hand, embodiments of this application provide an electronic device, which includes a processor and a memory. The memory stores a computer program, and the processor executes the video generation method as described in any of the above embodiments by calling the computer program stored in the memory.
[0007] On the other hand, embodiments of this application provide a computer-readable storage medium storing a computer program adapted for loading by a processor to execute the video generation method as described in any of the above embodiments.
[0008] The video generation method, electronic device, and storage medium provided in this application, in addition to acquiring single-frame environmental data collected by the vehicle to generate video based on the actual single-frame environmental data, can also introduce the driving data of the self-vehicle and other vehicles contained in the single-frame environmental data (i.e., the first driving data of the first vehicle and the second driving data of the second vehicle) to introduce traffic interaction between the self-vehicle and other vehicles. Furthermore, road topology information and scene description information can be introduced to further control the diversity of traffic scenarios, enabling the simulation of specific traffic scenarios. Since the first driving data, the second driving data, the road topology information, and the scene description information are all controllable, it is beneficial to generate more diverse and controllable autonomous driving simulation videos. Afterwards, feature vectors are obtained by extracting features from the input data. Then, the feature vectors are processed through a pre-trained video generation model to generate future autonomous driving simulation videos (i.e., future environment prediction videos).
[0009] Therefore, when generating a large number of simulation videos, this application only needs to collect single-frame environmental data at the target time to generate simulation videos containing multiple single-frame environmental data from the target time and subsequent times. If single-frame environmental data from multiple times are collected, the same number of simulation videos can be generated. The required single-frame environmental data is relatively small, and a large number of simulation videos can be generated using a small amount of real single-frame environmental data, greatly reducing data acquisition costs. Furthermore, this application can selectively introduce driving data and traffic scene information from both the user vehicle and other vehicles, enabling the controllable, high-fidelity generation of diverse autonomous driving simulation videos that conform to physical laws. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of an example video generation system provided in an embodiment of this application.
[0012] Figure 2 This is a first flowchart illustrating the video generation method provided in an embodiment of this application.
[0013] Figure 3 This is a second flowchart illustrating the video generation method provided in an embodiment of this application.
[0014] Figure 4 This is a schematic diagram of a first scenario for the video generation method provided in an embodiment of this application.
[0015] Figure 5 This is a schematic diagram of the third process of the video generation method provided in the embodiments of this application.
[0016] Figure 6 This is a schematic diagram of a second scenario for the video generation method provided in an embodiment of this application.
[0017] Figure 7 This is a schematic diagram of the structure of the video generation device provided in the embodiments of this application.
[0018] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] This application provides a video generation method, apparatus, storage medium, device, and program product. Specifically, the video generation method of this application can be executed by an electronic device, which can be a terminal or server, etc.
[0021] The terminal can be a vehicle, smartphone, tablet, laptop, smart TV, wearable smart device, smart vehicle terminal, etc.
[0022] A server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0023] For example, when this video generation method runs on a terminal device, the terminal device may include a display screen and a processor. The display screen is used to present an interface and receive user commands generated by the interface. The processor is used to store applications, run the applications, generate the interface, respond to commands, and control the display of the interface on the display screen. When the user operates the interface through the display screen, the interface can control the local content of the terminal device in response to the received operation commands. The terminal device can provide the graphical user interface to the user in various ways, such as rendering the display on the terminal device's screen or presenting the graphical user interface through holographic projection.
[0024] For example, when this video generation method runs on a server, it can be implemented and executed based on a cloud system. A cloud system refers to a program execution method based on cloud computing. A cloud system includes servers and client devices. The application's execution entity and interface presentation entity are separate; the storage and execution of the video generation method are completed on the server. Interface presentation is completed on the client, which is mainly used for data reception, transmission, and interface presentation. For example, the client can be a display device with data transmission capabilities located close to the user, such as a mobile terminal, television, computer, PDA, personal digital assistant, or head-mounted display device. However, the terminal device for data processing is the server in the cloud. When executing the video generation method, the user operates the client to send instructions to the server. The server processes the data according to the instructions, encodes and compresses the interface data, returns it to the client via the network, and finally, the client decodes and outputs the interface.
[0025] It should be noted that in this embodiment of the application, the execution subject of the video generation method can be a terminal device or a server, and this embodiment of the application does not limit the type of execution subject.
[0026] It is understood that in the specific implementation of this application, user object data, context data and other related data are involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0027] For example, in conjunction with the above description, Figure 1 This application illustrates a video generation system 1000 for implementing a video generation method, as provided in an embodiment of this application. The video generation system 1000 may include at least one terminal 1001, at least one server 1002, at least one database 1003, and a network. The terminal 1001 can be connected to different servers via the network. The terminal can be any device with computing hardware capable of supporting and executing software application tools corresponding to the program.
[0028] In the aforementioned video generation system 1000, terminal 1001 is used to install and run applications. In some cases, terminal 1001 may not need to have the application pre-installed, and users can directly access the application interface through a browser or other client. Users log in to the application using their registered accounts. When a user logs in, terminal 1001 sends a login request to server 1002. Server 1002 verifies the user's account, and if the verification is successful, server 1002 returns a login success notification to terminal 1001. During the execution of the video generation method, terminal 1001 and server 1002 interact with each other. Terminal 1001 sends various information to server 1002. Server 1002 determines the display data for terminal 1001 based on the stored program mechanism and the received information, and sends the display data back to terminal 1001 so that terminal 1001 can display the data sent by server 1002 to the user. For example, server 1002 can send the generated future environment prediction video to terminal 1001 for display.
[0029] In possible application scenarios, different terminals 1001 may be served by different servers 1002. Therefore, in order to distinguish the servers 1002 corresponding to different terminals 1001, the embodiments of this application will use the first and second methods for description.
[0030] Furthermore, when the video generation system 1000 includes multiple terminals, multiple servers, and multiple networks, different terminals can connect to each other through different networks and servers. The network can be a wireless network or a wired network; for example, wireless networks include Wi-Fi, LAN, cellular networks, 2G, 3G, 4G, and 5G networks. Additionally, different terminals can also connect to other terminals or servers using their own Bluetooth networks or hotspot networks. Moreover, the system 100 can include multiple databases coupled to different servers, and can continuously store game-related information in the databases while different users are playing multiplayer games online.
[0031] It should be noted that, Figure 1The schematic diagram of the video generation system shown is merely an example. The video generation system 1000 described in this application embodiment is intended to more clearly illustrate the technical solutions of this application embodiment and does not constitute a limitation on the technical solutions provided in this application embodiment. As those skilled in the art will know, with the evolution of video generation systems and the emergence of new business scenarios, the technical solutions provided in this application embodiment are also applicable to similar technical problems.
[0032] The technical solution of this application will be described in detail below through specific embodiments. It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0033] Please see Figure 2 It should be noted that the steps shown may be performed in a logical order different from that shown in the flowchart of this method. The video generation method of this application may include the following steps: Step 011: Obtain input data at the target time; wherein, the input data includes at least one of the following: single-frame environmental data collected by any acquisition device of the first vehicle, first driving data of the first vehicle, second driving data of the second vehicle contained in the single-frame environmental data, road topology information associated with the single-frame environmental data, and scene description information of the single-frame environmental data.
[0034] The input data at the target time can be input data acquired at any time. For example, the target time can be the current time. In this case, the input data includes at least a single frame of environmental data acquired by any acquisition device of the first vehicle (such as a camera, lidar, etc.) at the current time. The input data also includes the first driving data of the first vehicle acquired at the current time, the second driving data of the second vehicle contained in the single frame of environmental data, and the road topology information associated with the single frame of environmental data.
[0035] In one alternative embodiment, the first vehicle can upload the input data at the target time to the server, and the server can then obtain the input data at the target time.
[0036] In one optional embodiment, the input data at the target time also includes scene description information of the single-frame environmental data. The scene description information can be user-defined and edited, which limits the scene in the single-frame environmental data (such as weather, ambient temperature and humidity parameters). The user can also use the edited scene description information as part of the input data at the target time to jointly achieve subsequent video generation.
[0037] In this context, single-frame environmental data can be data representing the real-world scene collected by the vehicle's own acquisition devices during vehicle operation, after perceiving the surrounding real-world environment. For example, single-frame environmental data can be at least one of the following: scene images collected by each of the vehicle's individual cameras, point cloud data collected by the vehicle's LiDAR, and surround view images collected by a single camera. This yields realistic single-frame environmental data, providing a true data foundation for autonomous driving simulation and improving the realism of the simulation.
[0038] The first driving data of the first vehicle is used to characterize the state parameters (such as position, attitude, trajectory, etc.) of the first vehicle while it is driving on the road. It can be understood that the first driving data may include the target time and multiple frames of driving data collected by the first vehicle after the target time. Through the first driving data, the future movement of the first vehicle can be predicted, providing a real data basis for generating future videos.
[0039] The second driving data of the second vehicle included in the single-frame environmental data is used to characterize the state parameters (such as position, attitude, trajectory, etc.) of the second vehicle when it is driving on the road in the scene corresponding to the single-frame environmental data. It can be understood that the second driving data may include the driving data of the second vehicle collected by the first vehicle at the target time and multiple frames after the target time. Through the second driving data, the future movement of the second vehicle can be predicted, providing a real data basis for generating future videos.
[0040] Among them, the road topology information associated with single-frame environmental data is used to characterize the geometric layout, lane connectivity and driving rules of the road network in the scene corresponding to the single-frame environmental data, providing structured spatial constraints for driving scene generation and ensuring that the scene layout is realistic and reasonable.
[0041] It can be understood that road topology information can be road topology data representing the road network of a real scene, automatically generated after the vehicle's autonomous driving model collects single-frame environmental data, such as road topology vector maps and semantic segmentation maps. Alternatively, road topology information can also be high-precision map data obtained based on the vehicle's location at the target time. Or, road topology information can also be road description information input in natural language (such as "the road ahead changes from two lanes to three lanes, and the leftmost lane of the three lanes is a left-turn lane"). In this way, road topology information can be used to control road conditions in autonomous driving videos, which is beneficial for generating more diverse autonomous driving simulation videos.
[0042] The scene description information describes the scene corresponding to a single frame of environmental data to define the environmental parameters of the scene. For example, the scene description information can define environmental parameters such as weather, temperature, and humidity in traffic scenes, allowing for control of these environmental parameters and facilitating the generation of more diverse autonomous driving simulation videos.
[0043] In one optional embodiment, the first driving data of the first vehicle and the second driving data of the second vehicle can be preset driving parameters.
[0044] It is understandable that when generating autonomous driving simulation videos, actual driving data can be avoided. Instead, preset driving parameters can be used to limit the generation methods of the first and second driving data.
[0045] For example, preset driving parameters are used to define driving styles. Multiple driving styles are preset, each corresponding to a different preset driving parameter. By defining the driving style, when generating autonomous driving simulation videos, the model only needs to determine the initial positions of the first and second vehicles based on a single frame of environmental data. Then, based on the driving style defined by the preset driving parameters, it can predict the driving data of the first and second vehicles, so that the vehicle driving data matches the driving style.
[0046] This allows for control over vehicle driving data in the generated video, which is beneficial for generating more diverse autonomous driving simulation videos.
[0047] It should be noted that the preset driving parameters corresponding to the first driving data and the preset driving parameters corresponding to the second driving data can be the same or different, and no restriction is imposed here.
[0048] Step 012: Extract features from the input data to obtain feature vectors.
[0049] After obtaining the input data at the target time, since the input data may include data of different modalities and dimensions, it is necessary to extract features from the input data and unify the features of various input data so as to input them into the model for video generation.
[0050] Therefore, feature extraction can be performed on various types of input data to obtain feature vectors.
[0051] Please see Figure 3 and Figure 4 In one optional embodiment, step 012 includes: Step 0121: For each type of data in the input data, use the encoder corresponding to that type to extract features and obtain the feature vector corresponding to each type of data.
[0052] Since the input data contains data of different modalities or dimensions (collectively referred to as different types), in order to ensure the accuracy of feature extraction, an appropriate encoder can be set for each type of data to perform feature extraction.
[0053] For example, when the input data includes single-frame environment data, first driving data, second driving data, road topology information, and scene description information, the feature vector includes a first feature vector corresponding to the single-frame environment data, a second feature vector corresponding to the first driving data, a third feature vector corresponding to the road topology information, a fourth feature vector corresponding to the second driving data, and a fifth feature vector corresponding to the scene description information.
[0054] At this point, a variational autoencoder adapted to the single-frame environmental data (such as an image) is set up to extract features. The encoder of the variational autoencoder extracts features from the single-frame environmental data to obtain the corresponding first feature vector.
[0055] The Variational Autoencoder (VAE Encoder) maps scene images to the latent space through convolution and nonlinear transformation, and outputs the posterior distribution parameters (mean and log-variance) of latent variables. This enables the compression and probabilistic encoding of high-dimensional visual information into low-dimensional continuous latent representations, providing a compact and smooth feature foundation for subsequent temporal modeling and generative decoding.
[0056] In one alternative embodiment, the variational autoencoder can be a 3D VAE, which directly performs convolutional coding on the spatiotemporally integrated 3D video tensor (frame × height × width × channel); or the variational autoencoder can be a 2D VAE combined with a temporal attention module, first using the 2D VAE to spatially encode the single-frame environmental data (outputting single-frame latent features), and then using temporal attention (such as Transformer / Conv1D) to model inter-frame temporal dependencies; or, the variational autoencoder can be a 3D VAE combined with discretized representation coding (such as VQ-VAE / VQ-GAN), where the 3D VAE outputs continuous latent feature vectors, while VQ-VAE / VQ-GAN maps continuous features to a discrete codebook, outputting a discrete feature sequence.
[0057] For the first driving data, an adaptive vehicle encoder (EgoEncoder) is set up to process information such as position and attitude. The vehicle encoder can extract features from the first driving data to obtain the corresponding second feature vector.
[0058] The vehicle encoder is used to encode the position, attitude, speed, trajectory and other state information of the autonomous vehicle to generate vehicle conditional embeddings. This provides the model with accurate vehicle motion constraints, ensuring the consistency of the perspective and the authenticity and rationality of the driving trajectory in the generated driving scene.
[0059] For road topology information, a static encoder that adapts to "static information that does not change over time" is set up to extract features from the second driving data, thereby obtaining the third feature vector.
[0060] Static encoders are used to extract and embed features from fixed structural and environmental information (such as road topology, lane layout, background semantics, etc.) that do not evolve over time in a scene, generate static conditional features and inject them into the model to ensure the consistency and stability of the generated content in spatial structure.
[0061] For the second driving data, a special agent encoder is set up to extract features from intelligent agents that "can move, think, and have behavior" (such as the second vehicle, which belongs to this intelligent agent because it has a driver). This agent encoder can extract features from the second driving data to obtain the fourth feature vector.
[0062] The intelligent agent encoder is used to encode the category, position, posture, motion trajectory and interaction relationship of various traffic-participating intelligent agents (such as motor vehicles) in the scene, generate structured intelligent agent features, and inject them into the model as conditional guidance information to ensure the rationality, temporal consistency and physical authenticity of the dynamic target behavior generated in the driving scene.
[0063] A text encoder was set up to extract features from the scene description information, thereby generating a fifth feature vector.
[0064] Step 013: Input the feature vector into the pre-trained video generation model and output the future environment prediction video after the target time. The video generation model is trained based on the sample single-frame environment data and the sample future environment prediction video.
[0065] The video generation model is a model that generates video based on a series of information that limits and guides video generation. For example, the video generation model can be a DiT diffusion model or a U-Net-based diffusion model.
[0066] In one alternative embodiment, the video generation model can be trained based on sample single-frame environmental data and sample future environment prediction videos. For example, the sample single-frame environmental data is input into the initial model, and then the model parameters are adjusted by the sample future environment prediction videos. The initial model is trained until convergence, thereby obtaining the video generation model.
[0067] Among them, the DiT (Diffusion Transformer) diffusion model is a new generation image / video generation model that combines the diffusion model with the Transformer architecture. It can divide the input latent feature map into blocks and serialize it. It integrates time step and conditional information through adaptive layer normalization (AdaLN-Zero), uses the Transformer global self-attention mechanism to model long-range spatial dependencies, predicts and removes noise after multi-layer feature transformation, and outputs the final feature vector that integrates global structure and conditional semantics. It provides a highly scalable backbone architecture for high-resolution visual content generation.
[0068] Among them, the diffusion model based on U-Net is a diffusion denoising model based on a U-shaped network architecture. It adopts a symmetrical U-shaped convolutional architecture, realizes multi-scale feature extraction and fusion through downsampling and upsampling, and combines temporal step embedding and conditional information to perform stepwise denoising prediction of noise potential features. It has efficient local spatial modeling capabilities and is widely used in image and video generation tasks.
[0069] Alternatively, the core logic of the DiT diffusion model can be migrated to a diffusion model based on U-Net. The core of DiT is "Transformer sequence modeling + diffusion time step encoding". Migrating it to the U-Net diffusion model essentially preserves the overall structure of U-Net's "encoder-decoder + skip connections", replaces / enhances the convolutional modules of U-Net with DiT's Transformer attention modules, and adapts the time step encoding and position encoding logic of DiT into U-Net. This retains the advantages of U-Net in local feature extraction while incorporating DiT's global modeling capabilities.
[0070] After obtaining the feature vector, it can be input into the pre-trained video generation model to output a video predicting the future environment after the target time.
[0071] In one alternative embodiment, a future environment prediction video of a duration matching the user-inputted duration parameter can be generated. For example, the duration parameter could be 1 second, 10 seconds, 1 minute, etc., and is not limited thereto.
[0072] In one optional embodiment, the single-frame environmental data includes scene images captured by each of the vehicle's individual cameras, and the future environment prediction video includes the video corresponding to each scene image.
[0073] It is understandable that vehicles typically have cameras installed at different locations to capture scene images from different directions. Therefore, given that a single frame of environmental data includes scene images captured by each of the vehicle's cameras, the future environment prediction video needs to generate a corresponding video based on each scene image to obtain the future autonomous driving simulation video of the vehicle in each direction.
[0074] Please refer to it again. Figure 3 and Figure 4 In an optional embodiment, the video generation model 100 includes multiple cascaded feature processing modules 10, and step 013 includes: Step 0131: Perform feature processing operations on the first feature vector to the fifth feature vector through the first feature processing module 10 to output the intermediate feature vector; Step 0132: The intermediate feature vector is used as the first feature vector of the second feature processing module 10, and the feature processing operation is performed again through the second feature processing module 10 until the last feature processing module 10 completes the feature processing operation to output the target feature vector. Step 0133: The target feature vector is decoded by the decoder of the variational autoencoder to generate a future environment prediction video.
[0075] Specifically, the first feature processing module 10 performs feature processing operations on the first to fifth feature vectors, outputting an intermediate feature vector. This intermediate feature vector is then used as the first feature vector of the second feature processing module 10. The second feature processing module 10 then performs feature processing operations on the first to fifth feature vectors again, outputting another intermediate feature vector, which is then used as the first feature vector of the next feature processing module 10. This process is repeated until the last feature processing module 10 completes its feature processing operations on the first to fifth feature vectors, outputting the target feature vector. Finally, the target feature vector is decoded by the variational autoencoder decoder to generate a future environment prediction video.
[0076] It is understood that the video generation model 100 includes multiple serially connected feature processing modules 10. The output of the previous feature processing module 10 serves as the input of the next processing module, realizing the serial connection of multiple modules. Each layer of modules iteratively optimizes and refines the features, gradually merging the input data, so that the final features can more accurately capture the global structure and local details of the scene, thereby improving the simulation realism of the video generation.
[0077] In one alternative embodiment, the conventional DiT diffusion model generates text-to-image output by only receiving text input. To adapt to the input data of this application, the video generation model 100 of this application has been improved, and the structure of the video generation model 100 of this application is as follows: The video generation model 100 includes multiple cascaded feature processing modules 10. Each feature processing module 10 includes a first normalization unit 11, a self-attention unit 12, a second normalization unit 13, a cross-attention unit 14, a third normalization unit 15, and a feedforward neural network unit 16. The first normalization unit 11, the second normalization unit 13, and the third normalization unit 15 are used to perform normalization operations. The self-attention unit 12 is used to perform global feature association. The cross-attention unit 14 is used to perform conditional fusion of features. The feedforward neural network unit 16 is used to perform nonlinear transformations on the features.
[0078] In one alternative embodiment, the multiple cascaded feature processing modules 10 are identical.
[0079] Please refer to it again. Figure 4 In one optional embodiment, the feature processing operation includes: The first feature vector is processed sequentially through the first normalization unit 11, the self-attention unit 12, and the second normalization unit 13 to obtain the sixth feature vector. The sixth and fifth feature vectors are processed by the cross-attention unit 14 to obtain the seventh feature vector; The third normalization unit 15 and the feedforward neural network unit 16 process the seventh feature vector, the second feature vector, the third feature vector and the fourth feature vector in sequence to output the intermediate feature vector.
[0080] The processing procedures of each feature processing module 10 are basically the same, but the input first feature vectors differ. The following explanation uses the first feature processing module 10 performing feature processing operations on the first to fifth feature vectors as an example.
[0081] First, the first feature vector can be processed through the first normalization unit 11, the self-attention unit 12, and the second normalization unit 13 in one step to obtain the sixth feature vector. That is, the first feature vector is first normalized by the first normalization unit 11, then globally associated by the self-attention unit 12, and finally normalized again by the second normalization unit 13 to obtain the sixth feature vector.
[0082] Then, the fifth and sixth feature vectors corresponding to the scene description information are injected into the cross-attention unit 14 to conditionally fuse the features and obtain the seventh feature vector. After that, the second, third, and fourth feature vectors can be injected into the model, and at least one of the seventh, second, third, and fourth feature vectors is processed sequentially by the third normalization unit 15 and the feedforward neural network unit 16 to output the intermediate feature vector corresponding to the first feature processing module 10.
[0083] That is, after the third normalization unit 15 normalizes the seventh feature vector, the second feature vector, the third feature vector and the fourth feature vector, the feedforward neural network unit 16 performs nonlinear mapping on the normalized features to obtain the intermediate feature vector.
[0084] It is understandable that conventional DiT models typically only support a single or limited number of conditions (such as text or category). This application, however, injects five types of data in a hierarchical manner: single-frame environmental data, road topology, vehicle and other vehicle driving data, and scene description information. Scene description information undergoes global semantic constraints through a cross-attention unit 14. Road topology and vehicle / other vehicle driving data are injected into the model separately for structured scene constraints, simultaneously controlling "scene semantics + road structure + vehicle behavior + other vehicle interaction." The generated driving scene more closely resembles real-world physical rules, avoiding issues such as trajectory drift and road distortion. Furthermore, injecting different types of input data separately improves the model's scalability, allowing for the introduction of more diverse input data in the future, resulting in strong scalability.
[0085] Furthermore, conventional DiT models rely on global attention for modeling the temporal motion and spatial topology of driving scenarios, which can easily lead to the loss of local constraints. In contrast, the video generation model 100 of this application uses the first feature vector (corresponding to single-frame environmental data) as a temporal prior to ensure the continuity of motion between frames, uses road topology to fix the road network skeleton to avoid road deformation, and uses the driving data of the vehicle itself / other vehicles to accurately constrain the vehicle's trajectory, making the generated driving behavior more in line with traffic rules and physical common sense.
[0086] In the video generation method provided in this application embodiment, in addition to acquiring single-frame environmental data collected by the vehicle to generate video based on the actual single-frame environmental data, the driving data of the self-vehicle and other vehicles contained in the single-frame environmental data (i.e., the first driving data of the first vehicle and the second driving data of the second vehicle) can be introduced to introduce traffic interaction between the self-vehicle and other vehicles. Furthermore, road topology information and scene description information can be introduced to further control the diversity of traffic scenes, enabling the simulation of specific traffic scenarios. Since the first driving data, the second driving data, the road topology information, and the scene description information are all controllable, it is beneficial to generate more diverse and controllable autonomous driving simulation videos. Afterwards, feature vectors are obtained by extracting features from the input data. Then, the feature vectors are processed through a pre-trained video generation model to generate future autonomous driving simulation videos (i.e., future environment prediction videos).
[0087] Therefore, when generating a large number of simulation videos, this application only needs to collect single-frame environmental data at the target time to generate simulation videos containing multiple single-frame environmental data from the target time and subsequent times. If single-frame environmental data from multiple times are collected, the same number of simulation videos can be generated. The required single-frame environmental data is relatively small, and a large number of simulation videos can be generated using a small amount of real single-frame environmental data, greatly reducing data acquisition costs. Furthermore, this application can selectively introduce driving data and traffic scene information from both the user vehicle and other vehicles, enabling the controllable, high-fidelity generation of diverse autonomous driving simulation videos that conform to physical laws.
[0088] Please see Figure 5 and Figure 6 In some embodiments, the video generation model 100 further includes a timing processing unit 17, which includes a recurrent neural network or a state-space model. The method further includes: Step 014: Encode the historical single-frame environmental data collected before the target time to obtain the encoded feature vector; Step 015: The encoded feature vector and the first feature vector are fused in a time sequence by the time sequence processing unit 17 to output the updated first feature vector.
[0089] Among them, the timing processing unit 17 is a unit for processing the dependency relationship of multiple frames of single-frame environmental data that are sequential in time.
[0090] Recurrent neural networks (RNNs) are a type of neural network capable of remembering historical information and processing sequential data. By preserving hidden states within the network, they model temporal dependencies. RNNs transmit historical information through hidden states, naturally satisfying causality and making them suitable for modeling the temporal evolution of video frames.
[0091] State-space models are a class of models that efficiently model long sequences by dynamically evolving hidden states over time, simultaneously capturing historical dependencies and global trends. While maintaining linear complexity, state-space models achieve efficient long sequence modeling, capturing dependencies between distant frames in videos and naturally supporting causal reasoning, making them ideal for long video temporal tasks such as autonomous driving and continuous simulation.
[0092] Therefore, the historical single-frame environmental data collected before the target time is encoded to obtain an encoded feature vector, which is a vector representation in the same dimension as the first feature vector, facilitating the analysis of their temporal dependency. Then, the temporal processing unit 17 performs temporal fusion on the encoded feature vector and the first feature vector to obtain a fused feature vector, and updates the first feature vector based on the fused feature vector. For example, the fused feature vector is used as the updated first feature vector, thus making the first feature vector the current frame feature vector that incorporates historical temporal information.
[0093] Thus, by using the first feature vector of the current frame feature vector, which integrates historical time-series information, for subsequent video generation, the generated future environment prediction video can be more coherent, more realistic, and more in line with the laws of physical motion.
[0094] In some embodiments, the feature vector corresponding to a single frame of environmental data is updated based on the concatenated vector of the feature vector corresponding to the single frame of environmental data and the Gaussian noise.
[0095] Thus, by splicing Gaussian noise onto the feature vector corresponding to a single frame of environmental data, controllable random perturbations can be introduced to guide the model to gradually denoise and generate content under temporal constraints, while enhancing the diversity and generation stability of video content.
[0096] All of the above technical solutions can be combined in any way to form optional embodiments of this application, and will not be described in detail here.
[0097] To facilitate better implementation of the video generation method of this application, this application also provides a video generation apparatus. Please refer to... Figure 7 , Figure 7 This is a schematic diagram of the structure of a video generation apparatus provided in an embodiment of this application. The video generation apparatus 200 may include: The acquisition module 201 is used to acquire input data at a target time; wherein, the input data includes at least one of the following: single-frame environmental data acquired by any acquisition device of the first vehicle; the input data also includes at least one of the following: first driving data of the first vehicle; second driving data of the second vehicle contained in the single-frame environmental data; road topology information associated with the single-frame environmental data; and scene description information of the single-frame environmental data. Feature extraction module 202 is used to extract features from input data to obtain feature vectors; The generation module 203 is used to input the feature vector into the pre-trained video generation model and output the future environment prediction video after the target time. The video generation model is trained based on the sample single-frame environment data and the sample future environment prediction video.
[0098] Each module or unit in the aforementioned video generation device can be implemented entirely or partially through software, hardware, or a combination thereof. Each unit can be embedded in or independent of the processor in the electronic device in hardware form, or stored in the memory of the electronic device in software form, so that the processor can call and execute the corresponding operations of each unit.
[0099] The video generation device 200 can be integrated into a terminal or server that has storage and a processor and thus computing power, or the video generation device 200 can be the terminal or server.
[0100] Optionally, this application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0101] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may be a terminal or a server. Figure 8 As shown, the electronic device 300 includes a processor 301 with one or more processing cores, a memory 302 with one or more computer-readable storage media, and a computer program stored in the memory 302 and executable on the processor. The processor 301 and the memory 302 are electrically connected. Those skilled in the art will understand that the electronic device structure shown in the figures does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0102] The processor 301 is the control center of the electronic device 300. It connects various parts of the electronic device 300 through various interfaces and lines. By running or loading software programs and / or modules stored in the memory 302, and calling data stored in the memory 302, it executes various functions of the electronic device 300 and processes data, thereby performing overall processing of the electronic device 300.
[0103] Optional, such as Figure 8 As shown, the electronic device 300 also includes: a display screen 303, a radio frequency circuit 304, an audio circuit 305, an input unit 306, and a power supply 307. The processor 301 is electrically connected to the display screen 303, the radio frequency circuit 304, the audio circuit 305, the input unit 306, and the power supply 307. Those skilled in the art will understand that... Figure 8 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0104] The display screen 303 can be used to display a graphical user interface (GUI) and receive operation commands generated by the user interacting with the GUI. The display screen 303 may include a display panel and a touch panel. The display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the electronic device. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. The touch panel can be used to collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel), generate corresponding operation commands, and execute the corresponding program. Optionally, the touch panel may include a touch detection device and a touch controller. The touch detection device detects the user's touch location and the signal generated by the touch operation, and transmits the signal to the touch controller. The touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to the processor 301, and can receive and execute commands from the processor 301. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it transmits the information to the processor 301 to determine the type of touch event. Subsequently, the processor 301 provides corresponding visual output on the display panel according to the type of touch event. In this embodiment, the touch panel and the display panel can be integrated into the display screen 303 to achieve input and output functions. However, in some embodiments, the touch panel and the display screen 303 can be implemented as two independent components to achieve input and output functions. That is, the display screen 303 can also be used as part of the input unit 306 to achieve input functions.
[0105] The radio frequency circuit 304 can be used to transmit and receive radio frequency signals to establish wireless communication with network devices or other electronic devices, and to transmit and receive signals with network devices or other electronic devices.
[0106] Audio circuitry 305 can be used to provide an audio interface between a user and an electronic device via a speaker and a microphone. Audio circuitry 305 converts received audio data into electrical signals, transmits them to the speaker, and the speaker converts them into sound signals for output. Conversely, the microphone converts collected sound signals into electrical signals, which are then received by audio circuitry 305, converted back into audio data, and then processed by processor 301 before being transmitted via radio frequency circuitry 304 to, for example, another electronic device, or output to memory 302 for further processing. Audio circuitry 305 may also include an earphone jack to facilitate communication between peripheral headphones and electronic devices.
[0107] The input unit 306 can be used to receive input numbers, characters, or object feature information (such as fingerprints, irises, facial information, etc.), and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control.
[0108] Power supply 307 is used to supply power to various components of electronic device 300. Optionally, power supply 307 can be logically connected to processor 301 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. Power supply 307 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0109] although Figure 8 As not shown in the diagram, the electronic device 300 may also include a camera, sensor, wireless fidelity module, Bluetooth module, etc., which will not be described in detail here.
[0110] This application also provides a computer-readable storage medium for storing a computer program. This computer-readable storage medium can be applied to an electronic device, and the computer program causes the electronic device to execute the corresponding processes in the video generation method of the embodiments of this application; for the sake of brevity, these will not be elaborated further here.
[0111] This application also provides a computer program product including computer instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the corresponding processes in the video generation method of this application embodiment. For simplicity, further details are omitted here.
[0112] It should be understood that the processor in this application may be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor described above can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0113] It is understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0114] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0115] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0116] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0117] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0118] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0119] In addition, the functional units in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0120] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a personal computer or a server) to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0121] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A video generation method, characterized in that, include: Acquire input data at the target time; wherein the input data includes at least one of the following: single-frame environmental data collected by any acquisition device of the first vehicle; the input data also includes at least one of the following: first driving data of the first vehicle; second driving data of the second vehicle contained in the single-frame environmental data; road topology information associated with the single-frame environmental data; and scene description information of the single-frame environmental data. Feature extraction is performed on the input data to obtain a feature vector; The feature vector is input into a pre-trained video generation model, which outputs a future environment prediction video after the target time. The video generation model is trained based on sample single-frame environment data and sample future environment prediction videos.
2. The video generation method according to claim 1, characterized in that, The step of extracting features from the input data to obtain a feature vector includes: For each type of data in the input data, feature extraction is performed using the encoder corresponding to that type to obtain the feature vector corresponding to each type of data.
3. The video generation method according to claim 2, characterized in that, The video generation model includes a variational autoencoder, and one or more of a vehicle encoder, a static encoder, an agent encoder, and a text encoder; the single-frame environmental data corresponds to the variational autoencoder, the first driving data corresponds to the vehicle encoder, the road topology information corresponds to the static encoder, the second driving data corresponds to the agent encoder, and the scene description information corresponds to the text encoder.
4. The video generation method according to claim 3, characterized in that, The input data includes the single-frame environment data, the first driving data, the second driving data, the road topology information, and the scene description information; The step of extracting features from each type of data in the input data using the encoder corresponding to that type to obtain the feature vector corresponding to each type of data includes: The encoder of the variational autoencoder extracts features from the single-frame environmental data to obtain a first feature vector; The vehicle encoder extracts features from the first driving data to obtain a second feature vector; The road topology information is extracted using the static encoder to obtain a third feature vector. The intelligent encoder extracts features from the second driving data to obtain a fourth feature vector; The text encoder extracts features from the scene description information to obtain a fifth feature vector.
5. The video generation method according to claim 4, characterized in that, The video generation model includes multiple cascaded feature processing modules. The step of inputting the feature vectors into the pre-trained video generation model and outputting a predicted video of the future environment after the target time includes: The first feature processing module performs feature processing operations on the first to fifth feature vectors to output an intermediate feature vector. The intermediate feature vector is used as the first feature vector of the second feature processing module, and the feature processing operation is executed again through the second feature processing module until the last feature processing module completes the feature processing operation to output the target feature vector. The target feature vector is decoded by the decoder of the variational autoencoder to generate the future environment prediction video.
6. The video generation method according to claim 5, characterized in that, The feature processing module includes a first normalization unit, a self-attention unit, a second normalization unit, a cross-attention unit, a third normalization unit, and a feedforward neural network unit. The self-attention unit is used to perform global correlation of features, the cross-attention unit is used to perform conditional fusion of features, and the feedforward neural network unit is used to perform nonlinear transformation of features. The feature processing operations include: The first feature vector is processed sequentially through the first normalization unit, the self-attention unit, and the second normalization unit to obtain the sixth feature vector; The sixth feature vector and the fifth feature vector are processed by the cross-attention unit to obtain the seventh feature vector; The third normalization unit and the feedforward neural network unit sequentially process the seventh feature vector, the second feature vector, the third feature vector, and the fourth feature vector to output the intermediate feature vector.
7. The video generation method according to claim 4, characterized in that, The video generation model further includes a temporal processing unit, which includes a recurrent neural network or a state-space model. The method further includes: Encode the historical single-frame environmental data collected before the target time to obtain an encoded feature vector; The timing processing unit performs timing fusion on the encoded feature vector and the first feature vector to output the updated first feature vector.
8. The video generation method according to claim 1, characterized in that, The single-frame environmental data includes scene images captured by each of the vehicle's cameras, and the future environment prediction video includes videos corresponding to each of the scene images.
9. The video generation method according to claim 1, characterized in that, Before inputting the feature vector into a pre-trained video generation model and outputting a predicted video of the future environment after the target time, the method further includes: The feature vector corresponding to the single frame of environmental data is updated based on the concatenated vector of the feature vector and the Gaussian noise.
10. The video generation method according to claim 1, characterized in that, The single-frame environmental data includes at least one of the following: scene images captured by each camera of the vehicle, point cloud data captured by the vehicle's lidar, and surround view images captured by a single camera. The road topology information includes at least one of the semantic segmentation map generated by the first vehicle, high-precision map data, and road description information; The first driving data and the second driving data include at least one frame of driving data or preset driving parameters collected by the first vehicle.
11. The video generation method according to claim 1, characterized in that, The video generation model includes the DiT diffusion model or a U-Net-based diffusion model.
12. A video generation apparatus, characterized in that, include: An acquisition module is used to acquire input data at a target time; wherein the input data includes at least one of the following: single-frame environmental data acquired by any acquisition device of the first vehicle; the input data also includes at least one of the following: first driving data of the first vehicle; second driving data of the second vehicle contained in the single-frame environmental data; road topology information associated with the single-frame environmental data; and scene description information of the single-frame environmental data. The feature extraction module is used to extract features from the input data to obtain feature vectors; The generation module is used to input the feature vector into a pre-trained video generation model and output a future environment prediction video after the target time. The video generation model is trained based on sample single-frame environment data and sample future environment prediction videos.
13. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing a computer program, and the processor executing the video generation method according to any one of claims 1-11 by calling the computer program stored in the memory.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted for loading by a processor to perform the video generation method as described in any one of claims 1-11.