Decision-making model optimization methods, equipment, media, and products based on world models
By employing a closed-loop optimization framework for training the world model and decision model in two stages, the environmental prediction and generation quality of the world model is improved, solving the problem of poor environmental modeling performance in existing technologies and enhancing the safety and reliability of autonomous driving decision-making.
Patent Information
- Application Number
- CN202511241537.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-02
AI Technical Summary
Existing world models are not effective in environmental modeling in the field of autonomous driving, resulting in limited performance improvement of autonomous driving decision-making models based on world models.
A two-stage training method is adopted for the world model. First, the model is trained to understand structured traffic conditions, which improves the world model's ability to understand complex traffic scenarios. Then, based on the structured traffic conditions and driving actions, the model predicts future driving scenarios. A closed-loop optimization framework is constructed by combining the decision model with the model, and the decision model is updated through reward values.
It significantly improves the quality of environmental prediction and generation in world models, enhances the safety and reliability of autonomous driving decisions, and improves vehicle performance in complex traffic environments.
Smart Images

Figure CN120735801B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving technology, and in particular to methods, devices, media, and products for optimizing decision models based on world models. Background Technology
[0002] Based on the environmental understanding and interaction capabilities of world models, applying world models to assist autonomous driving decision-making has become one of the means to achieve autonomous driving. However, most current applications of world models in autonomous driving involve using them to generate future driving scenarios for autonomous driving decision-making models to assist them in making decisions. Furthermore, world models are not very effective at modeling real-world environments, resulting in limited performance improvements for autonomous driving decision-making models based on world models.
[0003] Improving the decision-making performance of autonomous driving decision-making models based on world models is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention
[0004] This invention provides a method, device, medium, and product for optimizing decision-making models based on world models, in order to at least solve the problem of poor decision-making performance of autonomous driving decision-making models based on world models in related technologies.
[0005] This invention provides a decision model optimization method based on a world model, comprising:
[0006] Obtain driving video samples;
[0007] The initial world model is trained using the driving video samples, and image frames corresponding to the time are generated based on the input structured traffic conditions to obtain the first world model;
[0008] The first world model is trained using the driving video samples to generate a sequence of future driving scenarios based on the input image sequence and the structured traffic conditions and driving actions at the corresponding time, thus obtaining the second world model.
[0009] The decision model is iteratively trained using the second world model. During training, the driving actions output by the decision model and their corresponding structured traffic conditions are input into the second world model. The reward value of the driving action is calculated based on the obtained future driving scenario sequence, and the decision model is updated based on the reward value.
[0010] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-described world-model-based decision model optimization methods when executing the computer program.
[0011] The present invention also provides a non-volatile storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described world model-based decision model optimization methods.
[0012] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described world model-based decision model optimization methods.
[0013] This invention employs a two-stage training method for the world model. The first stage trains the world model to understand structured traffic conditions, enhancing its ability to comprehend complex traffic scenarios. Based on this, the world model is further trained to predict future driving scenarios using structured traffic conditions and driving actions, improving the world model's environmental prediction and generation quality. A closed-loop optimization framework is collaboratively constructed based on the trained world model and the decision model. The driving actions output by the decision model and their corresponding structured traffic conditions are input into the pre-trained second world model. The reward value for the driving actions is calculated based on the obtained future driving scenario sequence, and the decision model is updated according to the reward value. This achieves efficient closed-loop optimization of the decision model, significantly improving the model's perception and decision-making capabilities, enhancing the safety and reliability of autonomous driving decisions, and improving the autonomous driving performance of vehicles in complex traffic environments. Attached Figure Description
[0014] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 A flowchart illustrating a decision model optimization method based on a world model, provided as an embodiment of the present invention;
[0016] Figure 2 A schematic diagram of the structure of an autonomous driving decision-making system based on a world model, provided for an embodiment of the present invention;
[0017] Figure 3 This is a schematic diagram illustrating the process of collecting a driving dataset according to an embodiment of the present invention;
[0018] Figure 4 This is a schematic diagram of the structure of a denoising network provided in an embodiment of the present invention;
[0019] Figure 5 This is a schematic diagram of the structure of a downsampling module of a denoising network provided in an embodiment of the present invention;
[0020] Figure 6 A schematic diagram of the structure of an intermediate coding module of a denoising network provided in an embodiment of the present invention;
[0021] Figure 7 This is a schematic diagram of the structure of an upsampling module of a denoising network provided in an embodiment of the present invention;
[0022] Figure 8 This is a schematic diagram of the structure of an attention module in a denoising network provided in an embodiment of the present invention;
[0023] Figure 9 A schematic diagram illustrating the generation and training process of a world model according to an embodiment of the present invention;
[0024] Figure 10 A schematic diagram illustrating the prediction training process of a world model provided in an embodiment of the present invention;
[0025] Figure 11 This is a schematic diagram of a closed-loop optimization process for a decision model based on a world model, provided as an embodiment of the present invention. Detailed Implementation
[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.
[0027] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0028] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0029] World models are attracting widespread attention in the field of intelligent decision-making due to their ability to understand the environment and its interactions. World models possess significant potential for generating high-quality driving videos and for end-to-end driving applications. In related technologies, the application of world model agents to autonomous driving technology primarily aims to predict future driving environment information through world models, thereby assisting autonomous driving decision-making. However, the modeling method of world models treats environmental dynamics as a series of discrete latent variables. While discretization of the latent space helps avoid compound errors over multiple time steps, this encoding may result in information loss, leading to a loss of generality and reconstruction quality. Therefore, world models suffer from several problems, including an inability to truly predict future driving environment information and an inability to support closed-loop optimization of autonomous driving decision-making models, resulting in limited performance improvements for autonomous driving decision-making tasks when applied to world models.
[0030] To improve the performance of autonomous driving decision-making based on world models, this invention provides a world model-based decision model optimization method, device, medium, and product. This method employs a two-stage training approach for the world model. The first stage trains the world model to understand structured traffic conditions, enhancing its ability to comprehend complex traffic scenarios. Based on this, the world model is further trained to predict future driving scenarios based on structured traffic conditions and driving actions, improving the environmental prediction and generation quality of the world model. A closed-loop optimization framework based on the trained world model and the decision model is collaboratively constructed. The driving actions output by the decision model and their corresponding structured traffic conditions are input into the pre-trained second world model. The reward value for the driving actions is calculated based on the obtained future driving scenario sequence, and the decision model is updated based on the reward value. This achieves efficient closed-loop optimization of the decision model, significantly improving the model's perception and decision-making capabilities, enhancing the safety and reliability of autonomous driving decisions, and improving the autonomous driving performance of vehicles in complex traffic environments.
[0031] The embodiments of the present invention provide a decision model optimization method based on a world model. The method is described in detail below, along with the execution flow of the decision model optimization method based on a world model.
[0032] Figure 1 A flowchart illustrating a decision model optimization method based on a world model, provided as an embodiment of the present invention; Figure 2 A schematic diagram of the structure of an autonomous driving decision-making system based on a world model, provided for an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the process of collecting a driving dataset according to an embodiment of the present invention.
[0033] like Figure 1 As shown, an embodiment of the present invention provides a decision model optimization method based on a world model, which may include: S101: acquiring driving video samples.
[0034] S102: Train an initial world model using driving video samples and generate image frames corresponding to the input structured traffic conditions to obtain the first world model.
[0035] S103: Train a first-world model using driving video samples. Generate a sequence of future driving scenarios based on the input image sequence and the structured traffic conditions and driving actions at the corresponding time, and obtain a second-world model.
[0036] S104: The decision model is iteratively trained using a second-world model. During training, the driving actions output by the decision model and their corresponding structured traffic conditions are input into the second-world model. The reward value of the driving action is calculated based on the obtained future driving scenario sequence, and the decision model is updated based on the reward value.
[0037] like Figure 2 As shown, the model training step in the decision model optimization method based on the world model provided in this embodiment of the invention can be applied to a model training device. The model training device can be one or more artificial intelligence servers, which can include basic settings such as artificial intelligence processors, storage, input / output devices, communication buses, and communication interfaces, and consists of software environments and application software such as operating systems, databases, and middleware.
[0038] In this embodiment of the invention, the environment is modeled as a standard partially observable Markov decision process. ,in It is a set of states. It is a collection of actions. It is a collection of environmental observations. State transition function. Describe environmental dynamics reward function Map state transitions to scalar rewards.
[0039] Because the intermediate process cannot directly access the state. Only through environmental observation values Perceiving environmental dynamics, these observations are based on observation probabilities. Emitted, described by the observation function The goal is to obtain a policy π that maps observations to actions to maximize expected revenue. ,in, It is a discount factor.
[0040] A world model is a generative model of the environment, that is... The model is used as a simulation environment to train the autonomous decision-making system in a sample-efficient manner.
[0041] Based on this, the decision model optimization process provided in this embodiment of the invention mainly consists of three steps: S101 is to collect driving data to construct a dataset; S102 and S103 are to train a world model on the driving dataset; and S104 is to optimize the decision model in the environment of the world model.
[0042] For S101, the driving video samples can be driving video samples obtained based on real driving scenarios. Figure 3 This demonstrates one method for obtaining driving video samples.
[0043] like Figure 3 As shown, a vehicle is used as the main intelligent agent, controlled by a human driver. During this process, driving data is collected using various onboard sensor modules and preprocessed, such as timestamp alignment, on an onboard edge server. Considering the limited computing and storage resources of the onboard edge server, the data is only temporarily cached there. After preprocessing, it is synchronously uploaded to the model training device to build a dataset for training the world model. Onboard sensor modules can include, but are not limited to, radar, cameras, onboard control units, inertial measurement units, and navigation systems. Radar and cameras are used to collect environmental perception data, and vehicle control signals are obtained from the onboard control unit. The inertial measurement unit and navigation system are used to collect vehicle motion states. The environmental perception data, vehicle motion states, and other data are input into the data preprocessing module of the onboard edge server. Multimodal data is timestamped, and structured traffic condition-related data (such as targets and lane lines) are detected. Multi-view images, road topology images, traffic target data, and vehicle control information are sent to the model training device and stored in the model training device's storage module to construct the driving dataset. The model training device receives manually labeled driving datasets or calls the automatic labeling module for automatic labeling to add scene description information. Using the world model update module of the model training device, the driving dataset is batch-sampled from the storage module, and the sampled multimodal data is encoded with conditional information and used for diffusion generation training.
[0044] Based on such Figure 3 The dataset collection process shown includes the following information for each driving video sample: ① Multi-view images: Multiple on-board cameras can be used to acquire multi-view images of the vehicle in real time during its movement. The acquired images can be in Red, Green, Blue, or RGB format. For example, using six on-board cameras can correspond to six different viewpoints of the vehicle: front (F), front left (FL), front right (FR), rear (B), rear left (BL), and rear right (BR). Each image is... The image is stored in the form of a graph, where H and W represent the height and width of the image, respectively. In addition, corresponding viewpoint and time step markers are included along with the stored image.
[0045] ② Road topology images: Road topology information, including lane boundaries, lane dividers, and pedestrian crossings, is extracted from multi-view images through manual annotation or related methods. Each image is... The images are stored in a format that is aligned with time step markers and multi-view images.
[0046] ③ Traffic target data: Road target and category information, including pedestrians, vehicles, bicycles, animals, etc., are extracted from multi-view images through manual annotation or related methods. Each target is labeled with a 3D bounding box and projected onto the image plane. Stored in the form of, where For each target category, each image is aligned with the multi-view images according to the time step marker.
[0047] ④ Vehicle control information: The vehicle control actions at each time step are recorded by the onboard controller, including two control quantities: steering angle and speed. The images are stored in a format that is aligned with the multi-view images according to time step markers.
[0048] ⑤ Scene description information: Add scene description information, including weather, time, etc., to the multi-view images through manual annotation or related methods. Store the information in text form and align it with the multi-view images according to the time step mark.
[0049] During the data acquisition phase, image data and other data can be acquired and processed at a frequency of 12Hz to obtain an image dataset. Further, multiple image data samples are aggregated into video data within a certain period of time according to a set number of frames, thereby dividing the image dataset into several video segment samples to obtain a driving video dataset. Each driving video sample contains a series of time-step aligned multi-view images, road topology images, traffic target data, vehicle control data, and scene description text data.
[0050] For S102 and S103, this embodiment of the invention trains the world model in two stages. The first stage is generative training, focusing on learning single-frame structured traffic conditions, training the world model to generate image frames of the driving scene at the corresponding time based on the input structured traffic conditions. Building on the training in the first stage, the second stage combines driving actions to train the world model to predict future driving scene information based on the input structured traffic conditions and driving actions. This gradually guides the world model to understand the impact of autonomous driving perception data and driving actions on the driving environment, improving the accuracy of the world model in generating future driving scene information, and providing a foundation for achieving closed-loop optimization of the decision model.
[0051] In this embodiment of the invention, the generation unit of the world model can be a diffusion model, a vector quantized variational autoencoder (VQ-VAE), or an autoregressor based on a Transformer-based generative model. If a variational autoencoder is used as the generation unit, continuous environment states can be compressed into vectors in a discrete codebook, with latent variables discretized through quantization. If an autoregressor based on a Transformer-based generative model is used as the generation unit, environment states (such as image frames) can be segmented into discrete token sequences, and future tokens can be predicted through autoregression.
[0052] For S104, such as Figure 2 As shown, during the training of the decision model, the decision model interacts with the second-world model in the model training device to conduct iterative training. The decision model makes decisions based on the prediction results output by the second-world model and optimizes the decision strategy based on the cumulative reward.
[0053] In the iterative training described in the various embodiments of the present invention, the iteration termination condition can be reaching a preset number of iterations or satisfying a preset convergence condition (loss value is less than a preset loss).
[0054] like Figure 2 As shown, the trained decision model and the feature encoding module of the second-world model are deployed to the vehicle edge server. During the actual operation of the autonomous driving system, the vehicle interacts with the real environment in real time, collecting environmental perception data from the real environment through the vehicle sensor module and transmitting the perception data to the vehicle edge server for processing via the data bus; at the same time, the vehicle receives action commands from the vehicle edge server through the Controller Area Network (CAN) bus and hands them over to the vehicle control drive module to control the vehicle to perform the corresponding driving actions.
[0055] After receiving the perception data sent by the vehicle, the vehicle edge server first performs preprocessing such as format conversion on the raw sensing data through the data preprocessing module. The resulting multimodal data is then sent to the feature encoding module to be processed into low-dimensional latent state features, which are then input into the decision model for action reasoning and output driving actions.
[0056] The world model-based decision model optimization method provided in this invention employs a two-stage training approach. The first stage trains the world model to understand structured traffic conditions, enhancing its ability to comprehend complex traffic scenarios. Based on this, the world model is further trained to predict future driving scenarios based on structured traffic conditions and driving actions, improving the world model's environmental prediction and generation quality. A closed-loop optimization framework is collaboratively constructed based on the trained world model and the decision model. The driving actions output by the decision model and their corresponding structured traffic conditions are input into the pre-trained second world model. The reward value for the driving actions is calculated based on the obtained future driving scenario sequence, and the decision model is updated according to the reward value. This achieves efficient closed-loop optimization of the decision model, significantly improving its perception and decision-making capabilities. Finally, an end-to-end autonomous decision-making system is constructed using the world model's feature encoding module for vehicle deployment, promoting the practical application of autonomous decision-making technology and improving vehicle performance in complex traffic environments.
[0057] Based on the above embodiments, in this embodiment of the invention, in S102, the initial world model is trained using driving video samples to generate image frames corresponding to the input structured traffic conditions, thus obtaining a first world model. This process may include extracting first image features from the first image frame in the driving video samples; encoding the structured traffic conditions corresponding to the first image frame to obtain a first conditional encoding result; generating a first generated image using the initial world model based on the first image features and the first conditional encoding result; calculating a first loss value between the first generated image and the first image frame; and updating the initial world model using the first loss value until the first iteration termination condition is met, thus obtaining the first world model.
[0058] In some optional embodiments of the present invention, the generation unit of the initial world model can be a diffusion model. In this case, generating the first generated image using the initial world model based on the first image features and the first conditional coding result may include: adding noise to the first image features to obtain the first noisy features; and performing denoising calculations using the initial world model based on the first image features and the first conditional coding result to obtain the first generated image.
[0059] To enable the world model to process multimodal data and accurately generate information about future driving scenarios, this invention provides a diffusion model structure.
[0060] Assume the data distribution is , To represent the latent features of an image, let This means adding variance to the input data. Given a sufficiently large noise variance, the data distribution after independent and identically distributed Gaussian noise is... ,have The diffusion model addresses high-variance noise. Initially, the noisy data was gradually denoised until... Its iterative optimization process can be expressed as the following probability flow differential equation:
[0061] ;
[0062] in, Let the scoring function be used. Training the diffusion model simplifies to using the scoring function. Training a model The model is parameterized in the following form:
[0063] ;
[0064] It is a tool for predicting the original The learnable denoiser is trained by matching the target with the following denoising scores:
[0065] ;
[0066] in, , Represents the probability distribution across noise levels. It is a weighting function. It is a conditional signal, including road topology features. Target conditions and text prompt embedding .
[0067] This invention employs a conditional preprocessing framework to denoise the device. The parameterization is as follows:
[0068] ;
[0069] in, It is the network to be trained. and A pre-conditioner used to adjust the network input and output at arbitrary noise levels. Maintaining unit variance can be expressed as:
[0070] , .
[0071] This is an empirical transformation of the noise level, which can be expressed as:
[0072] .
[0073] Is with The relevant noise figure can be expressed as:
[0074] .
[0075] This invention provides a condition processing module for driving video datasets. Each driving video sample can be represented as , including continuous The frame images are processed by an image encoder based on a variational autoencoder (VAE) to obtain the latent representation of each image frame in the video clip. Given the diffusion time step Parameterized noise timetable Noisy samples are obtained through a forward diffusion process. diffusion model take over And optimize the trainable parameters in the model by minimizing the following denoising score matching objective. .
[0076] This invention employs the UNet framework as the backbone network of the diffusion model, using structured traffic information projected onto the image plane as conditional input to construct the diffusion model. The denoising network of the diffusion model used in this embodiment is described below.
[0077] Figure 4 This is a schematic diagram of a denoising network provided in an embodiment of the present invention. Figure 4 As shown, the embodiment of the present invention adopts a diffusion model backbone network based on the UNet framework, which mainly consists of a downsampling module, an intermediate encoding module, and an upsampling module.
[0078] Figure 5 This is a schematic diagram of the downsampling module of a denoising network provided in an embodiment of the present invention. The downsampling module is responsible for performing dimensionality reduction encoding on the input data, such as... Figure 5 As shown, the noisy features and the diffusion time step are encoded separately and then processed by a ResNet network block. The results, along with the location embedding and cue embedding, are input into an attention module to obtain latent features. These latent features serve as input to subsequent network modules and are also stored as intermediate results for feature calculation in the upsampling module. Figure 5 In the diagram, ×2 represents two identical network structures cascaded together, and ×3 represents three identical network structures cascaded together. Identical network structures mean that the computational operations performed are consistent, but the network parameters of network structures at different locations may be different.
[0079] Figure 6This is a schematic diagram of the intermediate encoding module of a denoising network provided in an embodiment of the present invention. The intermediate encoding module is responsible for further processing the data in a low-dimensional feature space. The input data for this part includes latent features from the downsampling module, as well as encoded time-step embeddings, position embeddings, and cue embeddings. Figure 6 As shown, latent features and temporal embeddings are processed through a ResNet network block. The results, along with the location embeddings and cue embeddings, are input into an attention module and then processed again through a ResNet network block to obtain the latent features of this part of the output.
[0080] Figure 7 This is a schematic diagram of the structure of an upsampling module in a denoising network provided in an embodiment of the present invention. The upsampling module is responsible for restoring the encoded features from the low-dimensional space to the original input state space. The input data for this part includes latent features from intermediate encoding modules, encoded time-step embeddings, position embeddings, and cue embeddings, as well as intermediate features from the downsampling module. Figure 6 As shown, latent features are concatenated with intermediate features. The concatenated latent features and temporal embeddings are then processed through a series of ResNet network blocks. The result, along with the location embedding and cue embedding, is input into an attention module to obtain the latent features. Finally, the latent features are further processed through a convolutional layer to obtain the final predicted features.
[0081] Figure 8 This is a schematic diagram of the attention module structure of a denoising network provided in an embodiment of the present invention. Figure 8 As shown, in the denoising network of the diffusion model in this embodiment of the invention, the attention modules used in each network module all employ a gated self-attention mechanism to encode the latent features at each level. With position embedding Put together:
[0082] ;
[0083] in, These are learnable parameters. This indicates a self-attention operation. This indicates a label selection operation that only considers visual labels. Right now Figure 8 The activation function in [the context].
[0084] Furthermore, the attention module employs cross-attention operations to enable feature interactions between multimodal input data, such as feature interactions between text input and visual signals. This allows text descriptions to influence driving scene attributes, such as weather and time, thereby addressing the problem of world models handling multimodal input data.
[0085] Figure 9 This is a schematic diagram illustrating the generation and training process of a world model, as provided in an embodiment of the present invention.
[0086] In the world model based on the diffusion model, the model only considers single-frame inputs from structured traffic conditions in the first stage of generative training. This stage focuses on learning traffic structure constraints, which helps to accelerate convergence.
[0087] In embodiments of the present invention, such as Figure 9 As shown, encoding the structured traffic conditions corresponding to the first image frame to obtain a first conditional encoding result may include: extracting a first background feature from the first road topology map corresponding to the first image frame; concatenating the first noisy feature and the first background feature to obtain a second noisy feature; encoding the first target information of the first target included in the first image frame to obtain a second conditional encoding result; spatially aligning the second conditional encoding result with the second noisy feature to obtain a first position embedding feature; and using the first background feature and the first position embedding feature as the first conditional encoding result. Then, using an initial world model to perform denoising calculations based on the first image features and the first conditional encoding result to obtain a first generated image may include: using an initial world model to perform denoising calculations based on the second noisy feature and the second conditional encoding result to obtain the first generated image.
[0088] In other words, the structured traffic conditions used in the embodiments of the present invention may include road topology conditions. and target conditions The road topology conditions can be road topology maps (HDMaps). Target conditions refer to relevant information about traffic behavior targets (such as pedestrians and vehicles) in the driving scenario, which can include target location conditions and target category conditions. Road topology conditions and target conditions can be manually labeled information or obtained using perception methods (such as the Latent Action Value (LAV) model, the Bird's-eye View Annotation Model (BEVerse), and the Unified Autonomous Driving (UniAD) model). Road topology conditions can be three-lane (including lane boundaries, lane dividers, and pedestrian crossings), and target location conditions can be 3D bounding boxes. The three-lane road topology map and the 3D bounding boxes are projected onto the image plane to generate the corresponding conditions.
[0089] To reduce the complexity of condition processing, the number of targets in the scene at each time step can be limited, and the target conditions can be broken down into target (3D bounding box) position conditions. and target category conditions ,in It is a predefined maximum number of detection boxes with zero padding.
[0090] like Figure 9 As shown, a single frame of real image is obtained from the driving video sample, denoted as the first image frame. The first image feature (i.e., latent feature) is extracted from the first image frame using a variational autoencoder. Based on random noise, given a diffusion time step... Parameterized noise timetable This results in noise added to the features of the first image.
[0091] In the first stage of generation training of the world model, road topology conditions For the first road topology graph, the road topology conditions are... After processing by a convolutional network-based feature encoder, the first background feature and the second noisy feature are obtained. The features are then spliced together to obtain the second noisy feature.
[0092] In this embodiment of the invention, encoding the first target information of the first target included in the first image frame to obtain a second conditional encoding result may include: encoding the position information in the first target information to obtain a first position encoding result; encoding the category information in the first target information to obtain a first category encoding result; and using the first position encoding result and the first category encoding result as the second conditional encoding result. Spatially aligning the second conditional encoding result with a second noisy feature to obtain a first position embedding feature may include: concatenating the features of the first position encoding result and the first category encoding result, and aggregating the concatenated features using a linear layer to obtain the second conditional encoding result.
[0093] Due to target location conditions With the first noisy feature To address spatial misalignment, Fourier encoding and Contrastive Language-Image Pre-training (CLIP) encoding operations are first employed to conditionally define the target location. and target category conditions The features are processed and learnable parameters are introduced. After concatenating the features from both sources, a linear layer is used to aggregate the concatenated features to obtain the positional embedding.
[0094] ;
[0095] in, This indicates the computational operations of the MLP layer. This indicates the Fourier encoding operation. The bounding box category features are obtained by processing them using the CLIP network:
[0096] .
[0097] In this embodiment of the invention, the input driving data may also include text descriptions, such as text descriptions of weather, time, and location information for the driving scenario. Then, using the initial world model to perform denoising calculations based on the first image features and the first conditional encoding result to obtain the first generated image may include: encoding the first text description information corresponding to the first image frame to obtain a first cue embedding feature; and using the initial world model to perform denoising calculations based on the first image features, the first conditional encoding result, and the first cue embedding feature to obtain the first generated image.
[0098] Structured traffic conditions (first road topology map, 3D bounding box, target category) and text descriptions are encoded separately and input into a denoising network along with noisy features obtained by adding noise to the first image frame. Denoising is then performed to obtain the first predicted feature. Calculating the first loss value of the first generated image compared to the first image frame can include: calculating the loss value of the first predicted feature obtained from the denoising calculation compared to the features of the first image to obtain the first loss value.
[0099] Building upon the first phase of training, the model has grasped structured traffic information. However, an ideal world model should be able to predict the future and interact with the environment. In this embodiment, training a first-world model using driving video samples generates a sequence of future driving scenarios based on the input image sequence and its corresponding structured traffic conditions and driving actions, resulting in a second-world model. This process can include: encoding the structured traffic conditions corresponding to the first second image frame in the second image sequence from the driving video samples to obtain a third conditional encoding result; encoding the first driving action sequence corresponding to the second image sequence to obtain an action encoding result; generating future structured traffic conditions based on the third conditional encoding result and the action encoding result; generating a future driving scenario sequence using the first-world model based on the future structured traffic conditions; calculating a second loss value between the future driving scenario sequence and the corresponding third image sequence in the driving video samples, and updating the first-world model using the second loss value until a second iteration termination condition is met, thus obtaining the second-world model.
[0100] The first-world model obtained after the first stage of training can output image frames corresponding to the given time based on the input structured traffic conditions. If the input structured traffic conditions are sequential... , Then driving videos can be generated. However, in practical applications, structured traffic conditions beyond the current time step cannot be obtained, while the task requirement of this embodiment of the invention is to predict future driving scenario information, such as generating future videos.
[0101] To address this problem, embodiments of the present invention introduce a traffic prediction module, which utilizes driving behavior To iteratively predict future structured traffic conditions. In this embodiment of the invention, generating future structured traffic conditions based on the third conditional encoding result and the action encoding result includes: using the third conditional encoding result as the initial hidden state variable; performing cross-attention calculation based on the hidden state variable and the action encoding result at the corresponding time and converting it into a first latent variable with a Gaussian distribution; performing gated loop calculation based on the first latent variable and the hidden state variable at the corresponding time to obtain the hidden state variable at the next time, until the hidden state variable at the next time of the last second image frame of the second image sequence is obtained; concatenating the hidden state variable with the action encoding result at the previous time and performing decoding calculation to obtain the future structured traffic conditions.
[0102] The process of generating a future driving scene sequence using a first-world model based on future structured traffic conditions can include: outputting predicted features based on future structured traffic conditions using a first-world model; decoding the predicted features into the original image space corresponding to the driving video samples to obtain a future driving scene video including multiple future image frames; concatenating the predicted features with the action encoding results corresponding to the second image sequence and then performing decoding calculations to obtain a future driving action sequence for future moments in the second image sequence; and using the future driving scene video and the future driving action sequence as the future driving scene sequence.
[0103] In other words, a first-world model can be used to predict future driving scene videos and future driving action sequences, respectively. Based on the future driving scene videos and future driving action sequences, two types of losses can be calculated to accelerate the convergence of the second-stage training.
[0104] Calculating the second loss value of the future driving scene sequence relative to the third image sequence at the corresponding time in the driving video sample can include: calculating the third loss value of the future driving scene video relative to the third image sequence at the corresponding time in the driving video sample; calculating the fourth loss value of the future driving action sequence relative to the second driving action sequence at the corresponding time in the driving video sample; and calculating the second loss value based on the third and fourth loss values.
[0105] Similar to the embodiments described above, text-based prompts can also be introduced. In addition to the structured traffic conditions and text-based prompts generated by the traffic prediction module, reference image conditions can also be introduced. To enhance background consistency in the generated future videos, a first-world model is used to generate a sequence of future driving scenes based on future structured traffic conditions. This can include: encoding the reference background image corresponding to the driving video sample to obtain reference style embedding features; and using the first-world model to generate a sequence of future driving scenes based on the reference style embedding features and future structured traffic conditions. (Reference image conditions) The first second image frame in the second image sequence can be used to constrain the background of the future video generated by the first-world model. Alternatively, a background image of a specific traffic scene can be generated as a reference background image based on the current traffic scene.
[0106] Figure 10 This is a schematic diagram illustrating the prediction training process of a world model provided in an embodiment of the present invention.
[0107] like Figure 10 As shown, in the second stage of training the world model, a video prediction task is used to construct the world model. Initial observation reference image conditions are provided for the first second image frame in the second image sequence. Road topology conditions Target conditions and driving actions The desired outcome is future driving videos. and future driving behavior .
[0108] Introducing a traffic prediction module, utilizing driving behavior Iteratively predict future structural conditions.
[0109] First, the initial structured traffic conditions , Encode and expand into a one-dimensional latent space (i.e. Figure 10 (Spatial alignment conditions shown), then the encoded features are concatenated and processed sequentially through self-attention and linear layers to obtain the hidden state. .
[0110] Subsequently, using cross-attention layers To establish a connection between hidden states and driving behavior.
[0111] Then the latent variables Parameterized as a Gaussian distribution:
[0112] ;
[0113] in, and It is a network layer used to learn the parameters of a Gaussian distribution. It is an identity matrix.
[0114] To predict future hidden states, the traffic prediction module uses gated recurrent units (GRUs) for iterative updates:
[0115] .
[0116] Hide state With action characteristics The features are connected together and processed by a decoder, ultimately serving as the conditions for future traffic structure. Because the traffic prediction module predicts future traffic conditions at the feature level, it can mitigate pixel-level noise interference, resulting in more robust predictions.
[0117] In addition to the structured traffic conditions and textual prompts (prompt embedding) generated by the traffic prediction module, reference image conditions are also introduced. (Referencing style embedding) enhances background consistency in the generated future video.
[0118] Based on the above conditions, the first-world model obtained from the first stage of training is extended to enable the model to jointly generate future driving videos. and driving behavior Using the denoising network trained in the first stage, the input conditions are fused and the features for the next time step are predicted through temporal attention calculation, gated self-attention calculation, and cross-attention calculation. Specifically, video prediction features and action prediction features are calculated using Gaussian distributions. and Laplace distribution To model this, mean squared error and L1 loss are used to construct the loss function for video prediction training. , It is a learnable layer involved in the traffic prediction module, diffusion model, video decoder, and motion decoder.
[0119] For the video prediction part, a variational autoencoder image decoder is used to decode the predicted features into the original image space.
[0120] For the action prediction part, the predicted features are concatenated with historical action features, and then the future driving actions are generated through linear layer decoding.
[0121] To achieve closed-loop optimization of the decision-making agent within the world model, a reward predictor is needed to complete the final construction of the world model. Since reward estimation is a scalar prediction problem, and considering the difficulty in obtaining objectively evaluable reward labels in practical applications, this embodiment of the invention utilizes an image-based reward function to predict the reward state.
[0122] Figure 11 This is a schematic diagram of a closed-loop optimization process for a decision model based on a world model, provided as an embodiment of the present invention.
[0123] In embodiments of the present invention, such as Figure 11As shown, a reward module is constructed to calculate the driving safety coefficient of a vehicle at a future time using image frames from future videos generated based on a second-world model. Specifically, it can identify the positional relationship between the vehicle at a future time and lane lines, as well as other traffic behavior targets, to determine whether the vehicle is traveling in the designated lane and whether there is a risk of collision with other traffic behavior targets.
[0124] As described in the above embodiments, after training the world model using the collected driving video dataset, a virtual traffic environment is generated through the world model. A reinforcement learning interactive exploration training method is then used to iteratively optimize the autonomous decision-making system within the virtual environment, achieving learning through imagination. Considering that the world model includes an encoding module for driving perception data, this stage only requires training the decision-making module of the autonomous decision-making system. When deploying the model, the image encoding module in the second world model, used for encoding RGB video image data, can be interfaced with the decision-making module. The latent features encoded by the former are directly passed to the latter as state input, thus forming an end-to-end inference architecture.
[0125] like Figure 11 The illustrated agent-based closed-loop optimization framework, based on a world model, allows the latent predictive features output by the backbone network of the second world model to be directly used as input to the decision model during interactive training. This avoids redundant perceptual encoding operations during the decision model's inference process, simplifying the computation. However, in the reward calculation phase, the world model relies on the decoder to generate future prediction videos for environmental feedback and action evaluation.
[0126] In this embodiment of the invention, updating the decision model based on the reward value in S104 may include: calculating the cumulative reward value corresponding to the driving action based on the reward value and environmental state information in the future driving scenario sequence; calculating the value network loss value based on the value network of the decision model, environmental state information, and cumulative reward value; determining the weighted entropy of the policy network of the decision model based on the driving action parameters; calculating the policy network loss value based on the driving action parameters, the value network of the decision model, environmental state information, cumulative reward value, and weighted entropy; and updating the decision model based on the value network loss value and the policy network loss value.
[0127] Specifically, the cumulative reward value corresponding to the driving action is calculated based on the reward value and the environmental state information in the future driving scenario sequence. This can include: determining the imagined field of view of the vehicle; for the first time step in the future driving scenario sequence that is within the imagined field of view of the vehicle, determining the cumulative reward value at the corresponding time step based on the sixth reward value at the corresponding time step, the environmental state information at the next time step, and the cumulative reward value at the next time step; and for the second time step in the future driving scenario sequence that is equal to the imagined field of view, determining the cumulative reward value at the corresponding time step based on the environmental state information at the corresponding time step.
[0128] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0129] An embodiment of the present invention also provides a decision model optimization device based on a world model, comprising: an acquisition module for acquiring driving video samples; a first training module for training an initial world model using the driving video samples to generate image frames at corresponding times based on input structured traffic conditions, thereby obtaining a first world model; a second training module for training the first world model using the driving video samples to generate a future driving scenario sequence based on the input image sequence and its corresponding structured traffic conditions and driving actions, thereby obtaining a second world model; and a third training module for iteratively training a decision model using the second world model, wherein during training, the driving actions output by the decision model and their corresponding structured traffic conditions are input into the second world model, the reward value of the driving actions is calculated based on the obtained future driving scenario sequence, and the decision model is updated based on the reward value.
[0130] In this embodiment of the invention, the first training module trains an initial world model using driving video samples to generate image frames at corresponding times based on input structured traffic conditions, thereby obtaining a first world model. This process may include: extracting first image features from the first image frame in the driving video samples; encoding the structured traffic conditions corresponding to the first image frame to obtain a first conditional encoding result; generating a first generated image using the initial world model based on the first image features and the first conditional encoding result; calculating a first loss value between the first generated image and the first image frame, and updating the initial world model using the first loss value until a first iteration termination condition is met, thereby obtaining the first world model.
[0131] In this embodiment of the invention, the generation unit of the initial world model can be a diffusion model; the first training module uses the initial world model to generate a first generated image based on the first image features and the first conditional coding result, which may include: adding noise to the first image features to obtain a first noisy feature; and using the initial world model to perform denoising calculation based on the first image features and the first conditional coding result to obtain the first generated image.
[0132] In this embodiment of the invention, the first training module encodes the structured traffic conditions corresponding to the first image frame to obtain a first conditional encoding result, which may include: extracting a first background feature from the first road topology map corresponding to the first image frame; concatenating the first noisy feature and the first background feature to obtain a second noisy feature; encoding the first target information of the first target included in the first image frame to obtain a second conditional encoding result; spatially aligning the second conditional encoding result with the second noisy feature to obtain a first position embedding feature; and using the first background feature and the first position embedding feature as the first conditional encoding result. The first training module uses an initial world model to perform denoising calculations based on the first image features and the first conditional encoding result to obtain a first generated image, which may include: using an initial world model to perform denoising calculations based on the second noisy feature and the second conditional encoding result to obtain the first generated image.
[0133] In this embodiment of the invention, the first training module encodes the first target information of the first target included in the first image frame to obtain a second conditional encoding result. This may include: encoding the position information in the first target information to obtain a first position encoding result; encoding the category information in the first target information to obtain a first category encoding result; and using the first position encoding result and the first category encoding result as the second conditional encoding result. The first training module performs spatial alignment processing on the second conditional encoding result and the second noisy feature to obtain a first position embedding feature. This may include: concatenating the features of the first position encoding result and the first category encoding result, and aggregating the concatenated features using a linear layer to obtain the second conditional encoding result.
[0134] In this embodiment of the invention, the first training module calculates a first loss value of the first generated image relative to the first image frame, which may include: calculating the loss value of the first predicted feature obtained by denoising calculation relative to the first image feature to obtain the first loss value.
[0135] In this embodiment of the invention, the first training module uses an initial world model to perform denoising calculations based on the first image features and the first conditional coding result to obtain a first generated image. This may include: encoding the first text description information corresponding to the first image frame to obtain a first prompt embedding feature; and using the initial world model to perform denoising calculations based on the first image features, the first conditional coding result, and the first prompt embedding feature to obtain the first generated image.
[0136] In this embodiment of the invention, the second training module uses driving video samples to train a first-world model to generate a future driving scene sequence based on the input image sequence and its corresponding time-time structured traffic conditions and driving actions, thus obtaining a second-world model. This process may include: encoding the structured traffic conditions corresponding to the time of the first second image frame in the second image sequence of the driving video samples to obtain a third conditional encoding result; encoding the first driving action sequence corresponding to the time of the second image sequence to obtain an action encoding result; generating future structured traffic conditions based on the third conditional encoding result and the action encoding result; generating a future driving scene sequence based on the future structured traffic conditions using the first-world model; calculating a second loss value of the future driving scene sequence compared to the corresponding time-time third image sequence in the driving video samples, and updating the first-world model using the second loss value until a second iteration termination condition is met, thus obtaining the second-world model.
[0137] In this embodiment of the invention, the second training module generates future structured traffic conditions based on the third conditional encoding result and the action encoding result. This may include: using the third conditional encoding result as the initial hidden state variable; performing cross-attention calculation based on the hidden state variable and the action encoding result at the corresponding time and converting it into a first latent variable with a Gaussian distribution; performing gated loop calculation based on the first latent variable and the hidden state variable at the corresponding time to obtain the hidden state variable at the next time, until the hidden state variable at the next time of the last second image frame of the second image sequence is obtained; and then performing decoding calculation after concatenating the hidden state variable with the action encoding result at the previous time to obtain the future structured traffic conditions.
[0138] In this embodiment of the invention, the second training module generates a future driving scene sequence based on future structured traffic conditions using a first-world model. This can include: outputting predicted features based on future structured traffic conditions using the first-world model; decoding the predicted features into the original image space corresponding to the driving video samples to obtain a future driving scene video including multiple future image frames; concatenating the predicted features with the action encoding results corresponding to the second image sequence and then performing decoding calculations to obtain a future driving action sequence for future moments in the second image sequence; and using the future driving scene video and the future driving action sequence as the future driving scene sequence.
[0139] In this embodiment of the invention, the second training module calculates a second loss value for the future driving scene sequence compared to a third image sequence at a corresponding time in the driving video sample. This may include: calculating a third loss value for the future driving scene video compared to a third image sequence at a corresponding time in the driving video sample; calculating a fourth loss value for the future driving action sequence compared to a second driving action sequence at a corresponding time in the driving video sample; and calculating the second loss value based on the third and fourth loss values.
[0140] In this embodiment of the invention, the second training module uses a first-world model to generate a sequence of future driving scenarios based on future structured traffic conditions. This may include: encoding a reference background image corresponding to a driving video sample to obtain a reference style embedding feature; and using the first-world model to generate a sequence of future driving scenarios based on the reference style embedding feature and the future structured traffic conditions.
[0141] For a description of the features in the embodiment of the decision model optimization device based on the world model, please refer to the relevant description of the embodiment of the decision model optimization method based on the world model, which will not be repeated here.
[0142] Embodiments of the present invention also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above embodiments of the world model-based decision model optimization method.
[0143] Embodiments of the present invention also provide a non-volatile storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the decision model optimization method based on the world model.
[0144] In one exemplary embodiment, the aforementioned non-volatile storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0145] Embodiments of the present invention also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the decision model optimization method based on a world model.
[0146] Embodiments of the present invention also provide another computer program product, including a non-volatile storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above embodiments of the world-model-based decision model optimization method.
[0147] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0148] The above provides a detailed description of the world-model-based decision model optimization method, device, medium, and product provided by this invention. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of these embodiments are only intended to aid in understanding the method and core ideas of this invention. It should be noted that those skilled in the art can make various improvements and modifications to this invention without departing from its principles, and these improvements and modifications also fall within the protection scope of this invention.
Claims
1. A decision model optimization method based on a world model, characterized in that, include: Obtain driving video samples; The initial world model is trained using the driving video samples, and image frames corresponding to the time are generated based on the input structured traffic conditions to obtain the first world model; The first world model is trained using the driving video samples to generate a sequence of future driving scenarios based on the input image sequence and the structured traffic conditions and driving actions at the corresponding time, thus obtaining the second world model. The decision model is iteratively trained using the second world model. During training, the driving actions output by the decision model and their corresponding structured traffic conditions are input into the second world model. The reward value of the driving action is calculated based on the obtained future driving scenario sequence, and the decision model is updated based on the reward value.
2. The decision model optimization method based on the world model according to claim 1, characterized in that, Using the driving video samples, an initial world model is trained to generate image frames corresponding to the input structured traffic conditions, resulting in a first world model, including: Extract the first image feature from the first image frame in the driving video sample; The structured traffic conditions corresponding to the first image frame are encoded to obtain the first conditional encoding result; Using the initial world model, a first generated image is generated based on the first image features and the first conditional encoding result; Calculate a first loss value for the first generated image compared to the first image frame, and update the initial world model using the first loss value until the first iteration termination condition is met, thereby obtaining the first world model.
3. The decision model optimization method based on the world model according to claim 2, characterized in that, The initial world model is generated by a diffusion model; Using the initial world model, based on the first image features and the first conditional encoding result, a first generated image is generated, including: Add noise to the first image feature to obtain the first noisy feature; The first generated image is obtained by performing denoising calculations based on the first image features and the first conditional coding result using the initial world model.
4. The decision model optimization method based on the world model according to claim 3, characterized in that, The structured traffic conditions corresponding to the first image frame are encoded to obtain a first conditional encoding result, including: Extract the first background feature from the first road topology map corresponding to the first image frame; By concatenating the first noisy feature and the first background feature, a second noisy feature is obtained; The first target information of the first target included in the first image frame is encoded to obtain the second conditional encoding result; The second conditional encoding result is spatially aligned with the second noisy feature to obtain the first positional embedding feature; The first background feature and the first position embedding feature are used as the first conditional encoding result; The first generated image is obtained by performing denoising calculations based on the first image features and the first conditional coding result using the initial world model, including: The first generated image is obtained by performing denoising calculations based on the second noisy feature and the second conditional coding result using the initial world model.
5. The decision model optimization method based on the world model according to claim 4, characterized in that, Encoding the first target information of the first target included in the first image frame to obtain a second conditional coding result includes: The location information in the first target information is encoded to obtain the first location encoding result; The category information in the first target information is encoded to obtain the first category encoding result; The first position encoding result and the first category encoding result are used as the second conditional encoding result; The second conditional encoding result is spatially aligned with the second noisy feature to obtain the first positional embedding feature, including: The first position encoding result and the first category encoding result are concatenated, and the concatenated features are aggregated using a linear layer to obtain the second conditional encoding result.
6. The decision model optimization method based on the world model according to claim 3, characterized in that, Calculating the first loss value of the first generated image relative to the first image frame includes: The first loss value is obtained by calculating the loss value of the first predicted feature obtained from the denoising calculation compared with the first image feature.
7. The decision model optimization method based on the world model according to claim 3, characterized in that, The first generated image is obtained by performing denoising calculations based on the first image features and the first conditional coding result using the initial world model, including: Encode the first text description information corresponding to the first image frame to obtain the first prompt embedding feature; The first generated image is obtained by performing denoising calculations based on the first image features, the first conditional coding result, and the first cue embedding features using the initial world model.
8. The decision model optimization method based on the world model according to claim 2, characterized in that, The first world model is trained using the driving video samples to generate a sequence of future driving scenarios based on the input image sequence and its corresponding structured traffic conditions and driving actions, resulting in a second world model, including: The structured traffic conditions at the time corresponding to the first second image frame in the second image sequence of the driving video sample are encoded to obtain the third conditional encoding result; Encode the first driving action sequence at the time corresponding to the second image sequence to obtain the action coding result; Future structured traffic conditions are generated based on the third condition coding result and the action coding result; The first world model is used to generate the future driving scenario sequence based on the future structured traffic conditions; Calculate the second loss value of the future driving scene sequence compared to the third image sequence at the corresponding time in the driving video sample, and use the second loss value to update the first world model until the second iteration termination condition is met, thus obtaining the second world model.
9. The decision model optimization method based on the world model according to claim 8, characterized in that, Future structured traffic conditions are generated based on the third condition coding result and the action coding result, including: Using the third conditional encoding result as the initial hidden state variable, cross-attention calculation is performed based on the hidden state variable and the action encoding result at the corresponding time, and the result is converted into a first latent variable with a Gaussian distribution. Gated loop calculation is performed based on the first latent variable and the hidden state variable at the corresponding time to obtain the hidden state variable at the next time, until the hidden state variable at the next time of the last second image frame of the second image sequence is obtained. The hidden state variable is concatenated with the action encoding result of the previous time step and then decoded to obtain the future structured traffic conditions.
10. The decision model optimization method based on the world model according to claim 9, characterized in that, Based on the future structured traffic conditions, the first world model is used to generate the sequence of future driving scenarios, including: The first world model is used to output predicted features based on the future structured traffic conditions; The predicted features are decoded into the original image space corresponding to the driving video sample to obtain a future driving scene video including multiple future image frames; After concatenating the predicted features with the action encoding results corresponding to the second image sequence, decoding calculation is performed to obtain the future driving action sequence of the second image sequence at future moments; The future driving scenario video and the future driving action sequence are used as the future driving scenario sequence.
11. The decision model optimization method based on a world model according to claim 10, characterized in that, Calculating the second loss value of the future driving scene sequence relative to the third image sequence at the corresponding time in the driving video sample includes: Calculate the third loss value of the future driving scene video compared to the third image sequence at the corresponding time in the driving video sample; Calculate the fourth loss value of the future driving action sequence compared to the second driving action sequence at the corresponding time in the driving video sample; The second loss value is calculated based on the third loss value and the fourth loss value.
12. The decision model optimization method based on the world model according to claim 8, characterized in that, Based on the future structured traffic conditions, the first world model is used to generate the sequence of future driving scenarios, including: The reference background image corresponding to the driving video sample is encoded to obtain reference style embedding features; The first world model is used to generate the future driving scenario sequence based on the reference style embedding features and the future structured traffic conditions.
13. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the world-model-based decision model optimization method as described in any one of claims 1 to 12 when executing the computer program.
14. A non-volatile storage medium, characterized in that, The non-volatile storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the decision model optimization method based on the world model as described in any one of claims 1 to 12.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the decision model optimization method based on a world model as described in any one of claims 1 to 12.
Citation Information
Patent Citations
End-to-end automatic driving method and system based on reinforcement learning driving world model
CN117218618A
Keyframe-based compression for world model representation in autonomous systems and applications
US20230288223A1