World model driven decision model training method, system, equipment and product

By constructing a world model based on video data and combining it with a third-order motion prior loss function and uncertainty reward, the problem of the world model failing to accurately predict future states is solved, achieving closed-loop optimization and safety improvement of the autonomous driving system.

CN120756503APending Publication Date: 2025-10-10SHANDONG HAILIANG INFORMATION TECH RES INST
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510897191.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing technologies use world models to sample and synthesize training data, but fail to accurately predict and generate future states, and cannot support closed-loop optimization of decision models.

Method used

By collecting target video data to build a dataset, the pre-trained diffusion generative model is used to generate an initial world model. The initial world model is fine-tuned by combining the diffusion loss function, dynamic loss function and structure-preserving loss function of the third-order motion prior to generate a fine-tuned world model. The uncertainty of the world model prediction is used to automatically generate a reward function and close the loop to train the decision model.

Benefits of technology

It achieves physical consistency and high-frequency detail fidelity for short-term and long-term predictions, improves training efficiency, realizes the coordinated optimization of environmental cognition and strategy evolution, has low latency, high robustness and easy scalability, and improves the safety of autonomous driving systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120756503A_ABST
    Figure CN120756503A_ABST
Patent Text Reader

Abstract

The invention discloses a world model driven decision model training method, system, device and product, and relates to the technical field of artificial intelligence. According to the scheme, the initial world model is generated through the target video data and the diffusion generation model, and the initial world model is finely adjusted by using three different loss functions, namely the diffusion loss function, the dynamic loss function and the structure maintenance loss function based on the third-order motion prior; physical consistency and high-frequency detail fidelity of short-term and long-range prediction are realized; furthermore, a reward function is automatically generated by using the uncertainty of world model prediction, so that the training efficiency is improved; according to target video data and a world model closed-loop training decision model, collaborative optimization of environment cognition and strategy evolution is realized; and finally, the trained world model and the decision model can be integrated to the target server, closed-loop control of perception-decision-motion execution is realized, the method has low delay, high robustness and expansibility, and the safety of the automatic driving system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a world model-driven decision model training method, system, device and product. Background Art

[0002] World model technology is becoming increasingly important in autonomous driving decision-making systems because it effectively models the dynamics of the environment and predicts future states. By establishing state transition probability functions and observation generation models, world models create a virtual interactive environment for autonomous driving systems, enabling model-based reinforcement learning strategies to conduct large-scale gradient-guided exploration within a safe range.

[0003] However, current approaches focus on sampling and synthesizing training data through world models, but fail to truly predict and generate future states, thus failing to support closed-loop optimization of decision models. For example, some methods rely on explicit reward labels for environmental feedback, increasing the cost of data collection and annotation. Other methods, while combining immediate and future rewards for trajectory evaluation, essentially still only generate multiple candidate trajectories through world models, similarly failing to meet the requirements of closed-loop optimization of decision models.

[0004] In view of the above, how to solve the current problem that mainly relies on sampling and synthesizing training data through world models, but fails to achieve accurate prediction and generation of future states, and thus cannot support closed-loop optimization of decision models, is an urgent problem that needs to be solved by technical personnel in this field. Summary of the Invention

[0005] The present invention provides a world model-driven decision model training method, system, device and product to at least solve the problem that the current training data is mainly sampled and synthesized by the world model, but cannot achieve accurate prediction and generation of future states, and thus cannot support closed-loop optimization of the decision model.

[0006] The present invention provides a world model driven decision model training method, comprising: Collect target video data and construct a target video dataset based on the target video data; Generate an initial world model based on the target video dataset and the pre-trained diffusion generative model, and construct a diffusion loss function, a dynamic loss function, and a structure-preserving loss function based on the target video dataset. Fine-tune the initial world model according to the diffusion loss function, the dynamic loss function, and the structure preservation loss function to generate a fine-tuned world model; The reward function of the decision model is generated according to the world model, and the decision model is trained in a closed-loop based on the target video data and the world model.

[0007] The application further provides a motion trajectory generation method, comprising: target video data, and constructs a target video dataset according to the target video data; generating an initial world model according to the target video dataset and a pre-trained diffusion generation model, and constructing a diffusion loss function, a dynamic loss function and a structure preservation loss function based on a third-order motion prior according to the target video dataset; According to the diffusion loss function, the dynamic loss function and the structure preservation loss function, the initial world model is fine-tuned to generate a fine-tuned world model; According to the world model, a reward function of a decision model is generated, and the decision model is trained in a closed loop according to the target video data and the world model; The world model and the decision model are deployed to a target server, so that the target server generates a motion trajectory based on the world model and the decision model.

[0008] The application further provides a motion trajectory generation system, comprising: The cloud server is used for collecting target video data, and constructing a target video dataset according to the target video data; generating an initial world model according to the target video dataset and a pre-trained diffusion generation model, and constructing a diffusion loss function, a dynamic loss function and a structure preservation loss function based on a third-order motion prior according to the target video dataset; fine-tuning the initial world model according to the diffusion loss function, the dynamic loss function and the structure preservation loss function, to generate a fine-tuned world model; generating a reward function of a decision model according to the world model, and training the decision model in a closed loop according to the target video data and the world model; and deploying the world model and the decision model to a target server; The target server is used for generating a motion trajectory based on the world model and the decision model.

[0009] The application further provides an electronic device, comprising: a memory for storing a computer program; a processor for executing the computer program to realize the steps of the above-mentioned any one world model driven decision model training method.

[0010] The application further provides a computer program product, comprising a computer program, which realizes the steps of the above-mentioned any one world model driven decision model training method when executed by a processor.

[0011] The beneficial effects of the present invention lie in that an initial world model is generated through a target video dataset and a pre-trained diffusion generation model, and three different loss functions, namely a diffusion loss function based on third-order motion prior, a dynamic loss function, and a structure-preserving loss function, are used to fine-tune the initial world model, thereby achieving physical consistency and high-frequency detail fidelity of short-term and long-term predictions; on this basis, a reward function is automatically generated using the uncertainty of world model predictions without the need for manual or external detector labeling, thereby improving training efficiency; a decision model is trained in a closed-loop manner based on target video data and world models, thereby achieving coordinated optimization of environmental cognition and strategy evolution; finally, the trained world model and decision model can be integrated into the target server, so that the target server can generate motion trajectories based on the world model and the decision model, thereby achieving closed-loop control of perception-decision-motion execution, with low latency, high robustness, and easy scalability, thereby improving the safety of the autonomous driving system.

[0012] In addition, the present invention also provides a motion trajectory generation system, device and product, with the same effects as above. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0014] Figure 1 A flowchart of a world model-driven decision model training method provided by an embodiment of the present invention; Figure 2 A schematic diagram of a data set collection process provided by an embodiment of the present invention; Figure 3 A schematic diagram of a world model training process based on a video diffusion model provided by an embodiment of the present invention; Figure 4 A schematic diagram of a long temporal video frame generation process provided by an embodiment of the present invention; Figure 5 A schematic diagram of an end-to-end autonomous driving system based on a world model provided by an embodiment of the present invention; Figure 6 A schematic diagram of a motion trajectory generation system provided by an embodiment of the present invention; Figure 7 A schematic diagram of a world model-driven decision model training device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0015] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0016] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.

[0017] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0018] At present, world model technology focuses more on sampling and synthesizing training data through world models, but fails to truly realize the prediction and generation of future states, and thus cannot support the closed-loop optimization of decision models. For example, some methods rely on explicit environmental feedback reward labels, which increases the cost of data collection and annotation; other methods combine immediate rewards with future rewards for trajectory evaluation, but in essence they still only generate multiple candidate trajectories through world models, which also cannot meet the needs of closed-loop optimization of decision models. Therefore, in order to solve the above problems, the present invention provides a world model-driven decision model training method. It should be noted that the method provided by the present invention is applied to a cloud server, and the cloud server is communicatively connected to the target server. In this embodiment, there is no restriction on the specific type of the target server. For example, it can be an on-board edge server, an on-board computing unit, or an on-board artificial intelligence reasoning unit, depending on the specific implementation situation. The world model-driven decision model training method is described in detail below in conjunction with specific embodiments.

[0019] Figure 1 Flowchart of a world model driven decision model training method provided by an embodiment of the present invention. Figure 1 As shown, the method includes: S10: Collect target video data and construct a target video dataset based on the target video data.

[0020] Specifically, the driver controls the vehicle, and onboard sensors collect and acquire driving data. It should be noted that due to the limited computing and storage resources of the onboard target server, the driving data is only temporarily cached on the target server and will be uploaded to the cloud server after preprocessing.

[0021] Figure 2 Schematic diagram of the data set collection process provided by the embodiment of the present invention. Figure 2 As shown in the figure, when collecting driving data, at least the following information should be collected: front-view RGB image, which is used to obtain the front-view image information of the vehicle in real time during driving through the front-view vehicle camera. Each image is are stored in the form of The height and width of the image are respectively; in addition, the corresponding time step mark will be attached when the image is stored. Vehicle posture information, through the vehicle inertial measurement unit, records the vehicle posture data of each time step, including the short-term relative position and direction angle of the vehicle. The vehicle control information is stored in the form of and aligned with the front view RGB image according to the time step mark. The vehicle control action of each time step is recorded by the on-board controller, including the steering angle and speed. The target video data is stored in the form of a time step and aligned with the front view RGB image according to the time step mark. In this embodiment, the target video data is generated based on the above three types of information, and the target video dataset is constructed based on the target video data. In this embodiment, the specific process of constructing the target video dataset based on the target video data is not limited and depends on the implementation situation.

[0022] S11: Generate an initial world model based on the target video dataset and the pre-trained diffusion generative model, and construct a diffusion loss function, a dynamic loss function, and a structure-preserving loss function based on the target video dataset.

[0023] Furthermore, in this example, the pre-trained Stable Video Diffusion model (SVD) is used to initialize the world model. The data sample Converted to noise , and generate new samples from Gaussian noise by gradually denoising the latent space until the noise variance is 0. The training of SVD can be simplified to minimize ;in is a parameterized U-network (UNet) denoiser, is a reweighting function which is omitted for brevity. Based on this framework, SVD handles a range of noise potentials. , and generates a video containing K = 25 frames. The generation process is guided by a conditional image whose latent representation is concatenated to the input channel-wise as a reference for content generation.

[0024] Because the SVD predicted image differs from the true image, it cannot undergo autoregressive expansion due to inconsistent content, limiting the time series length of its future predictions. Furthermore, SVD has limitations in handling the complex dynamics of driving scenarios, resulting in irrational motion behavior in its generated results. To address these issues, the present invention employs a potential alternative approach to achieve consistent future predictions. Specifically, because world models aim to predict the future from the current state, their prediction starting point should be closely aligned with the conditional image. Therefore, SVD is customized as a dedicated prediction model, using the first frame as the conditional image and discarding noise enhancement during training. Leveraging this predictive capability, the world model is able to perform long-term extrapolation by iteratively predicting short-term segments and resetting the conditional image to the final segment. However, training using this setup often results in irrational dynamics relative to historical frames, particularly in long-term extrapolations. This is primarily due to ambiguity caused by insufficient prior information about future motion trends. It should be noted that this embodiment does not restrict the specific process of generating the initial world model.

[0025] To achieve consistent future scene estimation, the present invention introduces three fundamental prior conditions: position, velocity, and acceleration. Velocity and acceleration are the first- and second-order derivatives of position, respectively. These can be fully derived by conditioning on three consecutive frames. This allows a diffusion loss function based on a third-order motion prior to be constructed based on the target video dataset. This enables the world model to fully capture the state of surrounding instances and, through iterative expansion, to predict a more consistent and reasonable long-term future. The specific process for generating the diffusion loss function is not limited in this embodiment.

[0026] On the other hand, considering that in most driving videos, the monotonous area in the distance occupies most of the field of view, and the moving foreground instances only occupy a relatively small area, the latter usually exhibits higher randomness, which makes its prediction complicated. Since the diffusion loss function uniformly supervises all outputs, it cannot effectively distinguish the subtle differences between different areas, so the model cannot efficiently learn to predict the dynamics of reality in key areas. Since the difference between two adjacent frames provides rich motion patterns, the present invention introduces an additional supervision mechanism to promote the learning of the dynamics of key areas. The dynamic loss function is specifically constructed according to the target video dataset. By adaptively reweighting the standard diffusion loss, the dynamic loss function can improve the learning efficiency of dynamic areas. In this embodiment, there is no restriction on the specific generation process of the dynamic loss function.

[0027] Finally, in high-resolution dynamic driving scene predictions, the predicted structural details can be severely degraded, manifesting as overly smoothed or fragmented objects. For example, the outline of a vehicle rapidly disintegrates during movement. To alleviate this problem, greater attention must be paid to structural details. Therefore, the present invention also constructs a structure-preserving loss function based on the target video dataset to minimize the difference in high-frequency features between the prediction and the true value, thereby preserving more structural information. This embodiment does not restrict the generation process of the structure-preserving loss function.

[0028] S12: Fine-tune the initial world model according to the diffusion loss function, the dynamic loss function, and the structure preservation loss function to generate a fine-tuned world model.

[0029] After generating the three loss functions, the initial world model is fine-tuned using the three loss functions to generate a fine-tuned world model. It should be noted that this embodiment does not limit the fine-tuning method for the initial world model; fine-tuning can be performed continuously or in stages, depending on the specific implementation.

[0030] S13: Generate a reward function for the decision model based on the world model, and perform closed-loop training of the decision model based on the target video data and the world model.

[0031] In order to achieve closed-loop optimization of the autonomous driving decision-making system in the world model, it is necessary to combine the reward predictor to complete the final construction of the world model. Currently, rewards are generally established by using external detectors. However, these detectors are usually developed on specific data sets, and there is a reward estimation bottleneck in real complex dynamic scenes. Considering that the world model will generate observations with increased diversity based on out-of-distribution conditions, the present invention uses the prediction uncertainty of the world model itself as the source of reward, that is, estimates uncertainty through conditional variance. In this embodiment, there is no restriction on the specific process of generating the reward function of the decision model based on the world model, which depends on the specific implementation situation. Finally, the decision model is closed-loop trained based on the target video data and the world model. In this embodiment, there is no restriction on the specific structure and training process of the decision model, which depends on the specific implementation situation.

[0032] It should also be noted that, based on the above-mentioned decision model training process, the present invention also provides a motion trajectory generation method. Specifically, the method involves collecting target video data and constructing a target video dataset based on the target video data; generating an initial world model based on the target video dataset and a pre-trained diffusion generative model; and constructing a diffusion loss function, a dynamic loss function, and a structure-preserving loss function based on a third-order motion prior based on the target video dataset; fine-tuning the initial world model based on the diffusion loss function, the dynamic loss function, and the structure-preserving loss function to generate a fine-tuned world model; generating a reward function for the decision model based on the world model, and closed-loop training the decision model based on the target video data and the world model. After the world model and decision model are constructed and trained, the world model and decision model can be further deployed to a target server so that the target server can generate motion trajectories based on the world model and decision model. It is understood that the target server can be an on-board edge server or other server for generating motion trajectories, depending on the specific implementation. In addition, the specific process by which the target server generates motion trajectories based on the world model and decision model is not limited in this embodiment and depends on the specific implementation.

[0033] In this embodiment, an initial world model is generated using a target video dataset and a pre-trained diffusion generative model, and three different loss functions, namely a diffusion loss function based on third-order motion prior, a dynamic loss function, and a structure-preserving loss function, are used to fine-tune the initial world model, thereby achieving physical consistency and high-frequency detail fidelity for short-term and long-term predictions. On this basis, a reward function is automatically generated using the uncertainty of the world model prediction without the need for manual or external detector labeling, thereby improving training efficiency. A decision model is trained in a closed loop based on the target video data and the world model, thereby achieving coordinated optimization of environmental cognition and strategy evolution. Finally, the trained world model and decision model can be integrated into the target server so that the target server can generate motion trajectories based on the world model and the decision model, thereby achieving closed-loop control of perception-decision-motion execution, with low latency, high robustness, and easy scalability, thereby improving the safety of the autonomous driving system.

[0034] Based on the above embodiments, in some embodiments, such as Figure 2 As shown, target video data is collected and a target video dataset is constructed based on the target video data, including: S101: Collecting forward-view image information during vehicle driving according to a first preset frequency.

[0035] S102: Collecting vehicle posture data and vehicle control information according to a second preset frequency.

[0036] S103: Aggregating the forward-looking image information according to a preset number of frames to generate a plurality of target video data.

[0037] S104: Time-step alignment of each target video data with the corresponding vehicle posture data and vehicle control information to generate a target video dataset.

[0038] Specifically, in order to collect target video data and construct a target video data set, this embodiment specifically collects forward-view image information during vehicle driving according to a first preset frequency, and collects vehicle posture data and vehicle control information according to a second preset frequency.

[0039] It should be noted that in this embodiment, there is no restriction on the size of the first preset frequency and the second preset frequency. For example, the first preset frequency and the second preset frequency can be 12 Hz and 2 Hz respectively. Then, during data acquisition, the forward-view image information is acquired and processed at a frequency of 12 Hz to obtain an image data set; the vehicle posture data and vehicle control information are acquired and processed at a frequency of 2 Hz to obtain a vehicle status data set, that is, the vehicle posture and control information are stored once every 6 frames of images.

[0040] Subsequently, the forward-view image information is aggregated according to a preset frame count, thereby dividing the image dataset into a number of target video data. In this embodiment, the preset frame count is not limited and depends on the specific implementation. For example, each target video data set is 2 seconds long and contains 25 frames of images, along with the vehicle pose and control data aligned with the target video data set. Finally, each target video data set is time-aligned with the corresponding vehicle pose data and vehicle control information to generate a target video dataset.

[0041] In this embodiment, forward-view image information, vehicle posture data, and vehicle control information are collected at different frequencies during vehicle driving, and the forward-view image information is aggregated to generate multiple target video data. Finally, each target video data is time-step aligned with the corresponding vehicle posture data and vehicle control information, thereby achieving the complete construction of the target video dataset, so as to facilitate the subsequent training of the world model based on this dataset.

[0042] Figure 3 A schematic diagram of a world model training process based on a video diffusion model provided by an embodiment of the present invention. Based on the above embodiments, in some embodiments, such as Figure 3 As shown, generating an initial world model based on the target video dataset and the pre-trained diffusion generation model includes: S110: compressing each frame image in the target video dataset using a pre-trained autoencoder to generate a potential feature corresponding to each frame image; S111: Encode the prediction condition information in the target video data set to generate a conditional input, wherein the prediction condition information at least includes a conditional frame image, a trajectory action sequence, an action sequence number, and a sample sequence number.

[0043] S112: Add Gaussian noise to the latent features to generate a noisy latent sequence and label the conditional frames to generate an initial world model.

[0044] The conditional frame of the noise potential sequence corresponding to the initial generation period is marked as the first frame, and the conditional frame of the noise potential sequence corresponding to the non-initial generation period is marked as the first three frames.

[0045] In a specific implementation, for the target video dataset constructed in the above embodiment, each video sample contains 25 frames of front-view RGB images and vehicle posture data and vehicle control information at the corresponding time step. As the input of the adaptation model, the pre-trained autoencoder is first used to compress each frame of the image into a low-dimensional latent space to obtain the latent features. : ; in, is a video clip sample in the target video dataset, , For a single frame image, are the number of channels and resolution of the image, is an autoencoder model.

[0046] Then, the prediction condition information is encoded through a condition processing module. Figure 3 As shown, the prediction condition information includes four parts of condition information: the conditional frame image, the trajectory action sequence, the action sequence number, and the sample sequence number. Different conditions are processed using different encoding models to generate the conditional input. In this embodiment, there is no limitation on the method for generating the conditional input, which depends on the specific implementation.

[0047] Furthermore, the potential features Add Gaussian noise .in, Indicates that the mean is 0 and the covariance matrix is Gaussian distribution of , generates a noisy latent sequence and labels the conditional frames to generate the initial world model.

[0048] Figure 4 Schematic diagram of the long time sequence video frame generation process provided by the embodiment of the present invention. Figure 4 As shown, for the noise potential sequence in the initial generation cycle, the conditional frame is marked as the first frame. For the noise potential sequence generated by autoregression in subsequent generation cycles, the conditional frames are marked as the first three frames, that is, the last three frames corresponding to the last prediction generation result, which are used as prior injection. In this embodiment, there is no limit on the number of generation cycles. For example, the number of generation cycles can be set to 6, and a long-term future video frame sequence of length 135 can be generated through autoregression in one time.

[0049] In summary, the generation of an initial world model is achieved, which serves as the basis for the world model, so that it can be fine-tuned using the three loss functions to generate a complete and accurate world model.

[0050] In order to generate the conditional input, based on the above embodiment, in some embodiments, the prediction condition information in the target video dataset is encoded to generate the conditional input, including: S113: Encode the conditional frame image through an autoencoder and a contrastive learning image-text encoder to output image conditional features and semantic conditional features respectively.

[0051] S114: Encode the trajectory action sequence, action sequence number, and sample sequence number respectively through a time step encoding function to generate conditional features after the trajectory action sequence, action sequence number, and sample sequence number are encoded.

[0052] S115: Concatenate the conditional features obtained by encoding the action sequence number and the sample sequence number to generate a concatenated conditional feature.

[0053] S116: Aggregate image conditional features, semantic conditional features, conditional features after trajectory action sequence encoding, and splicing conditional features to generate conditional input.

[0054] It can be seen from the above embodiments that the prediction condition information includes conditional frame images, trajectory action sequences, action numbers and sample numbers, and different conditions are processed using different encoding models. In this embodiment, the conditional frame images are specifically encoded using an autoencoder and a contrastive learning image-text (CLIP) encoder to obtain two parts of conditional features: the autoencoder outputs image conditional features, and the CLIP encoder outputs semantic conditional features. The trajectory action sequence, action number and sample number are encoded using a time step encoding function, and the conditional features after the action number and sample number are encoded are spliced ​​to obtain spliced ​​conditional features. Finally, the image conditional features, semantic conditional features, conditional features after the trajectory action sequence is encoded and the spliced ​​conditional features are aggregated to generate the conditional input . In this way, the generation of conditional input is achieved.

[0055] Based on the above embodiments, in some embodiments, a diffusion loss function, a dynamic loss function, and a structure preservation loss function based on a third-order motion prior are constructed according to a target video dataset, including: S117: Construct a frame-level mask.

[0056] The mask is set in time order, and the number of elements that are 1 is not greater than 3.

[0057] S118: Determine an input latent variable according to the frame-level mask, the latent feature, and the noise latent sequence, and generate a diffusion loss function according to the frame-level mask, the latent feature, the noise latent sequence, and the input latent variable.

[0058] S119: Generate dynamic perception weights corresponding to each target video data according to the corresponding input latent variables and latent features, and generate a dynamic loss function according to each dynamic perception weight.

[0059] S120: Performing Fourier transform on each potential feature, and generating a structure preservation loss function based on each potential feature after Fourier transform and the frame-level mask.

[0060] Specifically, in order to construct a diffusion loss function based on the third-order motion prior, this embodiment specifically constructs a frame-level mask , whose length is K, is used to indicate the existence of the conditional frame. The mask is set sequentially in time order, and at most three elements are assigned to 1 to represent the three conditional frames. Subsequently, the clean latent variable encoded by the image encoder Replace the corresponding noise latent variables Instead of concatenating additional channels into the input, the input latent variable is constructed as Since there is no need to predict the observed conditional frame, the diffusion loss function is specifically as follows: ; in, represents the diffusion loss function, It is the UNet denoiser in SVD; represents the expectation operation; is the i-th mask element, is the multiplication operation, is the i-th input latent variable, is the noise variance, Denotes the denoiser input The output features obtained after processing, Encode the potential features for the i-th frame image, Denotes the squared difference operation. When the replaced latent variables have sufficient prior information, the world model can fully capture the state of the surrounding instances and, through iterative expansion, predict a more coherent and reasonable long-term future.

[0061] Furthermore, since the difference between two adjacent frames provides rich motion patterns, this embodiment introduces an additional supervision mechanism when constructing the dynamic loss function to promote the learning of key area dynamics. First, a dynamic perception weight is introduced , which highlights areas where the predictions have inconsistent motion compared to the true values: ; in, is the dynamic perception weight of the i-th frame, Denotes the denoiser input The output features obtained after processing, is the i-th input latent variable, is the i-1th input latent variable, Denotes the denoiser input The output features obtained after processing. Encode the potential features for the i-th frame image, Encode the potential features for the i-1th frame image. It should also be noted that in order to ensure numerical stability, the dynamic perception weights are normalized within each video clip. .

[0062] Based on the above dynamic perception weights, given the causal relationship of future predictions, that is, the subsequent frames should follow the previous frames, the dynamic loss function is defined by penalizing the subsequent frames of each adjacent frame pair. The formula is as follows: ; in, is the dynamic loss function, Indicates stopping the gradient operation. By adaptively reweighting the standard diffusion loss, the dynamic loss function can improve the learning efficiency in dynamic areas.

[0063] Finally, given that structural details (such as edges and textures) primarily reside in high-frequency components, it is necessary to perform a Fourier transform on each latent feature in order to generate a structure-preserving loss function based on the Fourier-transformed latent features and the frame-level mask. In some embodiments, a two-dimensional discrete Fourier transform and an inverse discrete Fourier transform are performed on each latent feature, as shown in the following formula: ; in, The potential features of the image encoding of the i-th frame after Fourier transform are: FFT and IFFT represent two-dimensional discrete Fourier transform and inverse discrete Fourier transform, respectively. H is an ideal high-pass filter used to cut off low-frequency components below a certain threshold. The Fourier transform is applied independently to the potential features of the image encoding of the i-th frame. Similarly, we can also get the predicted latent features Based on the extracted high-frequency features, the structure preservation loss function is as follows: ; It can be understood that the structure-preserving loss function is used to minimize the difference in high-frequency features between the prediction and the true value, thereby retaining more structural information.

[0064] In summary, this embodiment achieves physical consistency and high-frequency detail fidelity for short-term and long-term predictions by constructing three loss functions, injecting a customized video diffusion model into the third-order motion prior, and combining dynamic weighting with frequency domain structure preservation loss.

[0065] Based on the above embodiments, in some embodiments, the initial world model is fine-tuned according to the diffusion loss function, the dynamic loss function, and the structure preservation loss function, including: S121: Determine a total loss function based on the diffusion loss function, the dynamic loss function, and the structure preservation loss function.

[0066] S122: Perform preliminary fine-tuning of the initial world model based on the environment dynamics modeling framework and the total loss function.

[0067] S123: Introduce action condition information and freeze the pre-trained weights of the initial world model after preliminary fine-tuning.

[0068] S124: Add low-rank adaptation and projection layers to all attention blocks of the U-network of the initially fine-tuned initial world model, and fine-tune the initial world model again with a preset learning rate to generate a world model.

[0069] In order to balance generation quality and training efficiency, this embodiment adopts a staged fine-tuning scheme for the initial world model. First, based on the diffusion loss function, dynamic loss function and structure preservation loss function, the total loss function is determined as follows: ; in, is the total loss function, are weighed respectively.

[0070] Specifically, the first stage is used to learn high-fidelity future predictions. This stage disregards action-related conditional input and trains the parameters of all UNet components in the SVD model using a combined total loss function based on the Environmental Dynamic Modeling (EDM) framework. Dynamic priors are randomly sampled in different orders with increasing probability, i.e., the probability of the 0, 1, 2, and 3 conditional frames is 1 / 15, 2 / 15, 4 / 15, and 8 / 15, respectively.

[0071] The second phase is dedicated to motion control learning. This phase introduces action condition information, freezes pre-trained weights, and adds low-rank adaptation (LoRA) and projection layers to all attention blocks in the UNet network. The weight parameters in the SVD are continuously fine-tuned relative to a preset learning rate. After fine-tuning, a world model is generated and all model parameters are saved for subsequent use.

[0072] In this way, fine-tuning of the initial world model is achieved, taking into account both generation quality and training efficiency, making the generated world model more stable and accurate.

[0073] Based on the above embodiments, in some embodiments, generating a reward function of a decision model based on a world model includes: S131: Perform multiple rounds of denoising based on randomly sampled noise of the same conditional frames and actions to generate an initial reward function.

[0074] Here, the initial reward function is defined as the exponential of the mean negative conditional variance.

[0075] S132: Generate future motion trajectory parameters based on the current driving state, and map the future motion trajectory parameters to trajectory points.

[0076] S133: Calculate the Euclidean distance between the future trajectory point and the actual human driving trajectory.

[0077] S134: Determine the degree of difference between the future driving prediction and the actual driving process based on the potential features after Fourier transformation.

[0078] S135: Generate a reward function based on the initial reward function, the Euclidean distance, and the degree of difference.

[0079] To achieve closed-loop optimization within the world model for the autonomous driving decision-making system, a reward predictor is required to complete the final construction of the world model. In this embodiment, considering that the world model's out-of-distribution conditions will lead to increased diversity in generated observations, the prediction uncertainty of the world model itself is used as a source of reward, estimating uncertainty through conditional variance.

[0080] Specifically, in order to achieve a reliable approximation, the and actions Multiple rounds of denoising are performed on the random sampling noise of , and the initial reward function is defined as the exponential of the average negative conditional variance. The formula is as follows: ; ; in, is the initial reward function, To average all potential prediction feature values ​​within the video clip; Indicates that in the conditional frame and actions The denoiser pair is the i-th input latent variable under the generation condition of The output features obtained after processing; Indicates the Based on the above formula, unfavorable actions with greater uncertainty will result in lower rewards.

[0081] Although uncertainty-based rewards can eliminate the need for additional detection and labeling, the reward signal has limited supervision capabilities and cannot provide complete and reliable guidance information for decision optimization. Considering that the driving behavior information of human drivers will also be collected synchronously during the driving data collection, the reward signal can be supplemented by introducing expert behavior supervision information. Specifically, the decision model generates future motion trajectory parameters based on the current driving state, and then maps the trajectory parameters into a series of trajectory points through the trajectory generation module. The Euclidean (L2) distance between the future trajectory points and the real human driving trajectory is calculated. The formula is as follows: ; in, is the Euclidean distance, Future trajectory points generated by the decision model, Real human driving trajectory.

[0082] Furthermore, considering the inconsistency between the output action of the decision model and the actual driving action conditions, this embodiment also determines the degree of difference between the future driving prediction and the actual driving process based on the potential features after Fourier transformation. The specific formula is as follows: ; in, For the degree of difference, It is the real future feature sequence obtained by Fourier transform processing of the real future video.

[0083] Finally, the reward function is generated based on the initial reward function, Euclidean distance, and difference degree. The formula is as follows: ; in, is the reward function.

[0084] In summary, this embodiment automatically generates implicit rewards using world model prediction uncertainty. Combined with expert driving trajectory imitation rewards and frequency domain structure consistency, it provides reliable supervision signals for policy optimization without manual or external detector labeling, significantly reducing data collection and labeling overhead.

[0085] Based on the above embodiments, in some embodiments, closed-loop training of a decision model based on target video data and a world model includes: S141: Generate trajectory parameters according to the current observation state through the decision model.

[0086] S142: Generate a corresponding motion trajectory according to the trajectory parameters.

[0087] S143: Input the motion trajectory as action condition information and the current image frame into the world model to predict the potential feature sequence.

[0088] S144: Determine the final reward value based on the potential feature sequence and the reward function. Return to step S136 until the maximum capacity of the experience pool is reached.

[0089] The experience pool stores multiple sets of samples, each consisting of the current observation state, the corresponding trajectory parameters, the corresponding final reward value, and the corresponding next observation state.

[0090] S145: Randomly extract samples from the experience pool according to the preset batch size.

[0091] S146: Calculate and update the decision strategy network, strategy evaluation network, and value network based on each sample until the update number threshold is reached, and output the decision model.

[0092] In order to match the trajectory action conditions required in the future prediction process of the world model, this embodiment builds a decision model based on parameterized skills. Specifically, the action space output by the decision model is specifically the trajectory planning parameters , which corresponds to the trajectory end boundary condition, that is, the vehicle driving state at the end of the trajectory, including the lateral position corresponding to the end of the vehicle , heading angle ,speed and acceleration The input state space of the decision model corresponds to the video frame of the vehicle at the current moment. At the same time, the trajectory starting boundary conditions can be obtained according to the current state of the vehicle , which is the starting position of the vehicle , heading angle ,speed and acceleration .

[0093] It should be noted that during the vehicle's driving, the decision model is based on the vehicle's current state at each moment. Select trajectory planning parameters Therefore, the mapping relationship between driving state and action can be expressed as: ; in, For driving strategy.

[0094] Furthermore, the motion planning method is used to generate the motion trajectory according to the trajectory parameters. The formula is as follows: ; in, represents the t-th trajectory point in the motion trajectory X, and T represents the trajectory duration. In a specific implementation, a trajectory point can be set every 6 frames. That is, for a future video prediction period of 25 consecutive frames, a motion trajectory of length T = 4 needs to be generated.

[0095] It should also be noted that this solution is based on the Actor-Critic architecture to build a decision model training framework based on reinforcement learning. The framework contains two main networks to be learned: Actor and Critic. Specifically, Actor corresponds to the decision strategy network , used to select the corresponding trajectory parameter output according to the driving state of the vehicle at each moment; Critic corresponds to the strategy evaluation network , which is used to evaluate the quality of the parameters selected for the planning strategy.

[0096] It is worth noting that the actor network and the critic network have the same state input and main network structure, but due to the different output forms, different output coding layers need to be designed. In this embodiment, the network model is designed based on the convolutional layer and fully connected layer structure, where the first three layers of the main network are convolutional coding layers, which are used to extract state features through convolution operations. The middle two layers are fully connected layers, which are used to fuse and reduce the dimension of image features. The output layer is also a fully connected layer. The actor network output layer outputs trajectory parameters, and the critic network output layer encodes the fused features into a value scalar, which is used to estimate the expected return of the action taken by the planning strategy. The actor network and critic network constructed in the above manner are more reasonable and can perform prediction tasks more accurately.

[0097] Furthermore, the motion trajectory is used as action condition information and the current image frame is input into the world model to predict the potential feature sequence , and determines the final reward value based on the potential feature sequence and reward function. Specifically, the policy evaluation network critic measures the performance of the planning strategy by regressing the expected return; assuming is the strategy evaluation parameter, represents the trajectory parameters of the strategy output in a planning cycle, then The current video frame of the vehicle is Conditional planning strategy selection parameters The environment reward when generating the trajectory. Then return to step S136 until the maximum capacity of the experience pool is reached. It should be noted that the experience pool stores multiple sets of samples, which are composed of the current observation state, the corresponding trajectory parameters, the corresponding final reward value and the corresponding next observation state, that is, .

[0098] Finally, samples are randomly drawn from the experience pool according to the preset batch size. In this embodiment, there is no limit on the preset batch size. The decision strategy network, strategy evaluation network, and value network are calculated and updated based on each sample. Specifically, in order to achieve network updates and improve the exploration ability of subsequent strategy online training, and to reduce the overestimation of strategy evaluation values ​​and improve model stability, in this embodiment, two critic networks with the same structure are maintained. and , and set a parameter , target value network with the same structure as the Critic network , with minimizing the Bellman residual as the optimization objective, the formula is as follows: ; in, is the optimization function of the decision strategy network; is the state-action evaluation value calculated by the i-th strategy evaluation network for the trajectory parameter a selected at the current video frame I under the conditional frame c; is the discount factor, is the entropy regularization coefficient; is the logarithmic probability of the policy network choosing action a' in state I', is the target value network's valuation of the next state I', whose goal is to fit the expected Q value of the strategy. The formula is as follows: ; in, is the optimization function of the value network; Indicates taking the minimum evaluation value output by the two Critic networks.

[0099] Furthermore, the goal of the planning strategy network Actor is to maximize the expected cumulative reward of entropy regularization, and the following optimization objective is used to update the network parameters. The formula is as follows: ; in, is the optimization function of the policy evaluation network, For planning strategy network parameters, entropy term Used to encourage strategies to remain random.

[0100] In summary, the decision-making policy network, policy evaluation network, and value network are updated separately. This process is repeated until the update threshold is reached, outputting the final decision model. In this embodiment, the decision model is trained interactively in a virtual environment using a parameterized trajectory-planning action space. This collaborative optimization of environmental cognition and policy evolution is achieved through a multi-dimensional reward and experience replay mechanism.

[0101] Figure 5 Schematic diagram of an end-to-end autonomous driving system based on a world model provided by an embodiment of the present invention. Based on the above embodiments, in order to realize the autonomous driving control of the vehicle based on the trained world model and decision model, in some embodiments, such as Figure 5 As shown, after the decision model is closed-loop trained based on the target video data and the world model, the method further includes: S161: Loading network parameters of the image encoding modules of the decision model and the world model to the onboard edge server via Ethernet communication technology.

[0102] like Figure 5 As shown, the decision model is iteratively trained in the cloud server through interactive exploration with the world model. In this embodiment, the target server is specifically the vehicle-mounted edge server. Therefore, after training is complete, the cloud server loads the decision model and the network parameters of the image encoding portion of the world model to the vehicle-mounted edge server via Ethernet communication technology for execution.

[0103] Figure 6 Schematic diagram of a motion trajectory generation system provided by an embodiment of the present invention. Figure 6 As shown, the system includes: Cloud server 5 is used to collect target video data and construct a target video dataset based on the target video data; generate an initial world model based on the target video dataset and a pre-trained diffusion generative model, and construct a diffusion loss function, a dynamic loss function, and a structure-preserving loss function based on a third-order motion prior based on the target video dataset; fine-tune the initial world model based on the diffusion loss function, the dynamic loss function, and the structure-preserving loss function to generate a fine-tuned world model; generate a reward function for a decision model based on the world model, and conduct closed-loop training of the decision model based on the target video data and the world model; and deploy the world model and the decision model to the target server; The target server 6 is configured to generate a motion trajectory based on the world model and the decision model.

[0104] In this embodiment, the cloud server generates an initial world model using a target video dataset and a pre-trained diffusion generative model, and fine-tunes the initial world model using three different loss functions: a diffusion loss function based on third-order motion priors, a dynamic loss function, and a structure-preserving loss function, thereby achieving physical consistency and high-frequency detail fidelity for short-term and long-term predictions. On this basis, the reward function is automatically generated using the uncertainty of the world model prediction without the need for manual or external detector labeling, thereby improving training efficiency. The decision model is trained in a closed-loop based on the target video data and the world model, thereby achieving coordinated optimization of environmental cognition and strategy evolution. Finally, the trained world model and decision model are integrated into the target server so that the target server can generate motion trajectories based on the world model and the decision model, thereby achieving closed-loop control of perception-decision-motion execution, with low latency, high robustness, and easy scalability, thereby improving the safety of the autonomous driving system.

[0105] In some embodiments, the target server generates a motion trajectory based on the world model and the decision model, including: S162: Receive environmental perception data transmitted from the vehicle end via the data bus, and perform data preprocessing on the environmental perception data.

[0106] Among them, environmental perception data is collected in real time by the vehicle through on-board sensors.

[0107] S163: An image encoding module based on a world model converts image data in the environmental perception data into latent state features.

[0108] S164: Input the latent state features into the decision model for action reasoning to generate a motion trajectory.

[0109] Specifically, when the target server is an onboard edge server, upon receiving environmental perception data from the vehicle, the onboard edge server first performs preprocessing on the raw data, such as format conversion, through a data preprocessing module. The image data in the environmental perception data is then sent to an image encoding module for processing into low-dimensional latent state features, which are then passed to a decision model for action reasoning. When performing action reasoning, action instructions are generated based on the motion trajectory and sent to the vehicle control drive module on the vehicle side, which controls the vehicle to perform the corresponding driving action according to the action instructions.

[0110] In summary, this embodiment integrates the trained world model and decision model into the vehicle-side edge computing platform, supporting real-time perception-decision-motion execution closed-loop control with low latency, high robustness, and easy scalability, thus realizing autonomous driving control.

[0111] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0112] Figure 7 Schematic diagram of a world model driven decision model training device provided by an embodiment of the present invention. Figure 7 As shown, the device includes: The acquisition module 10 is used to acquire target video data and construct a target video data set according to the target video data.

[0113] The generation module 11 is used to generate an initial world model according to the target video dataset and the pre-trained diffusion generation model, and to construct a diffusion loss function, a dynamic loss function and a structure preservation loss function based on the third-order motion prior according to the target video dataset.

[0114] The fine-tuning module 12 is used to fine-tune the initial world model according to the diffusion loss function, the dynamic loss function and the structure preservation loss function to generate a fine-tuned world model.

[0115] The training module 13 is used to generate a reward function of the decision model based on the world model, and to perform closed-loop training of the decision model based on the target video data and the world model.

[0116] In some embodiments, the acquisition module 10 includes: A first acquisition submodule is configured to acquire forward-looking image information during vehicle travel according to a first preset frequency; A second acquisition submodule, configured to acquire vehicle posture data and vehicle control information according to a second preset frequency; a first aggregation submodule, configured to aggregate the forward-looking image information according to a preset number of frames to generate a plurality of target video data; The synchronization submodule is used to align the time steps of each target video data with the corresponding vehicle posture data and vehicle control information to generate a target video dataset.

[0117] In some embodiments, the generating module 11 includes: The first generation submodule is used to compress each frame image in the target video dataset through a pre-trained autoencoder to generate potential features corresponding to each frame image; The second generation submodule is used to encode the prediction condition information in the target video data set to generate a conditional input; wherein the prediction condition information at least includes a conditional frame image, a trajectory action sequence, an action sequence number, and a sample sequence number; a third generation submodule for adding Gaussian noise to the latent features to generate a noisy latent sequence and labeling the conditional frames to generate an initial world model; The conditional frame of the noise potential sequence corresponding to the initial generation period is marked as the first frame, and the conditional frame of the noise potential sequence corresponding to the non-initial generation period is marked as the first three frames.

[0118] In some embodiments, the second generation submodule includes: a first encoding submodule, configured to encode the conditional frame image through an autoencoder and a contrastive learning image-text encoder to output image conditional features and semantic conditional features, respectively; The second encoding submodule is used to encode the trajectory action sequence, action sequence number and sample sequence number respectively through the time step encoding function to generate the conditional features after the trajectory action sequence, action sequence number and sample sequence number are encoded; A splicing submodule is used to splice the conditional features encoded by the action sequence number and the sample sequence number to generate a spliced ​​conditional feature; The second aggregation submodule is used to aggregate image conditional features, semantic conditional features, conditional features after trajectory action sequence encoding, and splicing conditional features to generate conditional input.

[0119] In some embodiments, the generating module 11 includes: A first construction submodule is used to construct a frame-level mask; wherein the mask is set in a time order, and the number of elements that are 1 is not greater than 3; a fourth generation submodule for determining an input latent variable based on the frame-level mask, the latent features, and the noise latent sequence, and generating a diffusion loss function based on the frame-level mask, the latent features, the noise latent sequence, and the input latent variable; a fifth generation submodule, configured to generate a dynamic perception weight corresponding to each target video data according to the corresponding input latent variables and latent features, and to generate a dynamic loss function according to each dynamic perception weight; The sixth generation submodule is used to perform Fourier transform on each potential feature and generate a structure-preserving loss function based on each potential feature after Fourier transform and the frame-level mask.

[0120] In some embodiments, the sixth generation submodule includes: The Fourier transform module is used to perform a two-dimensional discrete Fourier transform and an inverse discrete Fourier transform on each potential feature to perform a Fourier transform on each potential feature.

[0121] In some embodiments, the fine-tuning module 12 includes: A first determination submodule is used to determine a total loss function based on a diffusion loss function, a dynamic loss function, and a structure preservation loss function; The first fine-tuning submodule is used to perform preliminary fine-tuning on the initial world model according to the environment dynamics modeling framework and the total loss function; A submodule is introduced to introduce action condition information and freeze the pre-trained weights of the initial world model after preliminary fine-tuning; The second fine-tuning submodule is used to add low-rank adaptation and projection layers to all attention blocks of the U-shaped network of the initial world model after preliminary fine-tuning, and fine-tune the initial world model again with a preset learning rate to generate a world model.

[0122] In some embodiments, the training module 13 includes: A denoising submodule, which performs multiple rounds of denoising based on randomly sampled noise of the same conditional frame and action to generate an initial reward function; where the initial reward function is defined as the exponential of the mean negative conditional variance; a seventh generation submodule, configured to generate future motion trajectory parameters based on the current driving state, and map the future motion trajectory parameters into trajectory points; The distance determination submodule is used to calculate the Euclidean distance between the future trajectory points and the real human driving trajectory; The second determination submodule is used to determine the degree of difference between the future driving prediction and the actual driving process based on the potential features after Fourier transformation; The reward function generation submodule is used to generate a reward function based on the initial reward function, Euclidean distance and difference degree.

[0123] In some embodiments, the training module 13 includes: The trajectory parameter generation submodule is used to generate trajectory parameters according to the current observation state through the decision model; The motion trajectory generation submodule is used to generate the corresponding motion trajectory according to the trajectory parameters; The prediction submodule is used to input the motion trajectory as action condition information and the current image frame into the world model to predict the potential feature sequence; The third determination submodule is used to determine the final reward value based on the potential feature sequence and the reward function; trigger the trajectory parameter generation submodule until the maximum capacity of the experience pool is reached; wherein the experience pool stores multiple sets of samples, each of which consists of the current observation state, the corresponding trajectory parameter, the corresponding final reward value, and the corresponding next observation state; The extraction submodule is used to randomly extract samples from the experience pool according to the preset batch size; The update submodule is used to calculate and update the decision strategy network, strategy evaluation network and value network according to each sample until the update number threshold is reached and the decision model is output.

[0124] In some embodiments, further comprising: The deployment module is used to load the network parameters of the image encoding module of the decision model and the world model to the vehicle edge server through Ethernet communication technology.

[0125] In some embodiments, the target server generates a motion trajectory based on the world model and the decision model, specifically by receiving environmental perception data transmitted from the vehicle side through a data bus and performing data preprocessing on the environmental perception data; wherein, the environmental perception data is collected in real time by the vehicle side through on-board sensors; an image encoding module based on the world model converts image data in the environmental perception data into latent state features; the latent state features are input into the decision model for action reasoning to generate a motion trajectory.

[0126] In some embodiments, the target server also generates action instructions based on the motion trajectory; and sends the action instructions to the vehicle control drive module on the vehicle side, so that the vehicle control drive module controls the vehicle to perform corresponding driving actions according to the action instructions.

[0127] For the description of the features in the embodiment corresponding to the world model-driven decision model training device, please refer to the relevant description of the embodiment corresponding to the world model-driven decision model training method, which will not be repeated here.

[0128] An embodiment of the present invention also provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned world model-driven decision model training method embodiments.

[0129] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned world model-driven decision model training method embodiments when running.

[0130] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0131] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned world model-driven decision model training method embodiments.

[0132] An embodiment of the present invention also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned world model-driven decision model training method embodiments.

[0133] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0134] The above is a detailed introduction to the world model-driven decision model training method, system, device and product provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, the present invention can also be improved and modified, and these improvements and modifications also fall within the scope of protection of the present invention.

Claims

1. A world model driven decision model training method, characterized in that: include: Collecting target video data, and constructing a target video dataset based on the target video data; Generate an initial world model according to the target video dataset and a pre-trained diffusion generative model, and construct a diffusion loss function, a dynamic loss function, and a structure preservation loss function based on a third-order motion prior according to the target video dataset; Fine-tuning the initial world model according to the diffusion loss function, the dynamic loss function, and the structure preservation loss function to generate a fine-tuned world model; A reward function of a decision model is generated according to the world model, and the decision model is trained in a closed loop according to the target video data and the world model.

2. The world model driven decision model training method according to claim 1, characterized in that: Collect target video data and construct a target video dataset based on the target video data, including: collecting forward-view image information during vehicle travel according to a first preset frequency; collecting vehicle posture data and vehicle control information according to a second preset frequency; Aggregating the forward-looking image information according to a preset number of frames to generate a plurality of target video data; Time-step alignment is performed on each of the target video data with the corresponding vehicle posture data and the vehicle control information to generate the target video dataset.

3. The world model driven decision model training method according to claim 1, characterized in that: Generating an initial world model according to the target video dataset and the pre-trained diffusion generation model, including: Compressing each frame image in the target video dataset through a pre-trained autoencoder to generate potential features corresponding to each frame image; Encoding the prediction condition information in the target video data set to generate a conditional input; wherein the prediction condition information at least includes a conditional frame image, a trajectory action sequence, an action sequence number, and a sample sequence number; Adding Gaussian noise to the latent features to generate a noise latent sequence and labeling conditional frames to generate the initial world model; The conditional frame of the noise potential sequence corresponding to the initial generation period is marked as the first frame, and the conditional frame of the noise potential sequence corresponding to the non-initial generation period is marked as the first three frames.

4. The world model driven decision model training method according to claim 3, characterized in that: Encoding the prediction condition information in the target video dataset to generate a conditional input includes: Encoding the conditional frame image by the autoencoder and the contrastive learning image-text encoder to output image conditional features and semantic conditional features respectively; Encoding the trajectory action sequence, the action sequence number, and the sample sequence number respectively through a time step encoding function to generate conditional features after encoding the trajectory action sequence, the action sequence number, and the sample sequence number respectively; Splicing the conditional features encoded by the action sequence number and the sample sequence number to generate a spliced ​​conditional feature; The image conditional features, the semantic conditional features, the conditional features after the trajectory action sequence is encoded, and the splicing conditional features are aggregated to generate the conditional input.

5. The world model driven decision model training method according to claim 3, characterized in that: According to the target video dataset, a diffusion loss function, a dynamic loss function and a structure preservation loss function based on a third-order motion prior are constructed, including: Constructing a frame-level mask; wherein the mask is set in time order, and the number of elements that are 1 is not greater than 3; determining an input latent variable based on the frame-level mask, the latent features, and the noise latent sequence, and generating the diffusion loss function based on the frame-level mask, the latent features, the noise latent sequence, and the input latent variable; Generating a dynamic perception weight corresponding to each target video data according to the corresponding input latent variable and the latent feature, and generating the dynamic loss function according to each dynamic perception weight; Performing Fourier transform on each of the potential features, and generating the structure preservation loss function according to each of the potential features after Fourier transform and the frame-level mask.

6. The world model driven decision model training method according to claim 5, characterized in that: Performing Fourier transform on each of the potential features includes: Performing a two-dimensional discrete Fourier transform and an inverse discrete Fourier transform on each of the potential features to perform a Fourier transform on each of the potential features.

7. The world model driven decision model training method according to claim 5, characterized in that: Fine-tuning the initial world model according to the diffusion loss function, the dynamic loss function, and the structure preservation loss function, comprising: Determining a total loss function based on the diffusion loss function, the dynamic loss function, and the structure preservation loss function; Performing preliminary fine-tuning on the initial world model according to the environment dynamics modeling framework and the total loss function; Introducing action condition information and freezing the pre-trained weights of the initial world model after preliminary fine-tuning; Low-rank adaptation and projection layers are added to all attention blocks of the U-shaped network of the initial world model after preliminary fine-tuning, and the initial world model is fine-tuned again at a preset learning rate to generate the world model.

8. The world model driven decision model training method according to claim 5, characterized in that: Generating a reward function of a decision model according to the world model includes: Perform multiple rounds of denoising based on randomly sampled noise of the same conditional frame and action to generate an initial reward function; wherein the initial reward function is defined as the exponential of the mean negative conditional variance; generating future motion trajectory parameters based on the current driving state, and mapping the future motion trajectory parameters into trajectory points; Calculate the Euclidean distance between the future trajectory points and the real human driving trajectory; Determining the degree of difference between the future driving prediction and the actual driving process based on each of the potential features after Fourier transformation; The reward function is generated according to the initial reward function, the Euclidean distance, and the degree of difference.

9. The world model driven decision model training method according to claim 8, characterized in that: A closed-loop training decision model is performed based on the target video data and the world model, including: Generate trajectory parameters according to the current observation state through the decision model; Generate a corresponding motion trajectory according to the trajectory parameters; Inputting the motion trajectory as action condition information and the current image frame into the world model to predict a potential feature sequence; determining a final reward value according to the potential feature sequence and the reward function; Returning to the step of generating trajectory parameters according to the current observation state through the decision model until the maximum capacity of the experience pool is reached; wherein the experience pool stores multiple sets of samples, each consisting of the current observation state, the corresponding trajectory parameters, the corresponding final reward value, and the corresponding next observation state; Randomly extracting the samples from the experience pool according to a preset batch size; The decision strategy network, strategy evaluation network and value network are calculated and updated according to each of the samples until a threshold of update times is reached, and the decision model is output.

10. The world model driven decision model training method according to any one of claims 1 to 9, characterized in that: After the decision model is closed-loop trained according to the target video data and the world model, the method further includes: The network parameters of the image encoding module of the decision model and the world model are loaded to the vehicle edge server through Ethernet communication technology.

11. A motion trajectory generation system, characterized in that: include: A cloud server is used to collect target video data and construct a target video dataset based on the target video data; An initial world model is generated based on the target video dataset and a pre-trained diffusion generative model, and a diffusion loss function, a dynamic loss function, and a structure-preserving loss function based on a third-order motion prior are constructed based on the target video dataset; the initial world model is fine-tuned based on the diffusion loss function, the dynamic loss function, and the structure-preserving loss function to generate a fine-tuned world model; a reward function of a decision model is generated based on the world model, and a decision model is closed-loop trained based on the target video data and the world model; and the world model and the decision model are deployed to a target server; The target server is configured to generate a motion trajectory based on the world model and the decision model.

12. The motion trajectory generation system according to claim 11, characterized in that: The target server generates a motion trajectory based on the world model and the decision model, including: Receiving environmental perception data transmitted by the vehicle end through a data bus and performing data preprocessing on the environmental perception data; wherein the environmental perception data is collected in real time by the vehicle end through on-board sensors; An image encoding module based on the world model converts image data in the environmental perception data into latent state features; The potential state features are input into the decision model for action reasoning to generate the motion trajectory.

13. The motion trajectory generation system according to claim 11, characterized in that: After the target server generates a motion trajectory based on the world model and the decision model, the method further includes: generating an action instruction according to the motion trajectory; The action instruction is sent to the vehicle control drive module on the vehicle side, so that the vehicle control drive module controls the vehicle to perform the corresponding driving action according to the action instruction.

14. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the world model-driven decision model training method according to any one of claims 1 to 10 when executing the computer program.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the world model-driven decision model training method according to any one of claims 1 to 10 are implemented.

Citation Information

Cited By

  • High-fidelity lightweight world model construction method for end-to-end automatic driving test

    CN120909949A

  • High-fidelity lightweight world model construction method for end-to-end autonomous driving test

    CN120909949B

  • Robotaxi-oriented VLA-world model fusion automatic driving system and closed-loop training platform

    CN121212218A

  • Robotaxi-oriented vla-world model fusion automatic driving system and closed-loop training platform

    CN121212218B

  • End-to-end automatic driving long tail identification method based on comparative learning pre-training

    CN121527727A