Model training method, video prediction method, device, equipment and storage medium

By using the initial video frame and trajectory data to determine the scene context and vehicle movement instructions, and training the video prediction model, the problem that traditional video prediction methods rely on labeled data sets and layout information is solved, and automatic labeling and video prediction of large-scale labeled unlabeled data sets are realized.

CN117764815BActive Publication Date: 2025-06-17SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311796343.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-17
Estimated Expiration
2043-12-25

AI Technical Summary

Technical Problem

Traditional video prediction methods rely on high-quality labeled data sets, are difficult to scale to large-scale labeled data sets, and tend to generate videos based on layout information.

Method used

By using the initial video frames and scene description metadata collected by the target vehicle to determine the scene context, combining the trajectory data to determine the vehicle movement instructions, and using control text to train the video prediction model, automatic labeling and video prediction of annotated large-scale data sets are achieved.

Benefits of technology

Automatic labeling of natural language instructions and scene contexts of unlabeled large-scale data sets is realized, which improves the universality and generalization of the video prediction model, and can generate training samples more freely, solving the problem that the video prediction model is limited by the annotation data set and layout information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117764815B_ABST
    Figure CN117764815B_ABST
Patent Text Reader

Abstract

The present invention discloses a model training method, a video prediction method, an apparatus, a device, and a storage medium. The method includes: determining the scene context of an initial video frame by using the initial video frame collected by a target vehicle and / or the scene description metadata of the initial video frame; determining a vehicle movement instruction of the target vehicle by using the initial video frame and / or the trajectory data of the target vehicle corresponding to the initial video frame; training an initial model by using the initial video frame and a corresponding control text to obtain a video prediction model, wherein the control text includes the scene context and the vehicle movement instruction. The technical solution of the embodiment of the present invention realizes the automatic annotation of natural language instructions and scene context for an unlabeled large-scale data (video frame) set, enables the model to output a frame image of future picture prediction according to the existing picture and the control text based on natural language, and solves the problem that the video prediction model is restricted by the labeled data set and layout information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent driving technology, and in particular, to a model training method, a video prediction method, a device, a device and a storage medium. Background Art

[0002] In recent years, scene generation in the field of autonomous driving has received extensive research attention. NeRF (Neural Radiance Field) can be used to construct a simulated environment based on real data. For example, by borrowing compliant objects from other scenes, rendering a video from a given picture, or constructing digital twins of scenes and agents in a dataset for closed-loop simulation and sensor screen rendering. Diffusion models can also be used to generate videos, and diffusion models can use layout information to control the rendered video. For example, the diffusion model GeoDiffusion can use two-dimensional bounding boxes to render videos, and the neural field diffusion model NeuralField-LDM and the conditional generation model BEVGen can use bird's-eye view segmentation maps to render videos.

[0003] However, in traditional video prediction methods, a high-quality manually annotated dataset is usually required to train a video prediction model, and larger-scale unannotated video data cannot be used to train the model, making it difficult to scale to larger datasets. Moreover, the video prediction model usually tends to generate videos based on given layout information. Summary of the Invention

[0004] The present invention provides a model training method, a video prediction method, a device, a device and a storage medium to solve the problem that the video prediction model is restricted by the annotated dataset and layout information.

[0005] In a first aspect, the present invention provides a model training method, including:

[0006] Determining the scene context of the initial video frame by using the initial video frame collected by the target vehicle and / or the scene description metadata of the initial video frame;

[0007] Determining the vehicle movement instruction of the target vehicle by using the initial video frame and / or the trajectory data of the target vehicle corresponding to the initial video frame;

[0008] Training an initial model by using the initial video frame and the corresponding control text to obtain a video prediction model, where the control text includes the scene context and the vehicle movement instruction, and the video prediction model is used to output predicted video frames.

[0009] In a second aspect, the present invention provides a video prediction method, including:

[0010] Obtain at least one conditional video frame, a target scene context, and a target movement instruction of a current vehicle, where the target scene context is the scene context of the current scene corresponding to the conditional video frame;

[0011] Input the encoding of the at least one conditional video frame and the encoding of a target control text into a preset video prediction model to obtain a predicted video, where the preset video prediction model is obtained by using the model training method described in the first aspect above, and the target control text includes the target scene context and / or the target movement instruction.

[0012] In a third aspect, the present invention provides a model training device, including:

[0013] A scene context determination module, configured to determine the scene context of an initial video frame by using the initial video frame collected by a target vehicle and / or the scene description metadata of the initial video frame;

[0014] A vehicle movement instruction determination module, configured to determine the vehicle movement instruction of the target vehicle by using the initial video frame and / or the trajectory data of the target vehicle corresponding to the initial video frame;

[0015] A model determination module, configured to train an initial model by using the initial video frame and a corresponding control text to obtain a video prediction model, where the control text includes the scene context and the vehicle movement instruction, and the video prediction model is used to output predicted video frames.

[0016] In a fourth aspect, the present invention provides a video prediction device, including:

[0017] A conditional information acquisition module, configured to obtain at least one conditional video frame, a target scene context, and a target movement instruction of a current vehicle, where the target scene context is the scene context of the current scene corresponding to the conditional video frame;

[0018] A predicted video determination module, configured to input the encoding of the at least one conditional video frame and the encoding of a target control text into a preset video prediction model to obtain a predicted video, where the preset video prediction model is obtained by using the model training method described in the first aspect above, and the target control text includes the target scene context and / or the target movement instruction.

[0019] In a fifth aspect, the present invention provides an electronic device, and the electronic device includes:

[0020] At least one processor;

[0021] And a memory communicatively connected to the at least one processor;

[0022] Among them, the memory stores a computer program that can be executed by at least one processor. The computer program is executed by at least one processor so that the at least one processor can execute the model training method of the first aspect described above, and / or execute the video prediction method of the second aspect described above.

[0023] In a sixth aspect, the present invention provides a computer-readable storage medium that stores computer instructions for causing a processor to implement the model training method of the first aspect described above when executed, and / or implement the video prediction method of the second aspect described above.

[0024] The model training solution provided by the present invention uses the initial video frame collected by the target vehicle and / or the scene description metadata of the initial video frame to determine the scene context of the initial video frame, uses the initial video frame and / or the trajectory data of the target vehicle corresponding to the initial video frame to determine the vehicle movement instruction of the target vehicle, and uses the initial video frame and the corresponding control text to train the initial model to obtain a video prediction model. Among them, the control text includes the scene context and the vehicle movement instruction, and the video prediction model is used to output a predicted video frame. By adopting the above technical solution, by using the initial video frame and / or the data of the initial video frame, automatic annotation of natural language instructions and scene context for an unlabeled large-scale data (video frame) set is realized, and by using the initial video frame and the corresponding control text to train the model, the model can output a frame image of the future picture prediction according to the existing picture and the control text based on natural language. This solution has higher universality and generalization ability, can generate training samples more freely, is not restricted by the camera settings and shooting scenes of small-scale public driving data sets, and solves the problem that the video prediction model is restricted by the labeled data set and layout information.

[0025] The video prediction solution provided by the present invention obtains at least one conditional video frame, a target scene context, and a target movement instruction of the current vehicle. Among them, the target scene context is the scene context of the current scene corresponding to the conditional video frame. The encoding of the at least one conditional video frame and the encoding of the target control text are input into a preset video prediction model to obtain a predicted video. Among them, the preset video prediction model is obtained by using the model training method described in the first aspect above, and the target control text includes the target scene context and / or the target movement instruction. By adopting the above technical solution, by inputting the encoding of the conditional video frame and the corresponding encoding of the target control text into the model, the predicted video can be obtained quickly and accurately.

[0026] It should be understood that the content described in this section is not intended to identify the key or important features of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become readily understood from the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0028] Figure 1 is a flowchart of a model training method provided in Embodiment 1 of the present invention;

[0029] Figure 2 is a schematic flowchart of generating a control text provided in Embodiment 1 of the present invention;

[0030] Figure 3 is a schematic diagram of the first-stage model training provided in Embodiment 1 of the present invention;

[0031] Figure 4 is a flowchart of a model training method provided in Embodiment 2 of the present invention;

[0032] Figure 5 is a schematic structural diagram of a timing inference module provided in Embodiment 2 of the present invention;

[0033] Figure 6 is a schematic structural diagram of a causal timing attention sub-module provided in Embodiment 2 of the present invention;

[0034] Figure 7 is a schematic diagram of the second-stage model training provided in Embodiment 2 of the present invention;

[0035] Figure 8 is a flowchart of a video prediction method provided in Embodiment 3 of the present invention;

[0036] Figure 9 is a schematic structural diagram of a model training device provided in Embodiment 4 of the present invention;

[0037] Figure 10 is a schematic structural diagram of a video prediction device provided in Embodiment 5 of the present invention;

[0038] Figure 11 is a schematic structural diagram of an electronic device provided in Embodiment 6 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0040] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In the description of the present invention, unless otherwise specified, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0041] Embodiment 1

[0042] Figure 1 The following is a flowchart of a model training method provided for Embodiment 1 of the present invention. This embodiment is applicable to the situation of training a video prediction model. This method can be executed by a model training device. The model training device can be implemented in the form of hardware and / or software. The model training device can be configured in an electronic device. The electronic device can be composed of two or more physical entities or one physical entity.

[0043] As Figure 1 shown, a model training method provided in Embodiment 1 of the present invention specifically includes the following steps:

[0044] S101. Determine the scene context of the initial video frame by using the initial video frame collected by the target vehicle and / or the scene description metadata of the initial video frame.

[0045] S102. Determine the vehicle movement instruction of the target vehicle by using the initial video frame and / or the trajectory data of the target vehicle corresponding to the initial video frame.

[0046] Figure 2 It is a schematic flowchart for generating control text. In this embodiment, as Figure 2 shown, the scene context of the unannotated initial video frame and the vehicle movement instruction can be automatically generated by using a vision-language model such as BLIP-2 (image-to-text model) and a model for classifying camera behavior based on video optical flow respectively. For the initial video frame with trajectory data, such as ground truth trajectory, and scene description metadata, a natural language model such as GPT (Generative Pre-Trained Transformer model) can be used to convert the metadata into natural language, and the vehicle behavior corresponding to each initial video frame can be determined from the trajectory by mathematical methods, and the vehicle movement instruction can be determined according to the behavior.

[0047] S103. Train the initial model by using the initial video frame and the corresponding control text to obtain a video prediction model, where the control text includes the scene context and the vehicle movement instruction, and the video prediction model is used to output a predicted video frame.

[0048] In this embodiment, the initial model can be the Stable Diffusion XL (hereinafter denoted as SDXL) model. The SDXL model is an image-to-image model for generating images according to the given text, which is implemented based on the diffusion model. In the process of training the SDXL model with image-text pairs (i.e., the initial video frame and the corresponding control text), the image and the text can be encoded by the encoder of VQVAE and the CLIP encoder respectively to obtain the feature map and the text encoding. The SDXL model can learn how to gradually denoise the noisy feature map under the control condition of the text encoding. The training objective is to make the image after each step of denoising as close as possible to the original image before adding noise at that step. After training, the SDXL model can gradually denoise the Gaussian noise map according to the given text, and finally generate a feature map that conforms to the description of the control text. The feature map can be converted into a driving perspective image that conforms to the description of the control text through the decoder of VQVAE.

[0049] The model training method provided by the embodiments of the present invention uses the initial video frames collected by the target vehicle and / or the scene description metadata of the initial video frames to determine the scene context of the initial video frames, uses the initial video frames and / or the trajectory data of the target vehicle corresponding to the initial video frames to determine the vehicle movement instructions of the target vehicle, and trains an initial model with the initial video frames and the corresponding control texts to obtain a video prediction model, where the control texts include the scene context and the vehicle movement instructions, and the video prediction model is used to output predicted video frames. The technical solution of the embodiments of the present invention uses the initial video frames and / or the data of the initial video frames to realize the automatic annotation of natural language instructions and scene context for an unlabeled large-scale data (video frame) set, and trains the model by using the initial video frames and the corresponding control texts, so that the model can output frame images for predicting future images according to the existing pictures and the control texts based on natural language. This solution has higher universality and generalization ability, can generate training samples more freely, is not restricted by the camera settings and shooting scenes of small-scale public driving data sets, and solves the problem that the video prediction model is restricted by the labeled data set and layout information.

[0050] Optionally, the training of the initial model with the initial video frames and the corresponding control texts to obtain a video prediction model includes: inputting the image encoding of at least one of the initial video frames and the text encoding of the corresponding control texts into the Stable Diffusion XL model to obtain a denoised image, and determining a loss function value according to the denoised image; determining an intermediate model according to the loss function value, and updating the intermediate model by using an attention mechanism to obtain a model to be trained; training the model to be trained with the image encodings of multiple initial video frames and the text encodings of the corresponding control texts to obtain a video prediction model. The advantage of this setting is that by using a large-scale and diverse data set to train the initial model in two stages, the obtained video prediction model has stronger zero-shot generalization ability than the existing driving video generation model, and the generated predicted video frames contain richer scenes and objects.

[0051] Specifically, the initial model can be trained in two stages. Figure 3 As shown in Figure 3As shown, first, the SDXL model can be trained multiple times using at least one pair of image-text encodings composed of the image encoding of the initial video frame and the text encoding of the corresponding control text. For each round of training, the denoised image generated by the SDXL model based on the text encoding can be obtained. According to the difference between the denoised image and the initial video frame, the value of the loss function can be obtained. If the value of the loss function is less than the preset value, it indicates that the first-stage training is completed, and an intermediate model (image generation model) is obtained. Next, an attention mechanism can be introduced into the intermediate model to process temporal information. For example, an attention mechanism for processing temporal information can be introduced into the Transformer model included in the intermediate model to obtain the model to be trained. During the second-stage training, multiple pairs of image-text encodings can be used to train the intermediate model multiple times, and based on the difference between all the generated denoised video frames and the initial video frames corresponding to them in the input at the corresponding time points, the value of the loss function and the video prediction model can be determined. Among them, the initial video frames corresponding to the image encodings in the multiple pairs of image-text encodings can be multiple adjacent video frames in the same video segment. Since the time intervals between the initial video frames corresponding to the image encodings in the multiple pairs of image-text encodings are very short, the text encodings of the control texts in the multiple pairs of image-text encodings can be the same. In particular, during the training process, the scene context or vehicle movement instructions can be discarded with a certain probability to cope with the situation where there is no control text input during application.

[0052] Exemplarily, the way to determine the value of the loss function includes:

[0053]

[0054] The diffusion model f in the SDXL model θ (where θ is the parameter to be trained) needs to predict the noise added to the original image (i.e., the initial video frame) during the noise addition process based on the result x after adding noise t , the given control condition c, and the current number of noise addition steps t. The training objective is to make the predicted noise addition value as close as possible to the actual noise addition ∈. Among them, the control condition c is the text encoding of the control text, represents the square of the second norm, and the process of adding noise to the image encoding during the noise reduction of the SDXL model is unknown.

[0055] Optionally, before training the model to be trained using the image encoding of a plurality of the initial video frames and the text encoding of the corresponding control text to obtain a video prediction model, it further includes: processing the trajectory data using a Fourier encoding layer to obtain a high-dimensional trajectory encoding, and encoding and projecting the high-dimensional trajectory encoding using a preset linear layer to obtain a projection encoding; wherein, training the model to be trained using the image encoding of a plurality of the initial video frames and the text encoding of the corresponding control text to obtain a video prediction model includes: training the model to be trained using the image encoding of a plurality of the initial video frames and the text encoding of the corresponding target control text to obtain a video prediction model, wherein the text encoding of the target text includes the text encoding of the scene context and the vehicle movement instruction, and the projection encoding. The advantage of this setting is that by introducing trajectory numerical information, the video prediction model can generate prediction video frames with more accurate video angles.

[0056] Specifically, the trajectory numerical values in the trajectory data can be encoded into a high-dimensional trajectory encoding using a Fourier encoding layer. Then, the projection encoding obtained by projecting the high-dimensional trajectory encoding using a learnable preset linear layer is added to the text encoding of the control text to obtain the target text encoding. The target text encoding is input into the model to be trained, and through multiple trainings of the model to be trained, the finally obtained video prediction model can generate prediction video frames with more accurate video angles according to the given initial frame (initial video frame) and the target control text.

[0057] Embodiment 2

[0058] Figure 4 The flowchart of a model training method provided by Embodiment 2 of the present invention. The technical solution of the embodiment of the present invention is further optimized on the basis of the above optional technical solutions, and a specific method for training a video prediction model is given.

[0059] Optionally, updating the intermediate model using the attention mechanism to enable it to process temporal information to obtain a model to be trained includes: inserting a preset temporal inference module before the spatial attention module, cross-attention module, and forward feedback neural network of the Transformer model in the intermediate model to obtain a model to be trained; wherein, the preset temporal inference module includes a causal temporal attention sub-module and a decoupled spatial attention sub-module, the causal temporal attention sub-module is used to perform weighted processing on the input of the Transformer model based on the attention mechanism, and the decoupled spatial attention sub-module is used to decouple the spatial attention of the output of the causal temporal attention sub-module. The advantage of this setting is that the video prediction model has a spaced structure with one layer of spatial interaction and one layer of temporal interaction, enabling the video prediction model to temporally associate the generated prediction frames and the given initial frames, and the decoupled spatial attention sub-module helps the video prediction model capture the associations between different pixels more efficiently, enabling the video prediction model to handle driving scenarios with large dynamic and complex agent movements well.

[0060] Optionally, training the model to be trained using the image encodings of multiple initial video frames and the text encodings of corresponding control texts to obtain a video prediction model includes: inputting the image encodings of multiple initial video frames and the text encodings of corresponding control texts into the model to be trained to obtain a noisy video frame, where the noisy video frame includes initial video frames with different degrees of noise; performing noise reduction processing on the image encoding of the noisy video frame through the model to be trained to obtain the same number of denoised video frames, and determining the video prediction model based on the denoised video frames. The advantage of this setting is that by training the model to be trained, the generated video prediction model can output more accurate prediction video frames according to the control text.

[0061] As Figure 4 shown, a model training method provided in Embodiment 2 of the present invention specifically includes the following steps:

[0062] S201. Determine the scene context of the initial video frame using the initial video frame collected by the target vehicle and / or the scene description metadata of the initial video frame.

[0063] S202. Determine the vehicle movement instruction of the target vehicle using the initial video frame and / or the trajectory data of the target vehicle corresponding to the initial video frame.

[0064] S203. Input the image encodings of at least one initial video frame and the corresponding target text encoding into the StableDiffusion XL model to obtain a denoised image, and determine the loss function value based on the denoised image.

[0065] S204. Determine an intermediate model according to the loss function value, and insert a preset temporal reasoning module before the spatial attention module, the cross-attention module, and the forward feedback neural network in the Transformer model in the intermediate model to obtain a model to be trained.

[0066] Among them, the preset temporal reasoning module includes a causal temporal attention sub-module and a decoupled spatial attention sub-module. The causal temporal attention sub-module is used to perform weighted processing on the input of the Transformer model based on the attention mechanism, and the decoupled spatial attention sub-module is used to decouple the spatial attention of the output of the causal temporal attention sub-module.

[0067] Optionally, a causal mask layer is provided before the Softmax layer of the causal temporal attention sub-module, and the initial parameters of the last linear layers in the causal temporal attention sub-module and the decoupled spatial attention sub-module are zero. The advantage of this setting is that it ensures that in the causal temporal attention sub-module, the features at the current moment cannot view the features at the prediction moment, so as to prevent the leakage of prediction features from affecting the model's learning of predicting the future, and prevent the intermediate model from being affected by the parameter initialization of the preset temporal reasoning module.

[0068] Specifically, the training parameters of the intermediate model can be fixed. On the basis of the intermediate model, a long temporal interaction (Deep Interaction, DI) mechanism is introduced. A preset temporal reasoning module (Temporal Reasoning Block, TRB) is added before the spatial attention module, the cross-attention module, and the forward feedback neural network in the Transformer model included in the intermediate model, so that the video prediction model has an alternating structure of one layer of spatial interaction and one layer of temporal interaction, and can temporally associate the generated prediction frames and the given initial frames.

[0069] Specifically, Figure 5 is a schematic structural diagram of a temporal reasoning module. As Figure 5 shown, the preset temporal reasoning module includes a causal temporal attention sub-module and a decoupled spatial attention sub-module. The function of the causal temporal attention (Causal Temporal Attention, Causal TA) sub-module is realized through the attention mechanism, and it can regard the feature input with the feature dimension of H*W*T as T features (each feature dimension is H*W) for processing. Figure 6 is a schematic structural diagram of a causal temporal attention sub-module. As Figure 6As shown, a causal mask layer is provided before the Softmax layer of the causal temporal attention sub-module. This causal mask layer can, before the Softmax layer, force the values that do not conform to the causal relationship in the result of multiplying the query value by the key value matrix to negative infinity, so that the result output after the Softmax layer is zero, thus realizing the constraint on the causal relationship. The causal mask layer ensures that in the causal temporal attention sub-module, the features at the current moment cannot view the features at the prediction moment, preventing the leakage of prediction features from affecting the model's learning of predicting future images. The final linear layers in both the causal temporal attention sub-module and the decoupled spatial attention sub-module are initialized with all zeros, so that the intermediate model can be directly equivalent to the model before adding the preset temporal inference module at the start of training, ensuring that the intermediate model will not be affected by the parameter initialization of the preset temporal inference module and that the unlearned preset temporal inference module will not interfere with the generation of predicted video frames.

[0070] Specifically, the function of the decoupled spatial attention (Decoupled SA) sub-module is realized through the attention mechanism, which can decouple the output of the causal temporal attention sub-module (i.e., the feature input of the decoupled spatial attention sub-module). For example, two decoupled spatial attention sub-modules can respectively decouple the feature input of the H*W*T feature dimension to two preset vertical directions, and process the input as H features (each feature dimension is W*T) and W features (each feature dimension is H*T).

[0071] S205. Process the trajectory data using the Fourier encoding layer to obtain high-dimensional trajectory encoding, and perform encoding projection on the high-dimensional trajectory encoding using a preset linear layer to obtain projection encoding.

[0072] S206. Input the image encoding of multiple initial video frames and the text encoding of the corresponding target control text into the model to be trained to obtain a noisy video frame.

[0073] Among them, the noisy video frame includes initial video frames with different degrees of noise.

[0074] Specifically, Figure 7 is a schematic diagram of the second-stage model training. As Figure 7 shown, the model to be trained can add noise to the image encoding v = {v m , v n} of T initial video frames to obtain a noisy video frame, where the first m frames (i.e., v m ) are image encodings without noise or with only a small amount of noise (as initial frames), and the last n = T - m frames (i.e., v n ) are obtained after adding noise (As the prediction frame to be predicted), where t is the noise addition step, and the noisy video frame includes v n and

[0075] S207. Denoise the image encoding of the noisy video frame through the model to be trained to obtain a denoised video frame, and determine a video prediction model based on the denoised video frame.

[0076] Specifically, the model to be trained will denoise the image encoding of the noisy video frame to obtain a denoised video frame. The loss function value of the model to be trained can be determined as follows:

[0077]

[0078] Among them, the diffusion model f with a preset temporal inference module (parameters φ to be learned) θ,φ needs to predict what kind of interaction exists between the features of the prediction frame to be predicted during the noise addition process through the initial frame v n , the given control condition c, and the current noise addition step number t. The training objective is to make the noise predicted by the model as close as possible to the noise ∈ actually added in the previous stage.

[0079] Optionally, a preset multi-layer neural network is further included after the Transformer model in the video prediction model. The preset multi-layer neural network is used to generate the predicted trajectory points of the target vehicle corresponding to the predicted video frame based on the feature map output by the Transformer model. The advantage of such a setting is that the prediction of the driving trajectory points of the vehicle can be realized by training a lightweight neural network.

[0080] Specifically, all parameters in the trained video prediction model can be fixed, and the features at the bottom layer of the U-Net structure in the video prediction model, that is, the feature map output by the Transformer model, are input into the preset two-layer neural network. The preset two-layer neural network, as a planner, can predict the future driving trajectory points based on the predicted feature map.

[0081] It should be noted that to verify the prediction effect of the video prediction model, the FID value and FVD value of the video prediction model can be determined using the NuScenes dataset. Among them, the FID value and FVD value are common metrics for evaluation in the field of video generation. FID is used to evaluate the closeness between the generated image and the real images in a certain dataset, that is, to evaluate the authenticity of the generated image. FVD is used to evaluate the closeness between the generated video and the real videos in a certain dataset, that is, to evaluate the authenticity of the generated video. The comparison of the prediction effect of the video prediction model with the effects of other prediction models is shown in the following Table 1 Effect Evaluation Comparison Table:

[0082] Table 1 Comparison Table of Effect Evaluation

[0083]

[0084] Among them, the smaller the values of FID and FVD, the better the effect and the closer to the real video. The DrivingDiffusion model is not a predictive video generation model, but a video generation model that renders according to given conditions and does not have the ability to predict the future. The other three methods are all trained only using the NuScenes dataset and should theoretically have better effects. However, in this case, the video prediction model is still superior to the above existing models in relevant metrics.

[0085] For the temporal inference module and long-term temporal interaction mechanism in the video prediction model, ablation experiments can be conducted on the OpenDV-2K evaluation set to test their effects. Note that the OpenDV-2K evaluation set is completely separated from the training set and there are no duplicate videos. The CLIPSIM metric is the average similarity of the generated predicted frames to the corresponding features of the given initial frame, and the CLIPSIM metric can be used to measure the consistency between the generated predicted video and the given initial frame. The results of the ablation experiments on the temporal inference module and long-term temporal interaction mechanism are shown in Table 2 Information Table of Ablation Experiment Results as follows:

[0086] Table 2 Information Table of Ablation Experiment Results

[0087]

[0088] Among them, the larger the CLIPSIM, the higher the consistency between the predicted video and the given initial frame. From the above experimental results, it can be seen that the temporal inference module and long-term temporal interaction mechanism in the video prediction model can enhance the authenticity of the predicted images and videos, enhance the consistency between the predicted video and the given initial frame, and effectively improve the effect of the generated predicted video.

[0089] To measure the consistency between the predicted video generated by the video prediction model according to the vehicle trajectory and the given vehicle trajectory, the video prediction model can be compared with the IDM (Inverse Dynamic Model). Using the IDM, the generated predicted video can be converted into a trajectory, and the Euclidean distance from the true value of the given vehicle trajectory can be calculated. This distance is denoted as Action Prediction Error (action prediction error) in Table 3 Information Table of Controllability Experiment Results below. The experimental result information on the consistency between the predicted video generated according to the vehicle trajectory and the given vehicle trajectory is shown in Table 3 Information Table of Controllability Experiment Results as follows:

[0090] Table 3 Information Table of Controllability Experiment Results

[0091]

[0092] As can be seen from Table 3 above, the video prediction model that inputs accurate trajectory information (i.e., inputs projection coding) can output a more accurate and controllable video, and the Action Prediction Error is reduced by 20.4%.

[0093] To measure the effectiveness of feature extraction of the video prediction model, the metrics of the end-to-end autonomous driving planning model on the NuScenes dataset can be compared with this video prediction model, and the comparison results are shown in Table 4, the information table of the experimental results of feature extraction effectiveness.

[0094] Table 4 Information Table of Experimental Results of Feature Extraction Effectiveness

[0095]

[0096] Among them, ST-P3* and UniAD* indicate that ST-P3 and UniAD use multiple-angle images of surround view as input, while the video prediction model only uses front view images. ADE (Average Displacement Error) represents the Euclidean distance between the predicted trajectory output by the preset multi-layer neural network in the video prediction model and the trajectory ground truth. FDE (Final Displacement Error) represents the Euclidean distance between the end point of the predicted trajectory point output by the preset multi-layer neural network and the ground truth of the end point of the trajectory point. As can be seen from Table 3, the preset multi-layer neural network in the video prediction model can achieve better results than ST-P3 with one-tenth of the learnable parameters to be learned. Compared with UniAD, the preset multi-layer neural network only uses one-seventieth of the learnable parameters to be learned, can achieve relatively good results, and only needs a shorter training time on NVIDIA Tesla V100, which is nearly 3400 times faster than the training of UniAD.

[0097] The model training method provided by the embodiment of the present invention trains the model to be trained, so that the generated video prediction model can output more accurate predicted video frames according to the control text, and enables the video prediction model to have an interval structure with one layer of spatial interaction and one layer of temporal interaction, so that the video prediction model can temporally associate the generated predicted frames and the given initial frames, and assist the video prediction model to more efficiently capture the associations between different pixels through the decoupled spatial attention sub-module, so that the video prediction model can well handle driving scenarios with large dynamic and complex agent movements.

[0098] Embodiment Three

[0099] Figure 8FIG. 0 is a flowchart of a video prediction method provided in Embodiment 3 of the present invention. This embodiment is applicable to the situation of generating a predicted video. This method can be executed by a video prediction device, which can be implemented in the form of hardware and / or software. The video prediction device can be configured in an electronic device, which can be composed of two or more physical entities or one physical entity.

[0100] As Figure 8 shown, a video prediction method provided in Embodiment 3 of the present invention specifically includes the following steps:

[0101] S301. Obtain at least one conditional video frame, a target scene context, and a target movement instruction of the current vehicle, where the target scene context is the scene context of the current scene corresponding to the conditional video frame.

[0102] In this embodiment, the target movement instruction of the current vehicle can be understood as the driving movement instruction of the current vehicle, and this instruction can be input by the user of the current vehicle.

[0103] S302. Input the encoding of the at least one conditional video frame and the encoding of the target control text into a preset video prediction model to obtain a predicted video, where the preset video prediction model is obtained by using the model training method described in the above embodiment, and the target control text includes the target scene context and / or the target movement instruction.

[0104] In this embodiment, the encoding of at least one frame, such as 1 to 2 frames, of conditional video frames and the encoding of the corresponding target control text can be used as the input of the preset video prediction model, and the preset video prediction model can output a predicted video including multiple video frames. Among them, the predicted video includes reconstructed conditional video frames regenerated by the encoder and multiple predicted video frames, such as 1 to 2 frames of reconstructed conditional video frames and 6 to 7 frames of predicted video frames.

[0105] The video prediction method provided in the embodiment of the present invention obtains at least one conditional video frame, a target scene context, and a target movement instruction of the current vehicle, where the target scene context is the scene context of the current scene corresponding to the conditional video frame, and inputs the encoding of the at least one conditional video frame and the encoding of the target control text into a preset video prediction model to obtain a predicted video, where the preset video prediction model is obtained by using the model training method described in the above embodiment, and the target control text includes the target scene context and / or the target movement instruction. The technical solution of the embodiment of the present invention can quickly and accurately obtain a predicted video by inputting the encoding of the conditional video frame and the encoding of the corresponding target control text including actual requirements into the model.

[0106] Embodiment 4

[0107] Figure 9 This is a schematic structural diagram of a model training device provided in the fourth embodiment of the present invention. As Figure 9 shown, the device includes: a scene context determination module 401, a vehicle movement instruction determination module 402, and a model determination module 403, where:

[0108] The scene context determination module is configured to determine the scene context of the initial video frame by using the initial video frame collected by the target vehicle and / or the scene description metadata of the initial video frame;

[0109] The vehicle movement instruction determination module is configured to determine the vehicle movement instruction of the target vehicle by using the initial video frame and / or the trajectory data of the target vehicle corresponding to the initial video frame;

[0110] The model determination module is configured to train an initial model by using the initial video frame and the corresponding control text to obtain a video prediction model, where the control text includes the scene context and the vehicle movement instruction, and the video prediction model is used to output a predicted video frame.

[0111] The model training device provided in the embodiment of the present invention uses the initial video frame and / or the data of the initial video frame to realize the automatic annotation of natural language instructions and scene context for an unlabeled large-scale data (video frame) set, and trains the model by using the initial video frame and the corresponding control text, so that the model can output a frame image for predicting future images according to the existing picture and the control text based on natural language. The solution of the present invention has higher universality and generalization ability, can generate training samples more freely, is not restricted by the camera settings and shooting scenes of a small-scale public driving data set, and solves the problem that the video prediction model is restricted by the labeled data set and layout information.

[0112] Optionally, the model determination module includes:

[0113] The loss function value determination unit is configured to input the image encoding of at least one of the initial video frames and the text encoding of the corresponding control text into the Stable Diffusion XL model to obtain a denoised image, and determine the loss function value according to the denoised image;

[0114] The model update unit is configured to determine an intermediate model according to the loss function value, and update the intermediate model by using the attention mechanism so that it can process temporal information to obtain a model to be trained;

[0115] The model determination unit is configured to train the model to be trained by using the image encodings of multiple initial video frames and the text encodings of the corresponding control texts to obtain a video prediction model.

[0116] Optionally, updating the intermediate model by using an attention mechanism to enable it to process temporal information to obtain a model to be trained, including: inserting a preset temporal inference module before the spatial attention module, cross-attention module, and forward feedback neural network of the Transformer model in the intermediate model to obtain a model to be trained; wherein, the preset temporal inference module includes a causal temporal attention sub-module and a decoupled spatial attention sub-module, the causal temporal attention sub-module is used to perform weighted processing on the input of the Transformer model based on the attention mechanism, and the decoupled spatial attention sub-module is used to decouple the spatial attention of the output of the causal temporal attention sub-module.

[0117] Further, a causal mask layer is provided before the Softmax layer of the causal temporal attention sub-module, and the initial parameters of the last linear layers in the causal temporal attention sub-module and the decoupled spatial attention sub-module are zero.

[0118] Optionally, training the model to be trained by using the image encodings of a plurality of the initial video frames and the text encodings of the corresponding control texts to obtain a video prediction model, including: inputting the image encodings of a plurality of the initial video frames and the text encodings of the corresponding control texts into the model to be trained to obtain a noisy video frame, wherein the noisy video frame includes initial video frames with different degrees of noise; performing noise reduction processing on the image encoding of the noisy video frame through the model to be trained to obtain a denoised video frame, and determining the video prediction model according to the denoised video frame.

[0119] Optionally, the model determination module further includes:

[0120] An encoding projection unit, configured to, before training the model to be trained by using the image encodings of a plurality of the initial video frames and the text encodings of the corresponding control texts to obtain a video prediction model, process the trajectory data by using a Fourier encoding layer to obtain a high-dimensional trajectory encoding, and perform encoding projection on the high-dimensional trajectory encoding by using a preset linear layer to obtain a projection encoding. The training of this unit can be performed separately after the training of the video prediction model is completed.

[0121] Optionally, training the model to be trained by using the image encodings of a plurality of the initial video frames and the text encodings of the corresponding control texts to obtain a video prediction model, including: training the model to be trained by using the image encodings of a plurality of the initial video frames and the text encodings of the corresponding target control texts to obtain a video prediction model, wherein the text encoding of the target text includes the text encodings of the scene context and the vehicle movement instruction, and the projection encoding.

[0122] Optionally, a preset multi-layer neural network is further included after the Transformer model in the video prediction model. The preset multi-layer neural network is used to generate predicted trajectory points of the target vehicle corresponding to the predicted video frame according to the feature map output by the Transformer model. Among them, the training of this unit can be carried out separately after the video prediction model is trained.

[0123] The model training device provided by the embodiments of the present invention can execute the model training method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0124] Embodiment 5

[0125] Figure 10 It is a schematic structural diagram of a video prediction device provided in Embodiment 5 of the present invention. As Figure 10 shown, the device includes: a condition information acquisition module 501 and a predicted video determination module 502, where:

[0126] The condition information acquisition module is used to acquire at least one conditional video frame, the target scene context, and the target movement instruction of the current vehicle, where the target scene context is the scene context of the current scene corresponding to the conditional video frame;

[0127] The predicted video determination module is used to input the encodings of the at least one conditional video frame and the target control text into a preset video prediction model to obtain a predicted video, where the preset video prediction model is obtained by using the model training method described in the above embodiments, and the target control text includes the target scene context and / or the target movement instruction.

[0128] The video prediction device provided by the embodiments of the present invention can quickly and accurately obtain a predicted video by inputting the encodings of the conditional video frame and the corresponding target control text into the model.

[0129] The video prediction device provided by the embodiments of the present invention can execute the video prediction method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0130] Embodiment 6

[0131] Figure 11The schematic structural diagram of an electronic device 60 that can be used to implement the embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0132] As Figure 11 shown, the electronic device 60 includes at least one processor 61, and a memory communicatively connected to the at least one processor 61, such as a read-only memory (ROM) 62, a random access memory (RAM) 63, etc. The memory stores a computer program executable by the at least one processor. The processor 61 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 62 or the computer program loaded from the storage unit 68 into the random access memory (RAM) 63. In the RAM 63, various programs and data required for the operation of the electronic device 60 can also be stored. The processor 61, the ROM 62, and the RAM 63 are connected to each other through a bus 64. The input / output (I / O) interface 65 is also connected to the bus 64.

[0133] A plurality of components in the electronic device 60 are connected to the I / O interface 65, including: an input unit 66, such as a keyboard, a mouse, etc.; an output unit 67, such as various types of displays, speakers, etc.; a storage unit 68, such as a magnetic disk, an optical disk, etc.; and a communication unit 69, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 69 allows the electronic device 60 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0134] The processor 61 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 61 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 61 executes the various methods and processes described above, such as the model training method, and / or the video prediction method.

[0135] In some embodiments, the model training method and / or the video prediction method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 68. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 60 via the ROM 62 and / or the communication unit 69. When the computer program is loaded into the RAM 63 and executed by the processor 61, one or more steps of the model training method and / or the video prediction method described above may be executed. Alternatively, in other embodiments, the processor 61 may be configured to execute the model training method and / or the video prediction method by any other suitable means (e.g., by means of firmware).

[0136] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0137] The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a dedicated computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0138] The computer device provided above can be used to execute the model training method and / or the video prediction method provided in any of the above embodiments, and has the corresponding functions and beneficial effects.

[0139] Embodiment VII

[0140] In the context of the present invention, a computer-readable storage medium can be a tangible medium, and the computer-executable instructions are used to execute the model training method and / or the video prediction method when executed by a computer processor. The model training method includes:

[0141] Determine the scene context of the initial video frame by using the initial video frame collected by the target vehicle and / or the scene description metadata of the initial video frame;

[0142] Determine the vehicle movement instruction of the target vehicle by using the initial video frame and / or the trajectory data of the target vehicle corresponding to the initial video frame;

[0143] Train an initial model by using the initial video frame and the corresponding control text to obtain a video prediction model, where the control text includes the scene context and the vehicle movement instruction, and the video prediction model is used to output a predicted video frame.

[0144] The video prediction method includes:

[0145] Obtain at least one conditional video frame, and determine the target scene context of the conditional video frame by using the conditional video frame and / or the scene description metadata of the conditional video frame;

[0146] Determine the target movement instruction of the current vehicle by using the conditional video frame and / or the trajectory data of the current vehicle corresponding to the conditional video frame;

[0147] Input the encoding of the at least one conditional video frame and the encoding of the target control text into a preset video prediction model to obtain a predicted video, where the preset video prediction model is obtained by using the model training method described in the above embodiments, and the target control text includes the target scene context and / or the target movement instruction.

[0148] In the context of the present invention, a computer-readable storage medium may be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0149] The computer device provided above can be used to execute the model training method and / or the video prediction method provided in any of the above embodiments, and has the corresponding functions and beneficial effects.

[0150] It should be noted that in the embodiments of the above model training device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.

[0151] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A model training method, characterized in that, Including: Determine the scene context of the initial video frame by using the initial video frame collected by the target vehicle and / or the scene description metadata of the initial video frame; Determine the vehicle movement instruction of the target vehicle by using the initial video frame and / or the trajectory data of the target vehicle corresponding to the initial video frame; Train an initial model by using the initial video frame and the corresponding control text to obtain a video prediction model, where the control text includes the scene context and the vehicle movement instruction, and the video prediction model is used to output a predicted video frame; Among them, the step of training the initial model by using the initial video frame and the corresponding control text to obtain a video prediction model includes: Input the image encoding of at least one of the initial video frames and the text encoding of the corresponding control text into the StableDiffusion XL model to obtain a denoised image, and determine the loss function value according to the denoised image; Determine an intermediate model according to the loss function value, and update the intermediate model by using an attention mechanism to obtain a model to be trained; Train the model to be trained by using the image encodings of multiple initial video frames and the text encodings of the corresponding control texts to obtain a video prediction model; The step of updating the intermediate model by using an attention mechanism to obtain a model to be trained includes: Insert a preset temporal reasoning module before the spatial attention module, cross-attention module, and forward feedback neural network of the Transformer model in the intermediate model to obtain a model to be trained; Among them, the preset temporal reasoning module includes a causal temporal attention sub-module and a decoupled spatial attention sub-module. The causal temporal attention sub-module is used to perform weighted processing on the input quantity of the Transformer model based on the attention mechanism, and the decoupled spatial attention sub-module is used to decouple the spatial attention of the output quantity of the causal temporal attention sub-module.

2. The method according to claim 1, characterized in that, A causal mask layer is arranged before the Softmax layer of the causal temporal attention sub-module, and the initial parameters of the last linear layers in the causal temporal attention sub-module and the decoupled spatial attention sub-module are zero.

3. The method according to claim 1, characterized in that, The step of training the model to be trained by using the image encodings of multiple initial video frames and the text encodings of the corresponding control texts to obtain a video prediction model includes: Input the image encodings of multiple initial video frames and the text encodings of the corresponding control texts into the model to be trained to obtain a noisy video frame, where the noisy video frame includes initial video frames with different degrees of noise; Perform noise reduction processing on the image encoding of the noisy video frame through the model to be trained to obtain a denoised video frame, and determine the video prediction model according to the denoised video frame.

4. The method according to claim 1, characterized in that, Before the step of training the model to be trained by using the image encodings of multiple initial video frames and the text encodings of the corresponding control texts to obtain a video prediction model, it further includes: Process the trajectory data by using a Fourier encoding layer to obtain a high-dimensional trajectory encoding, and perform encoding projection on the high-dimensional trajectory encoding by using a preset linear layer to obtain a projection encoding; Among them, training the to-be-trained model using the image encoding of multiple initial video frames and the text encoding of corresponding control texts to obtain a video prediction model includes: Training the to-be-trained model using the image encoding of multiple initial video frames and the text encoding of corresponding target control texts to obtain a video prediction model, where the text encoding of the target control text includes the text encoding of the scene context and the vehicle movement instruction, as well as the projection encoding.

5. The method according to claim 1, characterized in that, After the Transformer model in the video prediction model, there is also a preset multi-layer neural network, which is used to generate the predicted trajectory points of the target vehicle corresponding to the predicted video frame according to the feature map output by the Transformer model.

6. A video prediction method, characterized in that, Including: Obtaining at least one conditional video frame, the target scene context, and the target movement instruction of the current vehicle, where the target scene context is the scene context of the current scene corresponding to the conditional video frame; Inputting the encoding of the at least one conditional video frame and the encoding of the target control text into a preset video prediction model to obtain a predicted video, where the preset video prediction model is obtained by using the model training method described in any one of claims 1-5, and the target control text includes the target scene context and / or the target movement instruction.

7. A model training device, characterized in that, Including: A scene context determination module, which is used to determine the scene context of the initial video frame by using the initial video frame collected by the target vehicle and / or the scene description metadata of the initial video frame; A vehicle movement instruction determination module, which is used to determine the vehicle movement instruction of the target vehicle by using the initial video frame and / or the trajectory data of the target vehicle corresponding to the initial video frame; A model determination module, which is used to train an initial model by using the initial video frame and the corresponding control text to obtain a video prediction model, where the control text includes the scene context and the vehicle movement instruction, and the video prediction model is used to output predicted video frames; Among them, the model determination module includes: A loss function value determination unit, which is used to input the image encoding of at least one initial video frame and the text encoding of the corresponding control text into the Stable Diffusion XL model to obtain a denoised image, and determine the loss function value according to the denoised image; A model update unit, which is used to determine an intermediate model according to the loss function value, and update the intermediate model by using the attention mechanism so that it can process temporal information to obtain a to-be-trained model; A model determination unit, which is used to train the to-be-trained model by using the image encoding of multiple initial video frames and the text encoding of the corresponding control text to obtain a video prediction model; Updating the intermediate model by using the attention mechanism so that it can process temporal information to obtain a model to be trained, including: inserting a preset temporal inference module before the spatial attention module, cross-attention module, and forward feedback neural network of the Transformer model in the intermediate model to obtain a model to be trained; wherein, the preset temporal inference module includes a causal temporal attention sub-module and a decoupled spatial attention sub-module, the causal temporal attention sub-module is used to perform weighted processing on the input of the Transformer model based on the attention mechanism, and the decoupled spatial attention sub-module is used to decouple the spatial attention of the output of the causal temporal attention sub-module.

8. A video prediction device, characterized in that, Including: A conditional information acquisition module, configured to acquire at least one conditional video frame, a target scene context, and a target movement instruction of the current vehicle, wherein the target scene context is the scene context of the current scene corresponding to the conditional video frame; A predicted video determination module, configured to input the encoding of the at least one conditional video frame and the encoding of the target control text into a preset video prediction model to obtain a predicted video, wherein the preset video prediction model is obtained by using the model training method according to any one of claims 1-5, and the target control text includes the target scene context and / or the target movement instruction.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the model training method according to any one of claims 1-5, and / or execute the video prediction method according to claim 6.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the processor to implement the model training method according to any one of claims 1-5 when executed, and / or implement the video prediction method according to claim 6.

Citation Information

Patent Citations

  • Video prediction method and device, computer equipment and readable storage medium

    CN112465278A

  • System for automatic driving and deployment method thereof

    CN115167446A