Consistency long video generation method based on universal world model

By introducing potential state variables into the video generation model, using the combination of multimodal large model and video diffusion model, the consistency and scalability problems in long video generation are solved, and high-quality and long-term consistent video generation is achieved.

CN120075547APending Publication Date: 2025-05-30TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510090360.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing video generation models are difficult to maintain consistency when generating long videos, and the limitations of computing resources and lack of scalability lead to limited video duration.

Method used

A consistent long video generation method based on a general world model is proposed. Through the combination of multimodal large model and video diffusion model, potential state variables are introduced to achieve wide temporal receptive fields and consistency in video generation.

Benefits of technology

The generated videos perform excellently in logic, dynamics and aesthetic quality, can maintain consistency and coherence over a long period of time, and have the ability to complex scenarios and dynamic changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075547A_ABST
    Figure CN120075547A_ABST
Patent Text Reader

Abstract

The invention provides a universal world model-based consistent long video generation method, which comprises the following steps of: S1, receiving an initially input image and text description, encoding the initially input image and text description into a group of tokens through a word segmentation device network, and inputting the tokens into a multi-mode large model to generate an initial state variable; s2, generating a corresponding video clip by using a video diffusion model under the condition of the current state variable, sampling the video clip, and extracting a key frame to obtain an observation variable; s3, inputting the observation variable into a multi-modal large model, predicting a current dynamic factor by combining a current state variable, and updating the state variable according to the dynamic factor to realize dynamic evolution of the state variable; and S4, repeating the steps S2 and S3, iteratively generating a video clip, and finally generating a long video sequence with time sequence consistency and content richness. By constructing the universal world model, the problems of consistency and content richness in long video generation are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of machine learning, and particularly to a method for generating consistent long videos based on a general world model. Background Art

[0002] With the success of image generation models, video generation has also received increasing attention. Although existing video generation models have achieved commercial-grade performance, the generated video duration is still short. The emergence of long video generation methods has solved this problem, with the focus on increasing the duration and consistency of the generated videos, thus promoting the development of emerging tasks such as video expansion, movie generation, and world simulation.

[0003] Despite the broad application prospects, how to increase the video duration while maintaining consistency remains an urgent problem to be solved. Some work has studied three-dimensional variational autoencoders, which compress videos in both spatial and temporal dimensions to generate long videos during a single denoising process of a latent diffusion model. Although the consistency of the video can be naturally guaranteed during the diffusion process, the generated video duration is still limited by computing resources, and further expanding the video duration requires retraining the diffusion model. Another type of research achieves long video generation through a divide-and-conquer method, first generating key frames of the long video and then interpolating between consecutive key frames. However, these methods rely on the duration of the training video data and thus lack scalability. In addition, iteratively prompting a video diffusion model to generate short segments is also a promising paradigm for generating long videos. To achieve consistency, these methods design prompts based on historical segments and text in each iteration. However, current prompt construction methods usually only adopt the last few frames of directly adjacent segments, which only contain short-term information of the scene, resulting in insufficient consistency over long time spans. Summary of the Invention

[0004] The present application aims to at least solve one of the technical problems in the related art to some extent.

[0005] To this end, the first objective of the present application is to propose a method for generating consistent long videos based on a general world model.

[0006] The second objective of the present application is to propose a device for generating consistent long videos based on a general world model.

[0007] The third objective of the present application is to propose an electronic device.

[0008] The fourth objective of the present application is to propose a computer-readable storage medium.

[0009] The fifth objective of the present application is to propose a computer program product.

[0010] To achieve the above object, an embodiment of the first aspect of the present application proposes a method for generating consistent long videos based on a general world model, including:

[0011] S1. Receive the initially input image and text description, encode them into a set of tokens through a tokenizer network, and input the tokens into a multimodal large model to generate an initial state variable;

[0012] S2. Use a video diffusion model to generate a corresponding video segment conditional on the current state variable, sample the video segment, and extract key frames to obtain an observation variable;

[0013] S3. Input the observation variable into the multimodal large model, combine it with the current state variable, predict the current dynamic factor, and update the state variable according to the dynamic factor to achieve the dynamic evolution of the state variable;

[0014] S4. Repeat the above steps S2 and S3, iteratively generate video segments, and finally generate a long video sequence with temporal consistency and content richness.

[0015] Optionally, receiving the initially input image and text description, and encoding them into a set of tokens through a tokenizer network includes:

[0016] Perform word segmentation on the input text description T to convert it into corresponding text tokens;

[0017] Extract features from the input image I to convert it into corresponding image tokens;

[0018] Fuse the text tokens and image tokens to form a multimodal token sequence [I, T], which is used as input to generate an initial state variable.

[0019] Optionally, receiving the initially input image and text description, and encoding them into a set of tokens through a tokenizer network includes:

[0020] Perform word segmentation on the input text description T to convert it into corresponding text tokens;

[0021] Extract features from the input image I to convert it into corresponding image tokens;

[0022] Fuse the text tokens and image tokens to form a multimodal token sequence [I, T], which is used as input to generate an initial state variable.

[0023] Optionally, inputting the tokens into a multimodal large model to generate an initial state variable includes:

[0024] Input the multimodal token sequence [I, T] into the feature extraction module of the multimodal large model; use the pre-trained vision model to perform embedding processing on the image token I to generate image modality features; use the pre-trained natural language processing model to perform embedding processing on the text token T to generate text modality features.

[0025] In the multimodal large model, fuse the image modality features and the text modality features through the co-attention mechanism, and map the fused multimodal features to the initial state variable s 0 , for use in subsequent video generation processes.

[0026] Optionally, use the video diffusion model to generate a corresponding video segment conditioned on the current state variable, and sample the video segment to extract key frames to obtain the observation variable, including:

[0027] Take the current state variable s t and the observation variable at the previous moment as conditions and input them into the video diffusion model, and gradually denoise to generate the video segment o t , the formula is:

[0028] o t = D(s t , o t-1 )

[0029] where D is the video diffusion model, s t is the state variable, and o t-1 is the observation variable at the previous moment;

[0030] Extract key frames from the generated video segment to obtain key frames;

[0031] Take the key frames as the observation variable o t , for subsequent dynamic factor prediction and state variable update.

[0032] Optionally, input the observation variable into the multimodal large model, combine it with the current state variable, predict the current dynamic factor, and update the state variable according to the dynamic factor, including:

[0033] Take the current state variable s t and the observation variable o t input them into the multimodal large model, and predict the current dynamic factor d t , the formula is:

[0034] d t = f(s t , o t )

[0035] where f is the dynamic prediction function;

[0036] According to the dynamic factor d t , update the state variable s at the next moment through the state update function t+1 , and the formula is:

[0037] s t+1 = g(s t , d t )

[0038] where g is the state update function.

[0039] Optionally, the expression of the finally generated long video sequence is:

[0040] Seq = [I, T, s 0 , o 0 , d 0 , …, s t , o t , d t , …]

[0041] where Seq is the long video sequence.

[0042] Optionally, the video diffusion model and the multimodal large model together constitute a video world model, and the training process of the video world model includes:

[0043] Align the state variables output by the multimodal large model with the text conditional space of the video diffusion model, fine-tune the parameters of the multimodal large model using the LoRA technique, and initialize the multimodal large model and the video diffusion model;

[0044] Unfreeze the parameters of the video diffusion model, and jointly fine-tune the multimodal large model and the video diffusion model on a large-scale short video dataset to train the model's ability to decode state variables into video observations;

[0045] On a small number of high-quality long video datasets, jointly fine-tune the multimodal large model and the video diffusion model to train the model's ability to predict the dynamic factors of future video segments.

[0046] To achieve the above object, the second aspect embodiment of the present application proposes a consistent long video generation device based on a general world model, including:

[0047] The Token encoding module is used to receive the initially input image and text description, encode them into a group of tokens through the tokenizer network, and input the tokens into the multimodal large model to generate initial state variables;

[0048] A video generation module, configured to use a video diffusion model to generate a corresponding video segment based on the current state variable, sample the video segment, and extract key frames to obtain an observation variable;

[0049] A dynamic prediction and state update module, configured to input the observation variable into a multi-modal large model, combine it with the current state variable, predict the current dynamic factors, and update the state variable according to the dynamic factors to achieve the dynamic evolution of the state variable;

[0050] A video generation iteration module, configured to repeat the operations of the video generation module and the dynamic prediction and state update module to iteratively generate video segments, and finally generate a long video sequence with temporal consistency and rich content.

[0051] To achieve the above object, an embodiment of the third aspect of the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0052] The memory stores computer-executable instructions;

[0053] The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of the first aspect.

[0054] To achieve the above object, an embodiment of the fourth aspect of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to any one of the first aspect.

[0055] To achieve the above object, an embodiment of the fifth aspect of the present application provides a computer program product, and when the computer program is executed by a processor, it implements the method according to any one of the first aspect.

[0056] The technical solutions provided by the embodiments of the present application at least bring the following beneficial effects:

[0057] (1) The present application proposes a long video generation method based on the General World Model (Owl). By constructing the General World Model, it effectively solves the problems of consistency and rich content in the long video generation process. The Owl model combines the current information with the historical information by introducing latent state variables, achieving a wide temporal receptive field and video consistency in long video generation.

[0058] (2) This application has demonstrated significant technical advantages in experiments. The videos generated in benchmark tests such as VBench-Long perform excellently in multiple dimensions, including logic, dynamics, and aesthetic quality. Compared with existing video generation models, the videos generated by this application are not only longer but also maintain consistency and coherence over a long time range. In addition, this application can also generate long videos with complex scenes and dynamic changes under the condition of a given text prompt, providing a new solution for the technological progress in the field of video generation.

[0059] (3) This application constructs an autoregressive state-observation-dynamics model to simulate the closed-loop evolution process of the world. This design not only improves the consistency in long video generation (through consistent latent state variables) but also enhances the diversity of content through dynamic prediction. By introducing a pre-trained large-scale multimodal model (LMM) and a video diffusion model, this application effectively decodes the latent state variables into short video segments, achieving the efficiency and quality guarantee of long video generation.

[0060] Additional aspects and advantages of this application will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of this application. Brief Description of the Drawings

[0061] The above-mentioned and / or additional aspects and advantages of this application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0062] Figure 1 is a schematic flowchart of a method for generating consistent long videos based on a general world model provided by an embodiment of this application;

[0063] Figure 2 is a schematic architecture diagram of a method for generating consistent long videos based on a general world model provided by an embodiment of this application;

[0064] Figure 3 is a schematic diagram of the training and inference of the general world model provided by an embodiment of this application;

[0065] Figure 4 is a schematic structural diagram of a method for generating consistent long videos based on a general world model provided by an embodiment of this application. Detailed Description of the Embodiments

[0066] The embodiments of this application will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain this application and should not be construed as limiting this application.

[0067] In view of the problem of generating consistent long videos in the field of content generation existing in the prior art, the embodiments of the present application provide a method for generating consistent long videos based on a general world model. For the general world model, the embodiments of the present application use three sets of variables to model the evolution of the world, including state variables, observation variables, and dynamic factors, and define the evolution relationship among these three in mathematics. For the process of generating consistent long videos, the embodiments of the present application use a multimodal large model to predict the above three sets of variables in an autoregressive manner, and use a video diffusion model to decode the state variables into observation variables.

[0068] The following refers to Figure 1 、 Figure 2 and Figure 3 introduce the definition of the general world model in the embodiments of the present application, as well as the training and inference processes of the general world model. As Figure 1 shown, the method includes the following steps:

[0069] S1. Receive the initially input image and text description, encode them into a set of tokens through a tokenizer network, and input the tokens into the multimodal large model to generate initial state variables.

[0070] In the embodiments of the present application, for a given image and a descriptive text, first input them into a tokenizer network to obtain a set of tokens. Specifically, perform tokenization on the input text description T to convert it into corresponding text tokens; perform feature extraction on the input image I to convert it into corresponding image tokens. Then fuse the text tokens and image tokens to form a multimodal token sequence [I,T] as the input for generating initial state variables.

[0071] The generation of this multimodal token can fully combine image information and text information, providing rich multimodal feature inputs for subsequent state variable modeling.

[0072] Then, the embodiments of the present application use a multimodal large model to model the relationship among state variables, observation variables, and dynamic factors in an autoregressive manner, and this multimodal large model is responsible for the generation of state variables and dynamic factors.

[0073] For the part of generating initial state variables, after obtaining the multimodal token sequence, the embodiments of the present application first input it into the feature extraction module of the multimodal large model and perform the following feature extraction process:

[0074] (1) Embed the image token I using a pre-trained visual model to generate image modality features;

[0075] (2) Use a pre-trained natural language processing model to embed the text token T and generate text modality features.

[0076] Then, in the multimodal large model, fuse the image modality features and text modality features through a co-attention mechanism, and map the fused multimodal features to the initial state variable s 0 , which is used for the subsequent video generation process.

[0077] This initial state variable not only contains the current information of the input image and text, but also incorporates historical information, laying the foundation for the dynamic generation of long videos.

[0078] S2. Use the video diffusion model to generate corresponding video segments conditional on the current state variable, sample the video segments, and extract key frames to obtain the observation variable.

[0079] It should be noted that the state variable is used to represent the state of the world. It not only encodes the world information at the current moment but also incorporates historical information, forming an implicit representation that comprehensively reflects the current and historical conditions of the world. The state variable is decoded into an explicit video observation through the video diffusion model, and the observation variable corresponds to the world state visualized through the video, which is used to capture the specific details and changes of the world.

[0080] In the embodiment of this application, the conversion from the state variable to the observation variable is implemented by a video diffusion model. Specifically, take the current state variable s t and the observation variable at the previous moment as the conditional input to the video diffusion model, and generate the video segment o t through progressive denoising. The formula is:

[0081] o t = D(s t , o t-1 )

[0082] where D is the video diffusion model, s t is the state variable, and o t-1 is the observation variable at the previous moment.

[0083] Then extract the key frames from the generated video segments to obtain the key frames, and use this key frame as the observation variable o t , which is used for subsequent dynamic factor prediction and state variable update. The observation variable represents the world state at the current moment in a visual form, which can provide a more intuitive reference condition for the model and support the subsequent iterative generation process.

[0084] S3. Input the observation variable into the multimodal large model, combine it with the current state variable, predict the current dynamic factor, and update the state variable according to the dynamic factor to achieve the dynamic evolution of the state variable.

[0085] It should be noted that the driving factor represents the force or event that drives the evolution of the world, which is usually described in text form and is an important variable reflecting the internal change law of the world. In this application, the driving factor is predicted by a multimodal large model in combination with state variables and observation variables.

[0086] Specifically, when implementing, the current state variable s t and the observation variable o t are input into the multimodal large model, and the current driving factor d t is predicted through the driving force prediction function. The formula is:

[0087] d t = d(s t , o t )

[0088] where d is the driving force prediction function. The prediction of the driving factor is based on the state variable and observation variable at the current moment and can reflect the internal dynamic changes of the world at the current moment.

[0089] It can be understood that the predicted driving factor d t is the key to promoting the update of the state variable.

[0090] In the embodiment of this application, further through the state update function, according to the current state variable s t and the driving factor d t , the state variable s t+1 at the next moment is updated. The formula is:

[0091] s t+1 = g(s t , d t )

[0092] where g is the state update function. The state update process realizes the dynamic evolution of the state variable, enabling it to gradually adapt to the changes in the world.

[0093] S4. Repeat the above steps S2 and S3 to iteratively generate video segments, and finally generate a long video sequence with temporal consistency and content richness.

[0094] After obtaining the state variable s t+1 through step S3, in the embodiment of this application, according to step S2, it is input into the video diffusion model to realize the conversion of the state variable - observation variable, and then through step S3, the observation variable is input into the multimodal large model to realize the conversion of the observation variable - driving factor, and the state variable is updated again according to the driving factor. Repeat this iteration process, thus forming a closed-loop world evolution process. Through this iteration process, the model generates a series of video segments with good consistency, and finally constructs a long video sequence. The expression of the finally generated long video sequence is:

[0095] Seq = [I, T, s 0 , o 0 , d 0 , …, s t , o t , d t , …]

[0096] Among them, Seq is a long video sequence.

[0097] Based on the above process, it can be seen that the embodiment of the present application constructs a complete dynamic evolution closed loop. This method not only improves the consistency of long video generation, but also enhances the richness of video content through dynamic factors, providing important support for the modeling of complex dynamic scenes.

[0098] In addition, to solve the problem of lack of high-quality long video data in reality, the embodiment of the present application designs a multi-stage training process for the video world model. The specific process includes the following three stages:

[0099] In the first stage, align the state variables output by the multi-modal large model with the text conditional space of the video diffusion model. Through this alignment method, good initialization is provided for the joint fine-tuning of the multi-modal large model and the video diffusion model. In this stage, the LoRA technology is used to fine-tune the multi-modal large model so that it can more efficiently adapt to subsequent tasks.

[0100] In the second stage, unfreeze the parameters of the video diffusion model, and jointly fine-tune the multi-modal large model and the video diffusion model on a large-scale short video dataset. Through training on the short video dataset, the ability of the model to decode state variables into video observations is fully improved, laying a foundation for subsequent long video generation.

[0101] In the third stage, further jointly fine-tune the multi-modal large model and the video diffusion model using a small amount of high-quality long video data. The focus of this stage is to train the model's ability to predict the dynamic factors of future video segments, so that the model can maintain the temporal consistency and content richness of the generated video in a long time range.

[0102] Through the above multi-stage training process, the video world model of the present application realizes the decoding ability from state variables to video observations and the accurate prediction ability of dynamic factors. This process makes full use of the wide availability of short video data, and at the same time effectively overcomes the problem of lack of high-quality long video data, providing an efficient training strategy for constructing a consistent long video generation system.

[0103] To implement the above embodiment, the present application also proposes a consistent long video generation device based on a general world model. Figure 4The figure is a schematic structural diagram of a consistency long video generation device 10 provided by an embodiment of the present application. As Figure 4 shown, the device includes:

[0104] A Token encoding module 100, configured to receive an initially input image and text description, encode them into a set of tokens through a tokenizer network, and input the tokens into a multi-modal large model to generate an initial state variable;

[0105] A video generation module 200, configured to use a video diffusion model to generate a corresponding video segment based on the current state variable, sample the video segment, and extract key frames to obtain an observation variable;

[0106] A dynamic prediction and state update module 300, configured to input the observation variable into the multi-modal large model, combine the current state variable, predict the current dynamic factor, and update the state variable according to the dynamic factor to realize the dynamic evolution of the state variable;

[0107] A video generation iteration module 400, configured to repeat the operations of the video generation module and the dynamic prediction and state update module, iteratively generate video segments, and finally generate a long video sequence with temporal consistency and rich content.

[0108] To implement the above embodiments, the present application also proposes an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiments.

[0109] To implement the above embodiments, the present application also proposes a computer-readable storage medium storing computer execution instructions, and the computer execution instructions are used to implement the method provided in the foregoing embodiments when executed by a processor.

[0110] To implement the above embodiments, the present application also proposes a computer program product including a computer program, and the computer program implements the method provided in the foregoing embodiments when executed by a processor.

[0111] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the present application all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0112] It should be noted that personal information from users should be collected for legal and reasonable purposes and should not be shared or sold outside of such legal uses. In addition, such collection / sharing should be carried out after obtaining the informed consent of the users, including but not limited to notifying the users to read the user agreement / user notice and sign an agreement / authorization including authorizing relevant user information before the users use the function. In addition, any necessary steps should be taken to safeguard and protect access to such personal information data and ensure that others with access to the personal information data comply with their privacy policies and procedures.

[0113] This application is expected to provide an implementation for users to selectively block the use or access of personal information data. That is, this disclosure is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of the users.

[0114] In the description of the foregoing embodiments, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0115] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of such features. In the description of this application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0116] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred implementation of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of this application belong.

[0117] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered a definitional sequence list of executable instructions for implementing logical functions, and can be embodied specifically in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0118] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0119] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods of the above-described embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0120] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, may exist physically alone for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0121] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

[0122] It should be understood that various forms of the processes shown above may be used, steps may be reordered, added, or deleted. For example, the steps described in the present application may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present application can be achieved. No limitation is imposed herein.

[0123] The above specific embodiments do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present application shall be included within the protection scope of the present application.

Claims

1. A consistent long video generation method based on a universal world model, characterized in that: The following steps are involved: S1, receiving the initial input image and text description, encoding them into a set of tokens through a word segmentation network, and inputting the tokens into a multimodal large model to generate initial state variables; S2, using the video diffusion model to generate corresponding video clips with the current state variable as a condition, and sampling the video clips to extract key frames to obtain observed variables; S3, inputting the observed variables into the multimodal large model, combining the current state variables, predicting the current dynamic factors, and updating the state variables according to the dynamic factors to achieve dynamic evolution of the state variables; S4. Repeat the above steps S2 and S3 to iteratively generate video clips, and finally generate a long video sequence with temporal consistency and rich content.

2. The method according to claim 1, characterized in that Receive the initial input image and text description, and encode it into a set of tokens through the tokenizer network, including: Perform word segmentation on the input text description T and convert it into corresponding text tokens; Extract features from the input image I and convert it into corresponding image tokens; The text token and the image token are fused to form a multimodal token sequence [I, T], which is used as input to generate initial state variables.

3. The method according to claim 2, characterized in that The token is input into the multimodal large model to generate initial state variables, including: Input the multimodal token sequence [I, T] into the feature extraction module of the multimodal large model; use the pre-trained visual model to embed the image token I to generate image modal features; use the pre-trained natural language processing model to embed the text token T to generate text modal features; In the multimodal large model, the image modal features and the text modal features are fused through a collaborative attention mechanism, and the fused multimodal features are mapped to the initial state variable s0 for the subsequent video generation process.

4. The method according to claim 3, characterized in that The video diffusion model is used to generate corresponding video clips with the current state variables as conditions, and the video clips are sampled to extract key frames to obtain observed variables, including: Set the current state variable s t The observed variables at the previous moment are used as conditions to input the video diffusion model, and the video segment o is generated by step-by-step denoising. t , the formula is: o t =D(s t ,o t-1 ) Where D is the video diffusion model, s t is the state variable, o t-1 is the observed variable at the previous moment; Extracting key frames from the generated video clip to obtain key frames; The key frame is used as the observation variable o t , used for subsequent dynamic factor prediction and state variable update.

5. The method according to claim 4, characterized in that Inputting the observed variables into the multimodal large model, combining with the current state variables, predicting the current dynamic factors, and updating the state variables according to the dynamic factors, including: Set the current state variable s t and the observed variable o t Input the multi-modal large model and predict the current dynamic factor d through the dynamic prediction function t , the formula is: d t =f(s t ,o t ) Where, f is the power prediction function; According to the dynamic factor d t , update the state variable s at the next moment through the state update function t+1 , the formula is: s t+1 =g(s t ,d t ) Among them, g is the state update function.

6. The method according to claim 5, characterized in that The expression of the final generated long video sequence is: Seq=[I,T,s0,o0,d0,…,s t ,o t ,d t ,…] Wherein, Seq is the long video sequence.

7. The method according to claim 6, characterized in that The video diffusion model and the multimodal large model together constitute a video world model. The training process of the video world model includes: Align the state variables of the multimodal large model output with the text conditional space of the video diffusion model, fine-tune the parameters of the multimodal large model using LoRA technology, and initialize the multimodal large model and the video diffusion model; Unfreeze the parameters of the video diffusion model, and jointly fine-tune the multimodal large model and the video diffusion model on a large-scale short video dataset to train the model's ability to decode state variables into video observations; On a small amount of high-quality long video datasets, the fine-tuned multimodal large model is combined with the video diffusion model to train the model's ability to predict the dynamic factors of future video clips.

8. A consistent long video generation device based on a universal world model, characterized in that: include: The Token encoding module is used to receive the initial input image and text description, encode them into a set of tokens through the word segmentation network, and input the tokens into the multimodal large model to generate the initial state variables; A video generation module is used to generate corresponding video clips using a video diffusion model with current state variables as conditions, and to sample the video clips and extract key frames to obtain observed variables; The power prediction and state update module is used to input the observed variables into the multimodal large model, combine the current state variables, predict the current power factors, and update the state variables according to the power factors to achieve dynamic evolution of the state variables; The video generation iteration module is used to repeat the operations of the video generation module and the power prediction and state update module, iteratively generate video clips, and finally generate a long video sequence with temporal consistency and rich content.

9. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.