Video generation method and related apparatus

By embedding latent vectors of control conditions into the video generation model, the problem of the inability to precisely control video generation models in existing technologies is solved, enabling precise control over video structure, motion, and physical properties, and improving the controllability and adaptability of video generation.

WO2025260799A1PCT designated stage Publication Date: 2025-12-26HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/078222
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-29
Filing Date
2025-02-20
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing video generation models cannot meet users' needs for precise control over videos, especially in autonomous driving scenarios where multiple parameters of video attributes need to be precisely adjusted, resulting in poor controllability of the generated videos.

Method used

By embedding latent vectors corresponding to control conditions into the video generation model, the structural, motion, and physical properties of the video are controlled respectively. Video generation is performed using a diffusion model and a cross-attention layer, supporting precise user control over the video.

Benefits of technology

It enables more precise control over videos, meets users' fine-grained needs, improves the controllability and adaptability of video generation, and can generate videos that meet specified conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025078222_26122025_PF_FP_ABST
    Figure CN2025078222_26122025_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present application relate to the technical field of artificial intelligence, and particularly relate to a video generation method and an electronic device. By means of the method, an accurately controllable video by means of a control condition, thereby improving the controllability of video generation. Specifically, the method may comprise: receiving at least one control condition, wherein the at least one control condition is used for controlling at least one attribute of a target video to be generated, and the at least one attribute comprises one or more of the following: a structural attribute, a motion attribute and a physical attribute; coding the at least one control condition into at least one latent vector; and inputting the at least one latent vector into a pre-trained video generation model, and generating, by means of the video generation model, a target video meeting the at least one control condition.
Need to check novelty before this filing date? Find Prior Art

Description

Video generation method and related apparatus

[0001] This application claims priority to the Chinese Patent Application No. 2024107980555, filed on June 19, 2024, entitled "Video generation method and electronic device", the content of which is incorporated herein by reference in its entirety.

[0002] Also, this application claims priority to the Chinese Patent Application No. 2024113753012, filed on September 29, 2024, entitled "Video generation method and related apparatus", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0003] Embodiments of the present application relate to the field of AI technology, in particular to a video generation method and related apparatus. BACKGROUND

[0004] With the development of Artificial Intelligence Generated Content (AIGC) and other technologies, some video generation models can support text-to-video, image-to-video and video continuation functions, and can generate high realistic videos according to the text prompts input by users.

[0005] Although the video generation model in the related art can generate high realistic videos, it only supports input of text and / or image prompt information, and cannot meet the needs of controllable video generation under more control conditions in actual scenarios. For example, a user wants to generate a large amount of video data as training samples in an automatic driving scenario, which may require accurate regulation and control of various parameters in video attributes. However, the video generation model in the related art only supports generation of videos according to simple prompt information, and does not support more accurate control of videos. SUMMARY

[0006] Embodiments of the present application provide a video generation method and electronic device, which set corresponding control conditions for various attributes of videos for accurate control respectively, and can generate videos meeting the control conditions by embedding the hidden vectors corresponding to the control conditions in the video generation model, so as to meet the actual needs of users for accurate control of videos.

[0007] In a first aspect, an embodiment of the present application provides a video generation method, which is applied to an electronic device. The method receives at least one control condition, the at least one control condition being used to control at least one attribute of a target video to be generated, the at least one attribute including one or more of the following: a structural attribute, a motion attribute, and a physical attribute. The method encodes the at least one control condition into at least one latent vector. The method inputs the at least one latent vector into a pre-trained video generation model, and generates the target video conforming to the at least one control condition through the video generation model.

[0008] The at least one attribute of the target video to be generated can be controlled as the at least one attribute of an image sequence included in the target video to be generated. The target video is a video format data output by the video generation model. The target video to be generated can be understood as a control object of the control condition. The attribute of the video (which can also be referred to as a controllable attribute) can be understood as various controllable objects in the video. All controllable objects are classified, and at least one attribute can be obtained. The control condition is used to control the corresponding attribute, and the control condition and the attribute can be in a one-to-one correspondence. Different control conditions can control different attributes in the video.

[0009] In this way, by using the above method, a user can control at least one of the structural attribute, the motion attribute, and the physical attribute of the video by inputting the control condition. For example, for the structural attribute, the user can input a corresponding control condition to specify which materials are included in the video and the distribution positions of the materials. For the motion attribute, the user can input a corresponding control condition to specify a motion trajectory of a material in the foreground (or a material in the background). For the physical attribute, the user can input a corresponding control condition to control a physical phenomenon expected by the user to appear in the video, such as a non-rigid body motion phenomenon. In this way, the user can control the generated video more accurately by inputting the control condition, thereby meeting the more granular needs of the user and improving the controllability of the generated target video. In addition, the user can also control different attributes jointly to achieve a more complex control target. For example, a material can be controlled to move along a set trajectory and then produce a specified physical phenomenon with one of a plurality of other materials.

[0010] The at least one control condition can be received by providing a visual interface and receiving the control condition set by the user through the visual interface, such as receiving text input by the user or a file uploaded by the user. Alternatively, in other embodiments, a visual interface can not be provided, and an API is provided for the user to call. The at least one control condition set by the user is received in a calling request.

[0011] In a possible implementation, the structural attribute includes material structure and / or camera track of the target video, the material structure includes position or category of at least one material. The motion attribute includes motion track of at least one material in the target video. The physical attribute includes physical phenomenon when the at least one material moves in the target video.

[0012] The structural attribute, the motion attribute, and the physical attribute can be independent of each other, and the target video can be controlled independently.

[0013] In a possible implementation, the target video includes a first material and a second material, the at least one attribute includes a structural attribute and a motion attribute, the structural attribute includes material structure of the first material and the second material, the motion attribute includes motion track of the first material, and the target video includes a video in which the first material interacts with the second material along the motion track.

[0014] In this implementation, the motion attribute and the structural attribute can be controlled in combination, for example, the material structure of the first material and the second material can be controlled, and the motion track of the first material can be controlled, so that the first material interacts with the second material in the video. The interaction can be collision, or friction, rubbing, or other forms of interaction, where collision can cause a change in motion track and / or deformation.

[0015] In a possible implementation, the at least one attribute further includes a physical attribute, and the target video includes a video in which the first material interacts with the second material in a physical phenomenon along the motion track.

[0016] In this implementation, the motion attribute, the structural attribute, and the physical attribute can be controlled in combination, so that the first material and the second material in the video interact in a manner consistent with the specified physical phenomenon. The interaction in a physical phenomenon can be an interaction that is consistent with the characteristics or rules of a specified physical phenomenon, for example, the physical phenomenon can be non-rigid motion, and the interaction can be collision. Collision that is consistent with the characteristics of non-rigid motion can be understood as collision that causes deformation during the collision process. Collision that does not cause deformation is not collision that is consistent with the characteristics of non-rigid motion.

[0017] In a possible implementation, the at least one attribute further includes an environmental attribute, and the environmental attribute includes environmental parameter information of the target video, the environmental parameter including at least one of weather and illumination.

[0018] In the present implementation, the attributes in the video content can be divided into structure, motion, environment, and physics, and each of them can be controlled. In other embodiments, the video attributes can be divided in other ways, for example, they can be divided into two controllable attributes, intra-frame and inter-frame. Intra-frame attributes include various controllable objects within a frame of image, such as material structure, etc. Inter-frame attributes include the relative position relationship between multiple frames of images, such as the position change trajectory of the same material in multiple frames of images, etc. Alternatively, other classification methods can be used to obtain at least one controllable attribute, for example, static attributes or dynamic attributes, etc. The attribute classification method provided by the present implementation can achieve accurate control of the video, and is compatible with 3D simulation tools in related technologies, which is beneficial to the generation of training data.

[0019] In a possible implementation, determining the at least one control condition includes: displaying a first control, and receiving a control condition sent by a user through the first control; wherein the control condition indicates control of the structure attribute, and the control condition includes one or more of a text and a file in a predetermined format; or the control condition indicates control of the motion attribute, and the control condition includes one or more of a text and a file in a predetermined format; or the control condition indicates control of the physics attribute, and the control condition includes one or more of a text and a file in a predetermined format.

[0020] In the present implementation, the control condition can be a text instruction input by the user and / or a file in a predetermined format uploaded by the user. The file in a predetermined format can be an image, a table, a document, or other file formats that support uploading.

[0021] In a possible implementation, a second control can also be displayed. The second control is used to receive prompt information input by the user, and the prompt information is used to guide the generation of the target video. The prompt information includes a text and / or an image.

[0022] The present implementation can be compatible with existing text and / or image-based control methods, which means that the attribute-based control method is added on the basis of the existing text and / or image prompt information for controlling the generation of the video.

[0023] In a possible implementation, the video generation model is a diffusion model Diffusion; the diffusion model includes a cross-attention layer; the cross-attention layer includes at least one first variable; and inputting the at least one latent vector into the pre-trained video generation model includes: inputting the at least one latent vector into the cross-attention layer in the pre-trained video generation model, and assigning a value of the at least one latent vector to the at least one first variable; and / or, an input layer of the diffusion model includes a noise vector corresponding to random noise; the noise vector is embedded with at least one second variable; and inputting the at least one latent vector into the pre-trained video generation model includes: assigning a value of the at least one latent vector to the second variable, and inputting the noise vector embedded with the second variable into the pre-trained video generation model.

[0024] In the present implementation, the video generation model can be Diffusion. In other embodiments, the video generation model can be other self-regressive video generation models containing a cross-attention layer other than Diffusion. The cross-attention layer includes a first variable, and a value of the first variable is equal to a value of a latent vector corresponding to a control condition, which can be understood as embedding the latent vector corresponding to the control condition in the cross-attention layer. In some embodiments, at least one latent vector can be embedded only in the cross-attention layer. In other embodiments, at least one latent vector can be embedded only through a noise vector corresponding to random noise. In yet other embodiments, at least one latent vector can be embedded through the cross-attention layer and at least one latent vector can be embedded through a noise vector corresponding to random noise.

[0025] Embedding the latent vector in the video generation model can fully learn the influence of the control condition on the video by using the learning mechanism of the diffusion model itself, so that the trained video generation model can generate a video consistent with the control condition when running.

[0026] In a possible implementation, embedding at least one second variable in the noise vector can be connecting the noise vector with the at least one second variable; or the at least one second variable is the noise vector.

[0027] In the present implementation, the second variable is embedded in the noise vector. It can be understood that the at least one hidden vector is embedded in the noise vector. The specific embedding can be to use the concat function to concatenate the noise vector and the at least one hidden vector. Alternatively, in other embodiments, embedding the at least one hidden vector in the noise vector can be to perform feature fusion on the at least one hidden vector and the noise vector, for example, to perform weighted summation or dot multiplication operation on the at least one hidden vector and the noise vector, or to perform other operations that can fuse the features of the at least one hidden vector and the noise vector. Alternatively, the embedding manner can also be to use the at least one hidden vector to replace the original random noise vector, that is, to use the at least one hidden vector as the noise vector to participate in subsequent calculation of the Diffusion model.

[0028] The noise vector of the random noise is part of the features in the input data of the Diffusion model. By fusing the control condition corresponding hidden vector and the noise vector through embedding, the influence of the control condition on the video can be learned by the diffusion mechanism of the diffusion model, and a video that meets the control condition can be generated.

[0029] In a possible implementation, an encoder can be called to encode the at least one control condition to obtain the at least one hidden vector corresponding to the at least one control condition.

[0030] Encoding the control condition can obtain a hidden vector that can be input to the video generation model and is compatible with the video generation model.

[0031] In a possible implementation, before displaying the first control, the method can further obtain training data with labels; the labels are at least one first hidden vector corresponding to the at least one control condition; in one iteration of training the video generation model, the method further includes: extracting the at least one control condition from a target video output by the video generation model; obtaining at least one second hidden vector corresponding to the at least one control condition extracted from the target video; calculating a value of a loss function according to the at least one second hidden vector and the at least one first hidden vector, and optimizing trainable parameters in the video generation model according to the value of the loss function.

[0032] In the training phase, through effective training of the video generation model, a video generation model with the ability to generate a video meeting the control condition can be obtained.

[0033] In a possible implementation, at least one control condition is extracted from the target video output by the video generation model, including performing one or more of the following steps: calling a semantic segmentation model and / or a 3D detection model, and extracting a structure corresponding control condition from the target video output by the video generation model through the semantic segmentation model and / or the 3D detection model; calling a 3D tracking model, and extracting a motion corresponding control condition from the target video output by the video generation model through the 3D tracking model; calling a multi-modal understanding model, and extracting an environment corresponding control condition from the target video output by the video generation model through the multi-modal understanding model; calling the multi-modal understanding model, and extracting a physics corresponding control condition from the target video output by the video generation model through the multi-modal understanding model.

[0034] In the training phase, it is necessary to evaluate whether the video output by the video generation model in training meets the preset control condition, and the premise of evaluation is to extract the control condition from the video output by the video generation model. In the extraction, the extraction of the control condition can be achieved by means of the above-mentioned model.

[0035] In a possible implementation, the labeled training data is obtained, including: determining a material and a background used for generating a sample video; running a simulation model, and importing the material and the background into the simulation model; detecting at least one control condition set by a user, and generating a sample video meeting the at least one control condition through the simulation model; calling an encoder, and encoding the at least one control condition through the encoder to obtain at least one first hidden vector; and taking the at least one first hidden vector as a label of the sample video to obtain the labeled training data.

[0036] The simulation model can be a 3D simulation engine, so that the training data can be obtained through 3D simulation, solving the problem of the source of the training data.

[0037] In a possible implementation, after the target video meeting the at least one control condition is generated, a quality feedback module can be further called to obtain a quality score of the target video through the quality feedback module; the quality score is obtained based on evaluation results of subjective indexes and / or objective indexes; the evaluation results of the subjective indexes are obtained according to user operations; the objective indexes include a video quality evaluation index FVD and / or a multi-modal model similarity index CLIP Similarity; in a case where the quality score is lower than a threshold value, incremental training of the video generation model is triggered.

[0038] The incremental training can continue to improve the accuracy of the video generation model in actual applications, ensure that the prediction performance of the video generation model does not decrease, or in other words, according to the actual data in the actual application as a sample for incremental training, the video generation model can better fit the specific characteristics of the actual application scene, output a video that better meets the current application environment, and improve the ability to adapt to different environments.

[0039] It should be noted that in the incremental training process, the video generation model needs to be iteratively trained until the quality score meets the threshold requirement, for example, equal to or greater than the threshold.

[0040] In a second aspect, the embodiments of the present application also provide a video generation device, which can include: a first receiving module configured to receive at least one control condition, the at least one control condition being used to control at least one attribute of a target video to be generated, the at least one attribute including one or more of the following: a structure attribute, a motion attribute, and a physical attribute; an encoding module configured to encode the at least one control condition into at least one latent vector; and a generation module configured to input the at least one latent vector into a pre-trained video generation model, and generate a target video meeting the at least one control condition through the video generation model.

[0041] In a possible implementation, the target video includes at least one material, the structure attribute includes a material structure and / or a camera track in the target video, the material structure includes a position or a category of the at least one material, the motion attribute includes a motion track of the at least one material in the target video, and the physical attribute includes a physical phenomenon when the at least one material moves in the target video.

[0042] In a possible implementation, the target video includes a first material and a second material, the at least one attribute includes a structure attribute and a motion attribute, the structure attribute includes a material structure of the first material and the second material, the motion attribute includes a motion track of the first material, and the target video includes a video in which the first material interacts with the second material along the motion track.

[0043] In a possible implementation, the at least one attribute further includes a physical attribute, and the target video includes a video in which the first material interacts with the second material in a physical phenomenon along the motion track.

[0044] In a possible implementation, the at least one attribute further includes an environment attribute, and the environment attribute includes environment parameter information of the target video, the environment parameter including at least one of weather and illumination.

[0045] In a possible implementation, when the at least one control condition is received, the first receiving module is specifically configured to: display a first control, and receive the control condition sent by the user through the first control; the control condition indicates control over the structural attribute, and the control condition includes one or more of a text and a file in a predetermined format; or the control condition indicates control over the motion attribute, and the control condition includes one or more of a text and a file in a predetermined format; or the control condition indicates control over the physical attribute, and the control condition includes one or more of a text and a file in a predetermined format.

[0046] In a possible implementation, the apparatus further includes a second receiving module configured to: display a second control; and receive prompt information input by the user, the prompt information being used to guide generation of the target video, and the prompt information including a text and / or an image.

[0047] In a possible implementation, the video generation model is a diffusion model Diffusion; the diffusion model includes a cross-attention layer; the cross-attention layer includes at least one first variable; when the at least one latent vector is input into the pre-trained video generation model, the generation module is specifically configured to: input the at least one latent vector into the cross-attention layer in the pre-trained video generation model, and assign a value of the at least one latent vector to the at least one first variable; and / or, an input layer of the diffusion model includes a noise vector corresponding to random noise; the noise vector is embedded with at least one second variable; when the at least one latent vector is input into the pre-trained video generation model, the generation module is specifically configured to: assign a value of the at least one latent vector to the second variable, and input the noise vector embedded with the second variable into the pre-trained video generation model.

[0048] In a possible implementation, when the at least one second variable is embedded in the noise vector, the generation module is specifically configured to: connect the noise vector with the at least one second variable; or, the at least one second variable is the noise vector.

[0049] In a possible implementation, the apparatus further includes a training module configured to, before the at least one control condition is received, perform the following steps: obtain training data with labels; the labels are at least one first latent vector coded from the at least one control condition; in one iteration of training the video generation model, the training module is further configured to: extract the at least one control condition from a target video output by the video generation model; obtain at least one second latent vector corresponding to the at least one control condition extracted from the target video; calculate a value of a loss function according to the at least one second latent vector and the at least one first latent vector; and optimize trainable parameters in the video generation model according to the value of the loss function.

[0050] In a possible implementation, when the at least one control condition is extracted from the target video output by the video generation model, the training module is specifically configured to perform one or more of the following steps: calling a semantic segmentation model and / or a 3D detection model, and extracting the structure corresponding control condition from the target video output by the video generation model by using the semantic segmentation model and / or the 3D detection model; calling a 3D tracking model, and extracting the motion corresponding control condition from the target video output by the video generation model by using the 3D tracking model; calling a multi-modal understanding model, and extracting the environment corresponding control condition from the target video output by the video generation model by using the multi-modal understanding model; calling the multi-modal understanding model, and extracting the physics corresponding control condition from the target video output by the video generation model by using the multi-modal understanding model.

[0051] In a possible implementation, when the training data with labels is obtained, the training module is specifically configured to: determine the material and the background used to generate the sample video; run a simulation model, and import the material and the background into the simulation model; detect the at least one control condition sent by the user, and generate the sample video that meets the at least one control condition by using the simulation model; call an encoder, and encode the at least one control condition by using the encoder to obtain at least one first hidden vector; and obtain the at least one first hidden vector as the label of the sample video to obtain the training data with labels.

[0052] In a possible implementation, the apparatus further includes a triggering module configured to perform the following steps after the target video that meets the at least one control condition is generated: calling a quality feedback module, and obtaining a quality score of the target video by using the quality feedback module; the quality score is obtained based on evaluation results of subjective indexes and / or objective indexes; the evaluation results of the subjective indexes are obtained according to user operations; the objective indexes include a video quality evaluation index FVD and / or a multi-modal model similarity index CLIP Similarity; and in a case where the quality score is lower than a threshold value, triggering the incremental training of the video generation model.

[0053] In a third aspect, an embodiment of the present application further provides a computing device cluster, including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the method in any one of the above first aspects.

[0054] In a fourth aspect, an embodiment of the present application further provides a computer program product including instructions, which, when executed by a computing device cluster, cause the computing device cluster to perform the method in any one of the above first aspects.

[0055] In a fifth aspect, the embodiments of the present application further provide a computer readable storage medium, including computer program instructions, when the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of the first aspect.

[0056] In other aspects, the embodiments of the present application further provide an electronic device, including: a processor configured to execute computer programs or instructions in a memory to implement the method according to any one of the above.

[0057] In other aspects, the embodiments of the present application further provide a chip system, including: a communication interface configured to input and / or output data; and a processor configured to execute computer executable programs, so that the device installed with the chip system executes the method according to any one of the above. BRIEF DESCRIPTION OF DRAWINGS

[0058] FIG. 1 is an example of a flow of generating a video based on a traditional UE simulation engine in the related art;

[0059] FIG. 2 is an example of generating a video by a general video generation model in the related art;

[0060] FIG. 3 is a schematic diagram of a structure of an electronic device (terminal device, such as a smart phone);

[0061] FIG. 4 is a schematic diagram of a structure of an electronic device (server);

[0062] FIG. 5 is a schematic diagram of an application scenario of the video generation method provided by the embodiments of the present application;

[0063] FIG. 6 is an example of a system architecture of the video generation method provided by the embodiments of the present application;

[0064] FIG. 7 is a schematic diagram of a foreground structure and a background structure in some embodiments of the video generation method provided by the embodiments of the present application;

[0065] FIG. 8 is a schematic diagram of controlling a motion trajectory in some embodiments of the video generation method provided by the embodiments of the present application;

[0066] FIG. 9 is a schematic diagram of controlling an environment in some embodiments of the video generation method provided by the embodiments of the present application;

[0067] FIG. 10 is an example of a software module architecture in some embodiments of the video generation method provided by the embodiments of the present application;

[0068] FIG. 11 is an example of a flow processing in which the video generation method provided by the embodiments of the present application is divided into a training phase, a use phase and an incremental training phase;

[0069] FIG. 12 is an example diagram of embedding hidden vectors in a cross attention layer in some embodiments of the video generation method provided by the embodiments of the present application;

[0070] FIG. 13 is another example diagram of embedding hidden vectors in a cross attention layer in some embodiments of the video generation method provided by the embodiments of the present application;

[0071] FIG. 14 is an example diagram of embedding hidden vectors in a noise vector in some embodiments of the video generation method provided by the embodiments of the present application;

[0072] FIG. 15 is another example diagram of embedding hidden vectors in a noise vector in some embodiments of the video generation method provided by the embodiments of the present application;

[0073] FIG. 16a is an example diagram of an interface in some embodiments of the video generation method provided by the embodiments of the present application;

[0074] FIG. 16b is an example diagram of an interface in some other embodiments of the video generation method provided by the embodiments of the present application;

[0075] FIG. 17 is a diagram of embedding hidden vectors in a video generation model runtime in some embodiments of the video generation method provided by the embodiments of the present application;

[0076] FIG. 18 is a flowchart of the video generation method provided by the embodiments of the present application from the perspective of a terminal device;

[0077] FIG. 19 is a flowchart of the video generation method provided by the embodiments of the present application from the perspective of a server;

[0078] FIG. 20 is a diagram of a computing device provided by the embodiments of the present application;

[0079] FIG. 21 is a diagram of a computing device cluster provided by the embodiments of the present application;

[0080] FIG. 22 is a diagram of one or more computing devices in a computing device cluster provided by the embodiments of the present application that can be connected through a network. DETAILED DESCRIPTION

[0081] The technical solutions in the embodiments of the present application will be described below with reference to the drawings. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0082] In the description of the embodiments of the present application, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B. The "and / or" in the text only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist together, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.

[0083] Hereinafter, the terms "first", "second" are only for descriptive purposes, and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features.

[0084] In the embodiments of the present application, the words such as "exemplary" or "for example" are used to mean an example, illustration, or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes.

[0085] The traditional video generation method is to obtain a simulation video by using a simulation tool such as a virtual engine (Unreal Engine, UE), Carla, Gazebo, and the like, which needs to consume a large amount of manpower for scene production. For example, as shown in FIG. 1, visual simulation based on a simulation engine such as UE can usually adopt the following process:

[0086] Material modeling: manually modeling or obtaining a plurality of material models through 3D scanning;

[0087] Static scene editing: a static scene is built by editing the positions of a plurality of materials, adding lighting, weather, and the like, and the like;

[0088] Dynamic scene editing: the scene is made dynamic by editing the motion of the materials, the motion of the camera, and the like;

[0089] Rendering: rendering the edited dynamic scene into a video.

[0090] Such a visual simulation scheme based on a simulation engine such as UE needs to consume a large amount of manpower for each step of scene building, and the final rendering result still has a domain gap (a gap between a simulation system and a real world) from a real scene.

[0091] Currently, with the development of artificial intelligence (AI) technology, especially the rapid development of AIGC technology, some AI-based video generation tools have been able to support text-to-video, image-to-video, and video style conversion, and generate high-fidelity videos. For example, in the animation video production scenario, users can quickly produce an animation video with a complete plot by inputting prompt words and animation styles; in the creative video generation scenario, users can upload a picture and input prompt words to describe the picture, and generate a video with rich facial expressions and various motion gestures; in the virtual character production scenario, a lifelike dynamic video can be generated by uploading a static picture and audio.

[0092] Among them, as shown in FIG. 2, it is a text-to-video generation model, which can generate high-quality video content according to descriptive text prompts. The video generation model can control video generation through text and image, or perform operations such as video continuation and style conversion, to directly produce a high-fidelity video.

[0093] For a general video generation model such as that shown in FIG. 2, when a user wants to generate a video, the user generally needs to input a text prompt or upload an image, and the video generation model will generate a corresponding video according to the text and / or image provided by the user. Although a high-fidelity video can be directly produced, the control is limited to text and / or image, and there is a certain limitation on the control of the video.

[0094] In actual applications, users may have more fine-grained needs for video controllability. For example, users may need to accurately control various attributes of a video, such as 3D structure and motion, where the 3D structure can include materials in the foreground and material position information, and the motion can refer to the motion trajectory of the materials in the foreground. Some of these attributes are difficult to describe in natural language, for example, the position of the materials can be a set of coordinate data, which cannot be described by simple text prompts; or most current video generation models have certain requirements for the length of the input text prompt, and a too complex text prompt can cause the video generation model to fail to recognize, and thus fail to generate a corresponding video.

[0095] Due to the limitations of text and image control methods, the generated video has certain randomness, which may deviate from the actual needs of the user, resulting in poor video usability and failing to meet the user's needs for more fine-grained or more accurate control of video generation.

[0096] For example, the user inputs the text: "a dog is flying in the sky", and the video generation model can generate a high-fidelity video of a dog flying in the sky, but if the user wants to precisely control the dog to fly according to a specified irregular trajectory, or control the intensity of the light to change according to a preset value during the flying of the dog, or precisely control the distribution of the clouds in the sky, or whether a specified physical phenomenon occurs at a specified position during the flying of the dog, etc., some current video generation models may not be able to achieve this.

[0097] Or, in actual applications, the user needs to precisely control the parameters of the video attributes, for example, the user needs to precisely control the specific coordinate position of the material, the coordinate of the motion trajectory of the material, and the light parameters, but most current video generation models do not support parameter-level control. In the automatic driving scene, a large number of training videos can be obtained by precisely controlling various parameters.

[0098] Therefore, the embodiments of the present application propose a video generation method and an electronic device, which divide the video required by the user into multiple attributes for separate control, support the user to input corresponding control conditions for multiple attributes respectively, and generate a video meeting the specific control conditions, so as to more accurately match the user's demand and realize more accurate control (parameter level) of the video, thereby improving the controllability of the video.

[0099] For example, the method proposed in the embodiments of the present application can control the position of the dog in each frame of image in the video, so as to obtain an accurately controllable flight trajectory.

[0100] The electronic device proposed in the embodiments of the present application can be a terminal device on the user side, for example, can be a smart phone, a Tablet PC, a laptop, a Desktop computer, a wearable device, an augmented reality (AR) / virtual reality (VR) device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA) device, etc. The embodiments of the present application do not make any limitation on the specific type of the electronic device.

[0101] For example, FIG. 3 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. For example, the electronic device can be a smart phone. As shown in FIG. 3, the electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headset jack 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0102] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than shown, or combine certain components, or split certain components, or different arrangement of components. The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0103] The processor 110 can include one or more processing units. For example, the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices or integrated into one or more processors.

[0104] The controller can generate operation control signals according to instruction operation codes and timing signals to complete the control of fetching and executing instructions.

[0105] The processor 110 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can hold instructions or data that the processor 110 has just used or is cycling through. If the processor 110 needs to use the instructions or data again, it can be called directly from the memory. This avoids repeated access and reduces the latency of the processor 110, thus improving the efficiency of the system.

[0106] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0107] The USB interface 130 is an interface that conforms to the USB standard specification, and can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device 100, and can also be used to transmit data between the electronic device 100 and a peripheral device. It can also be used to connect earphones to play audio through the earphones. The interface can also be used to connect other electronic devices, such as AR devices, etc.

[0108] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a structural limitation on the electronic device 100. In some other embodiments of the present application, the electronic device 100 can also use different interface connection methods or combinations of multiple interface connection methods in the above embodiments.

[0109] The charging management module 140 is configured to receive charging input from a charger. The charger can be a wireless charger or a wired charger. In some embodiments with wired charging, the charging management module 140 can receive charging input from a wired charger through the USB interface 130. In some embodiments with wireless charging, the charging management module 140 can receive wireless charging input through a wireless charging coil of the electronic device 100. The charging management module 140 can charge the battery 142 and power the electronic device 100 through the power management module 141.

[0110] The power management module 141 is configured to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, the internal memory 121, the display 194, the camera 193, and the wireless communication module 160. The power management module 141 can also be configured to monitor parameters such as battery capacity, battery cycle count, battery health (leakage, impedance), and the like. In some other embodiments, the power management module 141 can also be disposed in the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 can also be disposed in the same device.

[0111] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, and the like.

[0112] The antenna 1 and the antenna 2 are configured to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be configured to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in combination with a tuning switch.

[0113] The mobile communication module 150 can provide a solution for wireless communication including 2G / 3G / 4G / 5G, etc. applied to the electronic device 100. The mobile communication module 150 can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves by the antenna 1, and perform filtering, amplification, etc. on the received electromagnetic waves, and transfer to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor, and radiate as electromagnetic waves through the antenna 1. In some embodiments, at least part of the function modules of the mobile communication module 150 can be disposed in the processor 110. In some embodiments, at least part of the function modules of the mobile communication module 150 can be disposed in the same device as at least part of the modules of the processor 110.

[0114] The modem processor can include a modulator and a demodulator. The modulator is configured to modulate a low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is configured to demodulate a received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. The low-frequency baseband signal processed by the baseband processor is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the microphone 170B, etc.), or displays an image or a video through the display screen 194. In some embodiments, the modem processor can be a separate device. In other embodiments, the modem processor can be independent of the processor 110, and disposed in the same device as the mobile communication module 150 or other function modules.

[0115] The wireless communication module 160 can provide a solution for wireless communication including wireless local area networks (WLAN) (e.g., wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc. applied to the electronic device 100. The wireless communication module 160 can be one or more devices that integrate at least one communication processing module. The wireless communication module 160 receives an electromagnetic wave via the antenna 2, frequency-modulates and filters the electromagnetic wave signal, and transmits the processed signal to the processor 110. The wireless communication module 160 can also receive a signal to be transmitted from the processor 110, frequency-modulate it, amplify it, and radiate it as an electromagnetic wave via the antenna 2.

[0116] In some embodiments, the antenna 1 and the mobile communication module 150 of the electronic device 100 are coupled, and the antenna 2 and the wireless communication module 160 are coupled, so that the electronic device 100 can communicate with a network and other devices through wireless communication technology. The wireless communication technology can include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-CDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS can include global positioning system (GPS), global navigation satellite system (GLONASS), beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS). That is, the electronic device 100 has a positioning function and a wireless communication function.

[0117] The electronic device 100 implements a display function through a GPU, a display 194, and an application processor, etc. The GPU is a microprocessor for image processing, which is connected to the display 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs, which execute program instructions to generate or change display information.

[0118] The display screen 194 is configured to display images, videos, and the like. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diodes (QLED), or the like. In some embodiments, the electronic device 100 can include one or N display screens 194, where N is a positive integer greater than 1.

[0119] The electronic device 100 can implement the photographing function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor.

[0120] The ISP is configured to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing to convert it into an image visible to the naked eye. The ISP can also optimize the noise, brightness, and skin color of the image. The ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene. In some embodiments, the ISP can be disposed in the camera 193.

[0121] The camera 193 is configured to capture still images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV, or the like format. In some embodiments, the electronic device 100 can include one or N cameras 193, where N is a positive integer greater than 1.

[0122] The digital signal processor is used to process digital signals, in addition to being able to process digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0123] The video codec is used to compress or decompress digital video. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple encoding formats, such as: moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.

[0124] The NPU is a neural-network (NN) calculation processor, which can quickly process input information by drawing on the structure of a biological neural network, such as drawing on the transmission mode between human brain neurons, and can also constantly self-learn. Through the NPU, the electronic device 100 can realize intelligent cognition applications such as image recognition, face recognition, voice recognition, text understanding, etc.

[0125] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to realize data storage functions. For example, music, video, etc. Files are saved in the external memory card.

[0126] The internal memory 121 can be used to store computer executable program codes, which include instructions. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), etc. The data storage area can store data created during the use of the electronic device 100 (such as audio data, a phonebook, etc.), etc. In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various function applications and data processing of the electronic device 100 by running instructions stored in the internal memory 121 and / or instructions stored in the memory disposed in the processor.

[0127] The electronic device 100 can realize audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the earphone interface 170D, and the application processor, etc. For example, music playing, recording, etc.

[0128] The keys 190 include a power key, a volume key, and the like. The keys 190 can be mechanical keys. Alternatively, the keys 190 can be touch keys. The electronic device 100 can receive a key input and generate a key signal input related to user settings and function control of the electronic device 100.

[0129] The motor 191 can generate a vibration prompt. The motor 191 can be used for incoming call vibration prompts and touch vibration feedback. For example, touch operations for different applications (e.g., taking a photo, playing audio, and the like) can correspond to different vibration feedback effects. Touch operations on different regions of the display screen 194 can correspond to different vibration feedback effects. Different application scenarios (e.g., time reminders, receiving a message, an alarm, a game, and the like) can also correspond to different vibration feedback effects. The touch vibration feedback effects can also be customizable.

[0130] The indicator 192 can be an indicator light and can be used to indicate a charging state, a power change, and the like. Alternatively, the indicator 192 can be used to indicate a message, a missed call, a notification, and the like.

[0131] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation with the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support a Nano SIM card, a Micro SIM card, a SIM card, and the like. Multiple cards can be inserted into the same SIM card interface 195. The types of the multiple cards can be the same or different. The SIM card interface 195 can be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external storage cards. The electronic device 100 interacts with a network through a SIM card to implement functions such as a call and data communication. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.

[0132] The electronic device proposed in the embodiments of the present application can also be a server on the service side. For example, FIG. 4 is a structural schematic diagram of a server in an embodiment of the present application. As shown in FIG. 4, the server 200 can include one or more processors 210, a communication interface 220, a memory 230, and a communication bus 240 connecting different components (including the memory 230, the communication interface 220, and the processor 210).

[0133] The communications bus 240 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration bus, or a local bus using any of a variety of bus architectures. By way of example, and not limitation, the communications bus 240 can include an industry standard architecture (ISA) bus, a micro channel architecture (MCA) bus, an enhanced ISA bus, a video electronics standards association (VESA) local bus, and a peripheral component interconnect (PCI) bus.

[0134] The electronic device typically includes a variety of computer system readable media. These media can be any available media that is accessible by the electronic device and includes both volatile and non-volatile media, removable and non-removable media.

[0135] The memory 230 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The memory 230 can include, without limitation, at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions and / or techniques of embodiments of the application, including the video generation method.

[0136] The program / utility, having a set (at least one) of program modules, can be stored in the memory 230 by way of example, and not limitation, including an operating system, one or more application programs, other program modules, and program data, each or any combination thereof, can include implementation of a network environment. The program modules are generally executed by the processor 210 to implement the functions and / or techniques of embodiments of the application.

[0137] The processor 210 performs a variety of functions and processing of data, such as implementing the video generation method of embodiments of the application, by executing programs stored in the memory 230.

[0138] It should be appreciated that the processor 210 in the server 200 shown in FIG. 4 can be a system on a chip (SOC), which can include a central processing unit (CPU), and can further include other types of processors, such as a graphics processing unit (GPU), etc.

[0139] As shown in FIG. 5, the video generation method proposed in the embodiments of the present application can be applied in various application scenarios such as automatic driving, embodied intelligence, city twinning, video editing, etc. For example, the video editing can be animation video production, creative video generation, virtual character production, etc.

[0140] For example, in the field of automatic driving, the method proposed in the embodiments of the present application supports user input of weather, illumination, driving scene, background semantic map, foreground motion trajectory, camera parameter, and other control conditions, and generates a special driving scene that meets the control conditions, such as generating a video of a vehicle rapidly cutting in, a "ghost probe" phenomenon, etc. The generated video can be used as a training sample to improve the performance of the automatic driving system by simulating different driving scenes.

[0141] For example, as shown in FIG. 5, the user inputs a text prompt "pedestrian walking in a non-zebra crossing area", and inputs control conditions corresponding to various attributes such as structure, motion, environment, and physics in the control conditions, to generate a corresponding video. The video can be used as a sample video to optimize the recognition and response performance of the automatic driving system for the scene of a pedestrian walking in a non-zebra crossing area. In the control conditions corresponding to the structure, the driving scene, the background semantic map, and the camera parameter (camera trajectory) can be input, in the control conditions corresponding to the motion, the foreground motion trajectory can be input, and in the control conditions corresponding to the environment, the weather and the illumination can be input.

[0142] For example, the video generation method proposed in the embodiments of the present application can be implemented based on the system architecture shown in FIG. 6.

[0143] As shown in FIG. 6, the video generation model can be deployed in a terminal device or a cloud server. According to the deployment location of the video generation model, the method proposed in the embodiments of the present application can be implemented based on a system architecture in which the video generation model is deployed in the cloud or a system architecture in which the video generation model is deployed in the terminal device.

[0144] Exemplarily, a system architecture in which the video generation model is deployed in the cloud can include a server 301 and one or more terminal devices 302 in remote communication with the server, for example, the terminal devices can be smartphones, desktop computers, tablet computers, etc. In this system architecture example, the server 301 can be a standalone server or a server cluster, for example, a highly available (HA) cluster, a load balancing (LB) cluster, or a high-performance computing (HPC) cluster. The video generation model is deployed in the server 301, and the terminal device can display controls for user input of control conditions and prompts (text and / or images). After detecting the user input of the control conditions and prompts, the terminal device can send a call request to the server 301, the call request carrying the control conditions and the prompts, etc. The server 301 responds to the call request of the terminal device, inputs the control conditions, prompts, etc. uploaded by the terminal device into the video generation model, runs the video generation model, generates and outputs a video, and distributes the generated video to the terminal device. The terminal device can display or play the video.

[0145] In a system architecture example in which the video generation model is deployed in the terminal device, the terminal device 304 has a lightweight video generation model (e.g., Diffusion) deployed therein, and other models are deployed in the server, for example, a multi-modal large model such as Video LLava. The terminal device 304 can locally call the video generation model to generate a video. It should be noted that when incrementally training the video generation model, other models can need to be called to participate in the training process. Due to the limitations of the computing power and storage resources of the terminal device 304 such as a smartphone, the multi-modal large model and other models are generally deployed in the cloud server 303, and the terminal device 304 needs to remotely call.

[0146] The video generation method proposed in the embodiments of the present application adopts a video generation mechanism for separately controlling at least one attribute of a video, supports a user to set at least one control condition, and each control condition is used to control a corresponding attribute in the video. The attribute of the video refers to an object that can be adjusted and controlled in the video generation or editing process. For example, the at least one attribute can include at least one of the following attributes: structure, motion, environment, and physics. The user can input a corresponding control condition for each attribute to control a corresponding attribute, for example, inputting weather and lighting (which can be text) as a control condition to control the environment attribute, or inputting a background semantic graph (which can be in image format) to control the structure attribute, etc.

[0147] Specifically, the structural attribute can include a material structure and / or a camera track.

[0148] The camera track is a track of changes in the positions of the camera corresponding to respective frames of images in a video. The position of the camera is the position of the camera corresponding to each frame of image if the image is simulated to be taken by the camera, and the positions of the plurality of cameras corresponding to the plurality of frames of images constitute the camera track. The material structure is used to set material information within a frame of image. The camera track is used to set the relative positional relationship (camera view angle) between the plurality of frames of images.

[0149] The material structure refers to various information for describing materials in a frame of image. For example, the material structure can include a material position and / or a material category. The material position is used to set the position information of each material in a frame of image of the video, for example, the boundary coordinates of each material. The material category is the category to which each material belongs, for example, the categories of materials are people, vehicles, animals, etc., and each category can be further divided into a plurality of subcategories, for example, people can be specifically a variety of types such as beauties, children, business men, etc., and vehicles can be divided into large trucks, small cars, tank cars, etc. Alternatively, in some embodiments, the material category can be specifically a material name or a material ID, etc. information, and according to the material name or the material ID, the corresponding material can be found in the material library. The material category can also include other information related to the material, for example, the other information related to the material can be a storage address of the material, etc., or additional information related to the material, for example, the additional information can be information of other items for decorating the material.

[0150] For example, the material structure can be divided into a foreground structure and / or a background structure. The foreground structure can include a foreground position and a foreground material category, and is used to represent the composition of materials in the foreground and the distribution positions of the materials in the image, for example, the foreground information can indicate which materials are included in the foreground, and the boundary coordinates of each material. The background structure is used to represent the composition of materials in the background and the distribution positions of the materials in the image, for example, the background information can indicate which materials are included in the background, and the boundary coordinates of each material.

[0151] For example, as shown in FIG. 7, FIG. 7 shows a frame of image in which a dog wearing sunglasses is flying in the sky, in which, in the foreground structure, the material category is a dog, and the specific position of the material can be the boundary coordinates of the dog or other data information representing the position. In the background structure, the material category is a blue sky and white clouds, and the specific position information of the material can be the boundary coordinates of the blue sky and the white clouds or other data information representing the position.

[0152] For example, when the structure attribute is controlled, control conditions in the form of an image format and / or a text format input by a user are supported. The image format can be a single image or a plurality of images (image sequence).

[0153] For example, in the present embodiment, for the control of the foreground structure, three-dimensional target detection and other technologies can be used to generate image data containing two-dimensional bounding boxes (2D Bounding Box, 2D Bbox) or three-dimensional bounding boxes (3D Bounding Box, 3D Bbox) as control conditions for controlling the foreground structure. Among them, the 2D Bbox or the 3D Bbox can represent the position of the material in the foreground structure. For example, as shown in FIG. 7, the specific position of the dog in the foreground structure can be indicated by an image with a 3D Bbox.

[0154] In the present embodiment, the background structure can be represented by a background semantic map, which can contain the category and coordinate information of each material, for example, the material category can be a vehicle or a human, etc. For example, as shown in FIG. 7, when a video of a dog flying in the sky needs to be generated, a background semantic map can be generated first, which contains the material categories of blue sky and white clouds, and the coordinate information of the blue sky and the white clouds. Then, the background semantic map is uploaded as the control condition of the background structure in the structure attribute.

[0155] For example, the background semantic map can be generated by using a semantic segmentation model or other semantic map model, for example, a semantic segmentation model is used to extract the semantic map corresponding to the sample image or the reference image, and the semantic map is uploaded as the control condition of the background structure.

[0156] In other embodiments, a semantic map of a bird's eye view, i.e., a Bird's Eye View map (BEV map) can be used to represent the background structure, for example, when a drone aerial video needs to be generated, a BEV format background semantic map can be used.

[0157] It should be noted that in other embodiments, the foreground structure can also be represented by a foreground semantic map.

[0158] In addition, the method proposed in the embodiments of the present application is suitable for both 3D video generation and 2D video generation. For the sake of brevity of description, the following description will be mainly based on 3D as an example, which should not be understood as being applicable only to 3D.

[0159] Motion attribute, used to set a motion trajectory of at least one target object in the foreground in the target video. The target object can be one or more of the foreground materials. The motion attribute, also referred to as the motion trajectory, is used to represent the position change trajectory of the foreground material in a video. The position of the foreground material in a frame of image can obtain a set of coordinate data, and the coordinate data corresponding to multiple frames of images can constitute a trajectory. The method proposed in the embodiments of the present application supports setting the position of the foreground material in each frame of image, or supports setting the position of the target object in the foreground every N frames. The value of N can be controlled according to the frame rate of video playing. For example, if the frame rate is 24 frames / s, N can take a value of 3, 5 or 11, and the position of the material is set once every N frames.

[0160] For example, as shown in FIG. 8, the foreground material is a walking young woman. It is assumed that the motion trajectory of the material needs to be controlled to move from P1 to P5 and pass through P2, P3 and P4 in the generated video. A plurality of frames of images containing 3D Bbox can be generated. The 3D Bbox in each frame of image is used to indicate the specific position of the material. For example, the position information of the foreground material is set once every N frames to generate the 3D Bbox corresponding to the 1st frame, the 1+Nth frame, the 1+2Nth frame, the 1+3Nth frame and the 1+4Nth frame of image. According to the boundary coordinates of the three-dimensional bounding box in each frame of image, the accurate position information of the foreground material in the corresponding frame of image can be determined to obtain a video in which the young woman walks along the trajectory P1→P2→P3→P4→P5. In this way, the method proposed in the embodiments of the present application can realize accurate control of the motion trajectory of the specified target object (one or more materials) in the image, and the position information can be set once every frame or every N frames.

[0161] It should be noted that, for the motion trajectory, the input control condition can be an image sequence, for example, an image sequence with 3D Bbox. Alternatively, in some embodiments, for the input control condition of the motion attribute, only the boundary coordinate information corresponding to the 3D Bbox in each frame of image can be used. When setting the control condition, the user can upload a file of a specified format, and the file stores array, vector, matrix and other data used to record the boundary coordinate information of the 3D Bbox in each frame of image.

[0162] Environment, used to set the environmental parameter information in the target video. For example, the environmental parameter can be the weather, the light intensity and other environmental parameters of each frame of image in the video, or the environmental parameter can be the video style and other parameter information.

[0163] Taking the environmental parameter including weather as an example, as shown in FIG. 9, when a video of a dog flying in the sky needs to be generated, the user can set the weather to change from cloudy to rainy and then to sunny, for example, the user can specify that the weather of the first to M-th frame is cloudy, the weather of the M+1-th to M+Q-th frame is rainy, and the weather of the M+Q+1-th to X-th frame is sunny. Wherein, M, Q, and X are all integers, and M+Q≤X.

[0164] It should be noted that in the related art, a corresponding video can be generated by inputting the text prompt “a dog is flying in the sky, and the weather changes from rainy to sunny”, but the position (i.e., which frame) of the weather change in the video cannot be accurately controlled. In some scenarios that require accurate control of environmental parameters, the user's demand cannot be accurately matched. The method proposed in the embodiments of the present application can accurately control the correspondence between the weather and each frame of image, and accurately control the environmental parameters of each frame of image in the video, or control the environmental parameters every N frames.

[0165] Physical property, used to set the occurrence of a specified physical phenomenon in the target video. A physical phenomenon is a kind of phenomenon that occurs based on a certain physical law, and the video generation model can generate a phenomenon that conforms to this physical law by learning the physical law.

[0166] For example, the physical phenomenon can be rigid body motion or non-rigid body motion. A rigid body is an object whose shape and size do not change when subjected to force or motion, and the relative positions of internal points do not change. A non-rigid body is an object whose shape and size change when subjected to force or motion, and the relative positions of internal points change. Rigid body motion has no deformation or deformation velocity. Non-rigid body motion, i.e., has deformation or deformation velocity, for example, linear deformation rate and angular deformation rate. The method proposed in the embodiments of the present application supports user setting whether non-rigid body motion or rigid body motion needs to occur in the video.

[0167] For example, in the automatic driving scene, in order to improve the safety performance of the automatic driving system, some video data needs to be used to train the automatic driving system, and it may be necessary to generate a video of a vehicle collision accident, which can be understood as a non-rigid body motion occurring in the video.

[0168] The physical phenomenon can also be other phenomena that satisfy the physical law, for example, whether an optical phenomenon (shadow of an object, rainbow, emission of a mirror, etc.) occurs, whether a mechanical phenomenon (motion and stillness of an object, deformation and vibration of an elastic object, flow of a fluid) occurs, whether an acoustic phenomenon (for example, whether there is an echo, a change in tone, and a reflection phenomenon of a sound wave) occurs, whether a thermal phenomenon (thermal expansion and contraction, etc., shape change of ice, metal, etc. caused by heating) occurs.

[0169] According to the above description, it can be seen that the control condition can be text, a single image or an image sequence, or a file in other formats that can be uploaded from the file system.

[0170] It should be noted that for the control of multiple attributes, each attribute can be independently controlled, or the control conditions can be associated for combined control of multiple attributes.

[0171] For example, in the automatic driving scenario, to train an automatic driving system, it may be necessary to generate a high-fidelity video simulating a real accident scene. For example, a traffic accident such as a collision, friction, or scratching needs to occur in the video.

[0172] For example, in the automatic driving scenario, to train an automatic driving system, it may be necessary to generate a high-fidelity video simulating a real accident scene. For example, a traffic accident such as a collision, friction, or scratching needs to occur in the video.

[0173] Through the first control condition, the material categories in the material structure include a car and a tree beside the road, and various materials have corresponding IDs. The material positions are the specific positions of the car and the tree beside the road in the image. The second control condition can be associated with the first control condition. According to each material category in the material structure, a car is selected as a control object, and the motion trajectory of the car is controlled. The motion trajectory is a motion trajectory that can interact with the tree. For example, the interaction can be a collision. In addition, through the third control condition, the physical phenomenon that needs to be generated in the video is non-rigid body motion, for example, a collision that causes deformation.

[0174] Through the joint control of the above-mentioned three attributes, the car in the generated video can interact with the tree beside the road when moving along the set motion trajectory, and the car deforms in the interaction.

[0175] For example, the first control condition is also used to control the material categories to include multiple pedestrians, and the second control condition is also used to control the motion trajectories of the multiple pedestrians. Through the association setting of the motion trajectories of the multiple pedestrians and the car, the generated video can make a car along a specified motion trajectory interact with multiple pedestrians moving along another motion trajectory, for example, a collision.

[0176] To facilitate understanding of the video generation method proposed in the embodiments of the present application, the electronic device 100 with the structure shown in FIG. 3 or the server 200 with the structure shown in FIG. 4 will be taken as an example, combined with the application scenario shown in FIG. 5 and the system architecture shown in FIG. 6, to exemplarily explain the video generation method provided by the embodiments of the present application.

[0177] As shown in FIG. 10, FIG. 10 shows a software module architecture, which can be built based on one or more servers 200, or can be built based on the system architecture shown in FIG. 6.

[0178] The embodiment of the application proposes a software module architecture for implementing the video generation method proposed in the embodiment of the application. From the perspective of software implementation, the software module architecture at least includes a controllable video generation module (referred to as a video generation module), and in some embodiments, can further include a quality feedback module. Alternatively, a simulation data generation module can be further included.

[0179] In the embodiment, the video attributes are divided into structure, motion, environment and physics, which are controlled respectively.

[0180] The video generation module is used for encoding (Encoder) the control conditions corresponding to the four attributes of structure (for example, 3D structure), motion, environment and physical law, embedding the video generation model according to the hidden vectors obtained by encoding after inputting the control conditions into the encoder, and training the processed video generation model. The training data can be simulation video data obtained by the simulation data generation module, or other data sets used for training.

[0181] The physical feedback module can be deployed in the video generation module, which is used for evaluating the video generated in the training process to determine whether it meets the preset control condition, giving a corresponding scoring order, and optimizing the trainable parameters in the controllable video generation model according to the score. For example, the training method is supervised learning, the training data is a video with control labels (i.e., control conditions as labels), and the scoring method can be to calculate the value of the loss function, and to optimize the trainable parameters in the video generation model according to the value of the loss function. After training, the video generation model can accept user control in 3D structure, motion, environment, physics and other aspects to generate videos meeting the control conditions.

[0182] The quality feedback module: The quality of the video generated by the video generation model is evaluated by objective indicators and / or subjective indicators. If the quality does not meet the application requirements, incremental training (Incremental Training) is triggered, which can be triggered by simulation data generation, and the generated simulation video data is used to perform incremental training on the video generation model until the quality of the controllable video generation module meets the application requirements.

[0183] The simulation data generation module: The simulation data generation module obtains materials from 3D reconstruction algorithms or existing material libraries, constructs 3D simulation scenes related to application scenarios, and then renders video data with control labels related to application scenarios under the control of control conditions corresponding to attributes such as 3D structure, motion, environment and physics.

[0184] In this way, a video generation model capable of generating videos meeting multiple control conditions can be obtained, and the controllability of video generation is enhanced.

[0185] In the above description of the software module architecture, three stages are involved: the training stage, the use stage, and the incremental training stage. To prevent confusion, as shown in FIG. 11, the entire process is divided into three stages: the training stage, the use stage, and the incremental training stage.

[0186] Training stage:

[0187] In combination with FIGS. 10 and 11, first, based on the training data, the training of the video generation model is performed. The process of the training stage can be implemented on the service side, for example, based on one or more servers 200 in the system architecture shown in FIG. 4.

[0188] The training data includes input data and control conditions (referred to as control labels) as labels. For example, the hidden vector obtained after encoding the control conditions can be used as the label. Alternatively, in other embodiments, the control conditions can be directly used as the label.

[0189] Exemplarily, the training data can be the simulation video data obtained by the simulation data generation module shown in FIG. 10, which can be obtained in the following manner:

[0190] According to the user operation, the video material and the 3D simulation scene are determined. For example, by calling a 3D reconstruction algorithm model, the material required for the video is generated; or according to the user's instruction, the user's required material is searched in the specified material library, or the material is determined according to the user's selection operation. Exemplarily, the 3D reconstruction algorithm model can be a neural radiation field (NeRF) or a 3D Gaussian splatting model.

[0191] After saving the video material, the 3D simulation scene required for the target video can be constructed based on the existing material and in combination with the actual application scene. It should be noted that the 3D simulation scene in the training data can be a 3D simulation scene corresponding to the user operation and / or instruction generated in response to the user operation and / or instruction, which can be understood as a 3D simulation scene manually constructed by the user based on the user operation. Alternatively, the 3D simulation scene can be automatically constructed by a 3D simulation model such as 3D-GPT. 3D-GPT, i.e., Procedural 3D Modeling With Large Language Models, is a 3D modeling model based on a large language model.

[0192] Save or cache the determined material and 3D simulation scene, import the 3D scene containing various 3D materials into the 3D simulation model, receive the control conditions set by the user through the 3D simulation model, and receive control conditions such as 3D structure, motion, environment, physics, and receive text, image, and other general input data that the user may input, render video data related to the application scene according to the input data and control conditions, take the video data as the target video sample, and take the corresponding control conditions as the label to obtain the target video sample with control label.

[0193] Among them, the 3D simulation model can use a rendering, physics engine such as UE and / or Blender, or use a field simulator such as Carla and Gazebo, and specific embodiments of the present application are not listed one by one.

[0194] Next, the control conditions are encoded, that is, the control conditions corresponding to the four attributes of structure, motion, environment, and physics are respectively encoded by an Encoder to obtain the hidden vectors corresponding to various control conditions.

[0195] For example, the encoder can use one or more of the following encoders: STC-encoder, Sparse Condition Encoder, multilayer perceptron (MLP, Multilayer Perceptron, MLP), or other encoders.

[0196] For different control conditions, the same encoder can be used, or different encoders can be used for encoding. The same encoder can encode different control conditions at different times; or multiple encoders are used in parallel to encode multiple control conditions respectively.

[0197] Encode various control conditions into hidden space vectors and embed them in the video generation model. The embedding method can be to add a first variable participating in cross-attention calculation and subsequent model calculation process in the cross-attention layer, and assign the value of the hidden vector to the first variable after obtaining the hidden vector corresponding to the control condition in the training stage or the inference stage. And / or, the embedding method can be to embed a second variable in the noise vector, and assign the value of the hidden vector to the second variable after obtaining the hidden vector corresponding to the control condition in the training stage or the inference stage. The first variable or the second variable can be a vector with the same dimension as the hidden vector, and the elements in the vector are variables.

[0198] In this embodiment, the video generation model can be a diffusion model (Diffusion) or an improved model based on Diffusion, for example, it can be a Video Diffusion Model.

[0199] Or, in other embodiments, the video generation model can also be other video generation models containing cross attention layers, or a combination of diffusion and other video generation models. For example, the video generation model can be an autoregressive model containing cross attention layers, and exemplarily, can be one or more combinations of Temporal Generative Adversarial Net (TGANv2), Video Generation using VQ-VAE and Transformers (VideoGPT), Dual Variational Generation (DVG), and A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2 (StyleGAN-V) based on the generative adversarial network.

[0200] Exemplarily, the embedding of the latent vector can be implemented by one or more of the following embedding methods:

[0201] Embedding method one: embedding through a cross attention layer.

[0202] The embodiments of the present application propose a computer mechanism for embedding control conditions in a cross attention layer.

[0203] The video generation model can be a Diffusion model containing cross attention. The cross attention layer can be used to process the association between multiple different modal sequences, for example, it can be used to perform cross-modal attention calculation between the image modal and the text modal. Specifically, in Diffusion, it can be used to perform cross-modal attention calculation on the intermediate features (or called latent vectors) corresponding to the generated object Image and the embedded vectors Context Embedding corresponding to the text.

[0204] In the embodiments of the present application, the cross attention layer is used to perform cross attention calculation (cross-modal attention calculation) between the latent vectors corresponding to the generated object and the latent vectors corresponding to the control conditions. Wherein, the generated object can be a single image or an image sequence (i.e. a video) including multiple images, and the control conditions can be text, a single image or an image sequence.

[0205] Specifically, in the present embodiment, the following two ways are provided to realize the cross attention calculation between the latent vector of the generated object and the latent vector corresponding to the control condition:

[0206] Cross attention method one:

[0207] If the control condition contains an image or a video (i.e., an image sequence), the latent vector corresponding to the image or video in the control condition can be embedded into the latent vector corresponding to the image or video of the generated object. For example, the embedding can be a Concat of the image or video in the control condition and the image or video of the generated object, followed by cross attention calculation.

[0208] For example, as shown in FIG. 12, the encoder Encoder is used to encode the control condition provided by the user to obtain the latent vector corresponding to the control condition. Assuming that the control condition contains text and image modalities, after encoding, the latent vector corresponding to the text in the control condition Control Context Embedding and the latent vector corresponding to the image in the control condition Control Image Embedding can be obtained. The latent vector corresponding to the image in the control condition Control Image Embedding is embedded into the latent vector of the image Image, which is the generated object. The embedding can be a Concat of Control Image Embedding and Image to obtain an Image embedded with the control condition (hereinafter referred to as the embedded Image).

[0209] It should be noted that the Concat can be a call to the Concat function to connect or combine two objects. The Concat function is only an example, and other functions or other methods can be used instead, such as the concatenate() function. Alternatively, the embedding can be understood as a fusion of image features of Control Image Embedding and Image, and other image feature fusion methods can also be used.

[0210] Next, cross attention calculation can be performed on the embedded Image and the latent vector corresponding to the text in the control condition Control Context Embedding.

[0211] Specifically, the embedded Image is processed by the reshaping layer to generate a Q matrix based on the embedded Image.

[0212] The control condition text corresponding to the hidden vector Control Context Embedding is taken as the text modal data, and is directly input into the linear (Linear) layer corresponding to K (key) and V (value), respectively, to generate K and V matrices based on the text modal data.

[0213] Next, as shown in FIG. 12, for each element in Q, the degree of association between it and the element at the corresponding position in K is calculated to obtain a similarity matrix, and then the calculation formula as described in FIG. 12 is used to calculate After normalizing the similarity matrix, a weight matrix is obtained, which represents the degree of association between the image Q and each element of the text K, and then the weight sum is calculated by multiplying V to obtain new image data. After linear layer processing, the Q matrix, the K matrix and the V matrix have the same dimension, and d represents the dimension of the matrix after linear layer processing.

[0214] Among them, the role of the linear layer (Linear Layer) is to perform linear transformation on the input data of the corresponding modal. The linear layer (Linear Layer) can also be a fully connected layer (Fully Connected Layer) or a dense layer (Dense Layer).

[0215] Cross attention method two:

[0216] As shown in FIG. 13, another way to embed the control condition hidden vector through the cross attention layer can be to set multiple cross attention layers, and embed multiple hidden vectors corresponding to multiple control conditions layer by layer.

[0217] As shown in FIG. 13, for example, it is assumed that the user sets 3 control conditions, including 2 text format control conditions and 1 image format control file, and obtains 3 hidden vectors, which are: Control Context Embedding1, Control Context Embeddin2, and Control Image Embedding.

[0218] At least three cross attention layers can be designed in Diffusion, which are: Cross Attention Layer1, Cross Attention Laye2, and Cross Attention Layer3.

[0219] In the Cross Attention Layer1, cross-modal attention calculation of the generated object Image I0 and the Control Context Embedding1 is performed to obtain the associated image feature Image I1. For details of the calculation process, refer to cross attention mode one.

[0220] In the Cross Attention Layer2, the Image I1 obtained in the Cross Attention Layer1 is taken as one of the modes participating in the cross-modal calculation of the layer, and cross-modal attention calculation of the Image I1 and the Control Image Embedding corresponding to the image in the control condition is performed to obtain the associated image feature Image I2.

[0221] In the Cross Attention Layer3, the Image I2 obtained in the Cross Attention Layer2 is taken as one of the modes participating in the cross-modal calculation of the layer, and cross-modal attention calculation of the Image I2 and the Control Context Embedding2 is performed to obtain the associated image feature Image I3.

[0222] It should be noted that in the above FIG. 12 or FIG. 13, for the sake of clear view, only the case where the generated object Image or the image in the control condition is a single image is shown. In fact, the single image can be an image sequence, and should not be understood as being limited to processing a single image.

[0223] Embedding mode two: concatenated into a noise vector, or replaced by a noise vector.

[0224] The noise vector, i.e., the embedding corresponding to random noise, can also be referred to as the embedding vector corresponding to random noise, or simply referred to as the noise vector.

[0225] As shown in FIG. 14, one implementation is to use the Concat function to connect the random noise embedding vector Noise embedding and the control condition hidden vector Control embedding, and the connected embedding vector is taken as the Noise embedding participating in the subsequent calculation of the model.

[0226] Similarly, the Concat function here is only an example, and other functions or other connection methods can be used for replacement, such as using the concatenate() function, or using other methods that can realize vector splicing or combination.

[0227] As shown in FIG. 15, another implementation manner is that the control condition is encoded to obtain a control embedding, and the random noise embedding is replaced, that is, the control embedding is used as the noise embedding to participate in subsequent calculation.

[0228] The embedding manner 1 and the embedding manner 2 can be implemented simultaneously, or only one of the embedding manners can be used. In the embodiment, the two embedding manners are used.

[0229] It should be noted that the above description of the embedding manner is described by taking diffusion as an example. For other video generation models, the embedding manner can also be embedding the hidden vector corresponding to the control condition in the cross-attention layer and / or embedding the hidden vector corresponding to the control condition in the input layer. The specific embedding manner can be designed according to the above exemplary description and in combination with the hierarchical architecture of the video generation model, and the present specification does not list them one by one.

[0230] After the video generation model is embedded as described above, the video generation model can be trained. It should be noted that embedding has two meanings. One meaning is to modify the hierarchical architecture of the model itself, which can be realized by modifying the code of the model. Another meaning is to input the specific value of the control condition input by the user during the running of the model. In the training stage, that is, in one iteration, the value of the control condition in the training sample is input.

[0231] The input data in the training data is input into the video generation model, and the target video is output by the video generation model.

[0232] As shown in FIGS. 10 and 11, in the training stage, the physical feedback module is used to evaluate whether the target video generated in the training process meets the control condition, and to give a corresponding score ranking, and to optimize the video generation model according to the score.

[0233] The physical feedback module can call a specified model to extract the control condition corresponding to the generated video, compare the extracted control condition with the control condition label, and calculate the value of the loss function. Then, the value of the loss function is used to optimize the trainable parameters in the diffusion.

[0234] For example, as described above, the background structure or the foreground structure in the structural attribute can be represented by a background semantic graph or a foreground semantic graph. The latter does not distinguish between foreground and background, and represents the image structure as a whole by a semantic segmentation graph. Therefore, when the control condition contains a semantic graph, a semantic segmentation model can be used to extract a semantic segmentation graph (referred to as a semantic graph) in the target video, such as a background semantic graph or a foreground semantic graph.

[0235] Wherein, for the foreground structure, a 3D detection model can also be used to detect the material position of each frame image in the target video.

[0236] Then, the extracted semantic graph is compared with the semantic graph in the control label, and the detected material position, for example, the material position can be the bounding box coordinate information detected by the 3D Bbox, is compared with the bounding box information of the 3D Bbox in the control label, to obtain the index score value corresponding to the structure attribute.

[0237] Regarding the motion attribute, the motion trajectory of the material in the foreground of each frame image in the video can be tracked by a 3D tracking model, and compared with the motion trajectory in the control label, to obtain the index score value corresponding to the motion.

[0238] Regarding the environment attribute, the environment parameters in the target video can be detected by a multi-modal understanding model, and compared with the environment parameters in the control label, to obtain the index score value corresponding to the environment. The multi-modal understanding model can include but is not limited to one or more of VideoLLaVA, GPT4Video, etc.

[0239] Regarding the physical attribute, a multi-modal understanding model can be used to detect whether a predetermined physical phenomenon, such as non-rigid body motion, occurs in the target video, and compare the detection result with the physical control condition in the control label, to obtain the index score value corresponding to the physical.

[0240] Alternatively, for both motion and physical attributes, the target video can be 3D reconstructed and then simulated physically to extract the corresponding control condition. For example, the target video is 3D reconstructed to obtain a 3D model, and the 3D model is predicted by physical simulation to predict the deformation of the object in the video or the motion trajectory of the object. For example, a Lagrangian mathematical model can be used to obtain a predicted differential equation, and the differential equation is used to predict whether deformation occurs.

[0241] The value of the loss function can be obtained according to the score values of the indicators corresponding to the plurality of control conditions respectively, for example, the plurality of score values of the indicators corresponding to the plurality of control conditions are weighted and summed, and by setting the weights of various control conditions, the influence degree of each control condition on the video relative to other control conditions can be adjusted. The score value of the indicator can be obtained by using a mean squared error loss function (MSE) or a cross-entropy loss function (CE) or the like to calculate two objects compared. The two objects compared can be the hidden vectors obtained after encoding, that is, the hidden vectors of the control labels after encoding the control conditions. The control conditions extracted from the target video generated by the above-mentioned model can be encoded, and the obtained hidden vectors and the hidden vectors as control labels are calculated by the loss function.

[0242] In this way, according to the above description, the video generation model with expected performance can be obtained by multiple iterations of training. The expected performance refers to the ability to generate a target video that meets the control conditions.

[0243] Usage stage:

[0244] The trained video generation model can be deployed in an electronic device on the user side, for example, in the terminal device 304 on the user side shown in FIG. 6, and the video generation model can be run in a local calling manner to obtain the target video desired by the user. Alternatively, the trained video generation model can be deployed in a server on the service side, for example, in the server 301 shown in FIG. 6, and the terminal device on the user side, for example, the terminal device 302 shown in FIG. 6, can remotely obtain the target video desired by the user through the video generation model deployed in the service 300 in a remote calling manner.

[0245] Exemplarily, as shown in FIG. 16, the interface example shown in FIG. 16 can be displayed in the terminal device on the user side.

[0246] In the interface, a first control 401 for inputting control conditions and a second control 402 for inputting prompt information are displayed.

[0247] The input box for inputting the control condition corresponding to each of the structure, motion, environment, and physics and the like is displayed through the first control 401, and the input box supports multi-modal input, for example, text input and uploading files. The files can be images in various image formats, for example, semantic segmentation images, 3D Bbox images, and the like, and other formats of data or various formats of files in the file system can also be supported.

[0248] For example, the coordinate data corresponding to the 3D Bbox image shown in FIG. 8 can be stored in a file in a specified format. The user only needs to click the right side "+" in the input box corresponding to the structure attribute shown in FIG. 16, select the file used to save the coordinate data, and upload it.

[0249] For another example, the motion trajectory can be represented in data formats such as array, tensor, or matrix. The input box corresponding to the motion can upload a file containing array or tensor, matrix, etc.

[0250] In this embodiment, the existing video generation method can be compatible. In the existing video generation method, the user can input text and / or image. For example, in the interface example shown in FIG. 16, the second control 402 for inputting text and / or uploading image is displayed, and the text prompt input by the user in the prompt input box is "a man and a woman cross the road without walking the zebra crossing".

[0251] It should be noted that in the scheme provided in the embodiments of the present application, the user can not set the control condition, or select one or more from a plurality of control conditions. For example, in the interface example shown in FIG. 16, the user only inputs the text information "non-rigid motion" in the input box corresponding to the physical attribute. In other embodiments provided by the present application, the user needs to input at least two of the control conditions corresponding to the structure attribute, the motion attribute, the environment attribute and the physical attribute, etc. For example, the user needs to input the control condition corresponding to the structure attribute and the control condition corresponding to the motion attribute, so that at least one material in the target video moves along the motion trajectory defined by the motion attribute according to the initial position defined by the structure attribute; or the user needs to input the control condition corresponding to the motion attribute and the control condition corresponding to the physical attribute, so that at least one material in the target video moves according to the motion trajectory defined by the motion attribute with the physical phenomenon defined by the physical attribute. And so on. The number of control conditions that the user needs to input is not limited in the embodiments of the present application.

[0252] According to the control condition and the text prompt input by the user, the encoder is called to encode the control condition "non-rigid motion" input by the user to obtain the latent vector. And the text prompt "a man and a woman cross the road without walking the zebra crossing" input by the user is processed to obtain the embedded vector.

[0253] Next, the electronic device 303 (desktop computer) on the user side calls the local video generation model or remotely calls the video generation model on the server side to obtain the target video output by the video generation model.

[0254] In the local running of the video generation model, the hidden vector corresponding to the control condition and the embedding vector corresponding to the text prompt can be input into the video generation model. In the remote call, the hidden vector corresponding to the control condition and the embedding vector corresponding to the text prompt and other data can be carried in the call request.

[0255] In response to the call request, the video generation model starts running, as shown in FIG. 17. During the running process, the hidden vector corresponding to the control condition is input into the input layer and the cross attention layer. Specifically, in the input layer of Diffusion, the hidden vector of the control condition is concatenated with the noise vector through the concat function to obtain the embedded noise vector; and in the cross attention layer of Diffusion, the value of the hidden vector control context embedding obtained by encoding the control condition is embedded into the context embedding.

[0256] The video generation model outputs the target video that the user wants.

[0257] Next, the quality feedback module is used to score the quality of the target video.

[0258] Specifically, the quality feedback module can comprehensively evaluate the video quality generated by the controllable video generation model and whether the compliance of the control condition meets the application requirements through subjective indicators and objective indicators. If the application requirements are met, the process ends; if the application requirements are not met, as shown in FIG. 11, the incremental training of the video generation model is triggered. Specifically, the subsequent simulation data generation can be triggered for the incremental training of the controllable video generation model.

[0259] For example, the subjective evaluation index can be calculated based on user operations through an interactive interface. For example, multiple videos generated under controllable conditions are displayed on the interactive interface, and the user selects the available videos, and the available proportion is calculated. If the proportion is higher than a certain threshold, it is considered that the application requirements are met.

[0260] The objective evaluation index can include a video quality evaluation index (Frechet Video Distance, FVD), a CLIP (Contrastive Language-Image Pre-training) Similarity, and other control condition compliance indicators. The FVD can use Inflated-3D Convnets (I3D) pre-trained on Kinetics to extract features from video clips, and calculate the FVD score by calculating the combination of the mean and covariance matrix.

[0261] The CLIP similarity can be a cosine similarity between each frame image in the target video and the text input by the user, calculated using the CLIP model, to realize a cross-modal similarity measurement method between the text input by the user and the output video.

[0262] The total score value obtained by the quality feedback module can be a comprehensive calculation result of the subjective evaluation index score value and the objective evaluation index score value, for example, a weighted sum of the subjective evaluation score value and the objective evaluation score value.

[0263] It should be noted that in some embodiments, the objective index can also include an evaluation index of the physical feedback module used in the training stage for the control condition, that is, in the operation of the quality feedback module, the physical feedback module can be called to extract the control condition corresponding to the current output target video, compare it with the control condition input by the user, and further evaluate whether the generated video can meet the control condition set by the user.

[0264] If the quality score evaluated by the quality feedback module does not reach the threshold value, the incremental training process is triggered, and the video generation model is iteratively trained in the incremental training stage until the quality score is equal to or greater than the threshold value.

[0265] Incremental training stage:

[0266] In the incremental training stage, according to the control condition and text actually input by the user, new training data is obtained through the 3D simulation module, the actual input control condition of the user is taken as a label, and the embedding vector corresponding to the text, the noise vector of the embedding control condition, and the vector corresponding to the time step are input into the video generation model to output the target video.

[0267] In the incremental training stage, the training process of the video generation model can refer to the training stage, which will not be described here.

[0268] It should be noted that in this embodiment, in the training stage, the value of the loss function is calculated by the physical feedback module, and the trainable parameters in the video generation model are optimized according to the value of the loss function until the model converges. In the incremental training stage, a quality score is evaluated by the quality feedback module, and if the quality score is lower than the threshold value, the video generation model is iteratively trained until the quality score is equal to or greater than the threshold value, and the incremental training is stopped. When the next obtained quality score fails to reach the threshold value, the incremental training process is triggered again.

[0269] It can be seen that the embodiment of the present application proposes a method capable of stimulating the controllability of the video generation model, so that a controllable video generation model can simultaneously accept various combinations of control conditions, generate videos that are more matched with user needs, and match user needs in a more fine-grained manner.

[0270] In addition, the iterative optimization system of controllable video generation and visual simulation proposed in the embodiment of the present application provides effective guarantee for the controllability compliance degree of the controllable video generation model, and guarantees the performance of the video generation model and the usability of the generated video.

[0271] According to the above exemplary description, from the perspective of the electronic device on the user side, the video generation method provided by the embodiment of the present application can include the following flow as shown in FIG. 18:

[0272] S10: receiving at least one control condition.

[0273] Exemplarily, in some embodiments, receiving at least one control condition can be through a visual interface, for example, the electronic device displays a first control, and receives the text input by the user or the uploaded file through the first control.

[0274] The electronic device can be the electronic device 100 shown in FIG. 3, or the terminal device 302 or the terminal device 304 shown in FIG. 6.

[0275] As shown in FIG. 16a, the first control 401 is used to receive at least one control condition set by the user, and the user operation includes inputting text and uploading files. For example, various attributes such as structure attributes correspond to input boxes respectively displaying “Please enter or upload files here”, the user can select a control condition in the form of input text, or select a file and upload it by clicking the “+” button on the right.

[0276] It should be noted that in other embodiments, a visual interface can not be set, but an application programming interface (API) can be set, and the user can indicate at least one control condition in the process of calling the API, and a visual interface such as shown in FIG. 16a or 16b can not be displayed.

[0277] S11: The electronic device encodes the at least one control condition into at least one latent vector.

[0278] After determining the at least one control condition set by the user through the interface shown in FIG. 16a or in other manners, an encoder can be invoked to encode the various control conditions respectively to obtain at least one latent vector. The encoder can be deployed on the server side or in the electronic device on the user side. The electronic device can invoke the encoder locally or remotely. The encoder that can be invoked can refer to the example description above, which will not be repeated here.

[0279] The at least one control condition is used to control at least one attribute of the target video.

[0280] S12: input the at least one latent vector into the pre-trained video generation model, and generate the target video meeting the at least one control condition through the video generation model.

[0281] It has been mentioned in the above example description that the video generation model can be deployed in the server or in the electronic device on the user side, so that the electronic device on the user side can invoke the pre-trained video generation model locally or remotely and input the at least one latent vector into the video generation model.

[0282] According to the above example description, the video generation model can be Diffusion, and the at least one latent vector can be embedded in the cross attention layer or embedded in the noise vector corresponding to the random noise. The input data of Diffusion can include the embedding vector corresponding to the text and / or image, the noise vector and the time step vector. The noise vector corresponding to the random noise, in which the at least one latent vector is embedded, can be connected with the noise vector by using a function such as concat, or the at least one latent vector can replace the original random noise noise vector as the noise vector.

[0283] In some embodiments, the at least one attribute in the target video can be one or more of the following attributes: structural attribute, motion attribute, environmental attribute and physical attribute.

[0284] The structural attribute can be the material structure and / or camera track in the target video, for example, the background semantic graph and 3D Bbox in FIG. 7, which can be used to control the material structure.

[0285] The motion attribute can be a motion trajectory of at least one material in the foreground in the target video, for example, the coordinate data corresponding to the 3D Bbox of each frame in the first frame to the first+4N frame in FIG. 8, which can be used as a control condition to control the motion trajectory of the female character material shown in FIG. 8.

[0286] The environmental attribute can be environmental parameter information in the target video, for example, the environmental parameter can include weather, and in the M+1th to M+Qth frame shown in FIG. 9, the weather is controlled to be rainy, and in the M+Q+1th to Xth frame, the weather is controlled to be sunny.

[0287] The physical attribute can be a physical phenomenon appearing in the target video. For example, by inputting a text of the physical phenomenon, a physical phenomenon corresponding to the text is controlled to appear in the video. For example, in FIG. 16a, a text "non-rigid body motion" is input in the input box corresponding to the physical attribute, and a non-rigid body motion is controlled to appear in the video. For example, in the generated video, an accident occurs when a man and a woman cross the road without passing through the zebra crossing, and the vehicle is deformed.

[0288] It should be noted that the above-mentioned various controllable attributes are only examples, and other division methods of various controllable attributes can also be obtained in other ways. For example, it can be divided into two kinds of controllable attributes, intra-frame and inter-frame. The intra-frame attribute is mainly used to set the characteristics of various materials in a picture frame, such as material structure, environment, etc. The inter-frame is mainly used to set the relative position relationship between multiple frames, such as motion trajectory, camera trajectory, etc. which can be used as inter-frame attribute. Or, it can also be divided into static attribute or dynamic attribute, etc. The division method of the attribute is not listed one by one in this specification.

[0289] In addition, the video generation method proposed in the embodiments of the present application can be compatible with the existing text / image control method in some embodiments, that is, it supports text / image control and also supports setting control conditions. As shown in FIG. 16a, a first control 4021 and a second control 402 are displayed on the interface. The second control is used to determine the input data of the video generation model according to the user operation, and the input data includes the text prompt input in the "prompt" input box and / or the image uploaded through the "image" button.

[0290] In other embodiments, as shown in FIG. 16b, only the user setting control condition can be supported, and the user setting text / image is not supported, and only the first control 401 is displayed. In the interface example shown in FIG. 16, for the input box of the structure attribute, the user can input the text "foreground: two pedestrians, a man and a woman; background: a street with a zebra crossing, trees on both sides of the street", and in the input box corresponding to the motion attribute, the user can input "two pedestrians crossing the street" or upload coordinate data or images for accurately controlling the motion trajectory of the two pedestrians.

[0291] Thus, the method proposed in the embodiments of the present application displays at least the first control 401 on the electronic device.

[0292] It should be noted that the interface shown in FIG. 16a or FIG. 16b is only an example, and other various interface layout styles can be set in actual applications, and are not limited to the examples shown in FIG. 16a or FIG. 16b. It should be noted that in actual applications, an interactive interface can also not be set, but a data interface for calling by a user can be provided, for example, an API calling interface is provided, and the user can obtain a video meeting the control condition by calling the API after obtaining the control condition.

[0293] In the embodiments of the present application, the setting of the control condition can involve 3D space and motion, and thus a tool with a 3D editing interface can be used to generate the control condition, for example, the blender can be used to edit the motion track.

[0294] In addition, it should be noted that the method proposed in the embodiments of the present application can be used as a plug-in compatible with some existing tools. For example, the software program product obtained based on the method proposed in the embodiments of the present application can be a plug-in, which can be installed into an existing 3D simulation tool with a 3D editing interface or a video generation tool with a video editing interface. When the user uses these tools, the user sets the control condition through the editing interface of the tool itself, and then runs the plug-in. The plug-in has the permission to read the control condition. After the plug-in reads the control condition, the video data meeting the control condition is generated, and the video data can be played in the 3D simulation tool or the video generation tool.

[0295] In the method flow shown in FIG. 18, the video generation model is a trained model. In some embodiments, the video generation model needs to be trained in advance. Specifically, the video generation model can be trained in the following manner before the first control is displayed:

[0296] In combination with FIG. 10 and FIG. 11, the training data with labels is first obtained. The label can be at least one first hidden vector corresponding to at least one control condition.

[0297] After obtaining the training data, in a step of training the video generation model, the value of the loss function corresponding to the target video output by the video generation model can be obtained by calling the physical feedback module, and the trainable parameters in the video generation model are optimized according to the value of the loss function. Specifically, at least one control condition can be extracted from the target video output by the video generation model, and at least one second hidden vector corresponding to the at least one control condition extracted from the target video can be obtained, for example, the at least one control condition extracted from the target video can be encoded to obtain the at least one second hidden vector.

[0298] The value of the loss function is calculated according to the at least one second hidden vector and the at least one first hidden vector, for example, the loss function can be Cross Entropy Loss or other loss functions, which are not listed one by one here. The trainable parameters in the video generation model are optimized according to the value of the loss function. Such iteration is performed until the model converges.

[0299] Exemplarily, at least one control condition can be extracted from the target video output by the video generation model, which can be calling a specified model to extract the corresponding control condition, for example, a semantic segmentation model and / or a 3D detection model can be called to extract the structure corresponding control condition from the target video output by the video generation model through the semantic segmentation model and / or the 3D detection model. For example, a 3D tracking model can be called to extract the motion corresponding control condition from the target video output by the video generation model through the 3D tracking model.

[0300] For example, a multi-modal understanding model can be called to extract the environment corresponding control condition from the target video output by the video generation model through the multi-modal understanding model.

[0301] For another example, a multi-modal understanding model can be called to extract the physical corresponding control condition from the target video output by the video generation model through the multi-modal understanding model.

[0302] It should be noted that in the training stage or the incremental training stage shown in FIG. 11, the training data with labels can be training data obtained through 3D simulation. For example, the material and the background used to generate the sample video are determined first, and the background can be a 3D simulation background or a 2D background. The simulation model is called to import the material and the background into the simulation model, and the simulation model receives at least one control condition set by the user to generate a sample video (also referred to as a target video sample) conforming to the at least one control condition. Next, the encoder is called to encode the at least one control condition to obtain at least one first hidden vector, and the at least one first hidden vector is used as a label of the sample video to obtain the training data with labels.

[0303] In addition, in the use stage shown in FIG. 11, the generated target video is also evaluated (assessed) by a quality feedback module according to subjective indicators and / or objective indicators to obtain a quality score. In the case where the quality score is lower than a threshold, the incremental training of the video generation model is triggered. The evaluation result of the subjective indicators is obtained according to user operations. For example, the evaluation opinions of the user on the generated video can be collected by interacting with the user through an interactive interface. For example, the evaluation opinions can be whether the user is satisfied or whether the generated video meets the expectations. The objective indicators include a video quality evaluation indicator FVD and / or a multi-modal model similarity indicator CLIP Similarity. For specific evaluation methods, please refer to the above description.

[0304] As shown in FIG. 19, the embodiment of the present application also provides a video generation method. The method can be applied to a server. The server is deployed with a pre-trained video generation model. For example, the server can be the server 200 shown in FIG. 4, or the server 301 or the server 303 shown in FIG. 6. The method can include the following steps.

[0305] S20: The server receives a calling request sent by a terminal device.

[0306] The calling request can include at least one control condition.

[0307] S21: In response to the calling request, a video generation model is run to generate a target video that meets the at least one control condition.

[0308] The video generation model is embedded with at least one hidden vector. The at least one hidden vector can be obtained based on the at least one control condition, for example, by encoding the control condition. The at least one control condition is used to control at least one attribute in the target video.

[0309] S22: The target video is sent to the terminal device.

[0310] The present application also provides a video generation device, which includes the following modules.

[0311] A first receiving module is configured to receive at least one control condition. The at least one control condition is used to control at least one attribute of a target video to be generated. The at least one attribute includes one or more of the following: a structural attribute, a motion attribute, and a physical attribute.

[0312] An encoding module is configured to encode the at least one control condition into at least one hidden vector.

[0313] A generation module is configured to input the at least one hidden vector into a pre-trained video generation model, and generate a target video that meets the at least one control condition through the video generation model.

[0314] The first receiving module, the encoding module, and the generating module can be implemented by software or by hardware. For example, the implementation of the first receiving module is described below. Similarly, the implementation of the encoding module and the generating module can refer to the implementation of the first receiving module.

[0315] As an example of a software functional unit, the first receiving module can include code running on a computing instance. The computing instance can include at least one of a physical host (computing device), a virtual machine, and a container. Further, the computing instance can be one or more. For example, the first receiving module can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code can be distributed in the same region, or in different regions. Further, the multiple hosts / virtual machines / containers for running the code can be distributed in the same availability zone (AZ), or in different AZs. Each AZ includes one data center or multiple data centers in close geographical proximity. Generally, one region can include multiple AZs.

[0316] Similarly, the multiple hosts / virtual machines / containers for running the code can be distributed in the same virtual private cloud (VPC), or in multiple VPCs. Generally, one VPC is set in one region, and a communication gateway needs to be set in each VPC for cross-region communication between two VPCs in the same region or between VPCs in different regions, and the interconnection between VPCs is realized through the communication gateway.

[0317] As an example of a hardware functional unit, the first receiving module can include at least one computing device, such as a server. Alternatively, the first receiving module can be a device implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD can be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0318] The multiple computing devices included in the first receiving module can be distributed in the same region or in different regions. The multiple computing devices included in the first receiving module can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the first receiving module can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0319] It should be noted that in other embodiments, the first receiving module can be configured to perform any step of the video generation method (e.g., the method shown in FIG. 18), the encoding module can be configured to perform any step of the video generation method shown in FIG. 18, and the generating module can be configured to perform any step of the video generation method shown in FIG. 18. The steps implemented by the first receiving module, the encoding module, and the generating module can be specified as needed, and the entire function of the video generation device can be implemented by the first receiving module, the encoding module, and the generating module implementing different steps of the video generation method shown in FIG. 18.

[0320] As an example of a software functional unit, the video generation device can include code running on a computing instance. The computing instance can be at least one of a physical host (computing device), a virtual machine, a container, or the like. Further, the computing device can be one or more. For example, the video generation device can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the application can be distributed in the same region or in different regions. The multiple hosts / virtual machines / containers used to run the code can be distributed in the same AZ or in different AZs, and each AZ includes one data center or multiple data centers in close geographical proximity. Typically, one region can include multiple AZs.

[0321] Similarly, the multiple hosts / virtual machines / containers used to run the code can be distributed in the same VPC or in multiple VPCs. Typically, one VPC is set up in one region. Communication between two VPCs in the same region and between VPCs in different regions requires a communication gateway to be set up in each VPC to achieve interconnection between VPCs.

[0322] As an example of a hardware functional unit of a module, the video generation apparatus can include at least one computing device, such as a server or the like. Alternatively, the video generation apparatus can also be a device implemented with an ASIC, or a device implemented with a PLD, or the like. The PLD can be a CPLD, an FPGA, a GAL, or any combination thereof.

[0323] The plurality of computing devices included in the video generation apparatus can be distributed in the same region, or can be distributed in different regions. The plurality of computing devices included in the video generation apparatus can be distributed in the same AZ, or can be distributed in different AZs. Similarly, the plurality of computing devices included in the video generation apparatus can be distributed in the same VPC, or can be distributed in multiple VPCs. The plurality of computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, and GALs, or the like.

[0324] The present application also provides a computing device 400. As shown in FIG. 20, the computing device 400 includes a bus 402, a processor 404, a memory 406, and a communication interface 408. The processor 404, the memory 406, and the communication interface 408 communicate with each other through the bus 402. The computing device 400 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 400.

[0325] The bus 402 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, only one line is shown in FIG. 3, but it does not mean that there is only one bus or only one type of bus. The bus 402 can include a path for transmitting information between various components (e.g., the memory 406, the processor 404, the communication interface 408) of the computing device 400.

[0326] The processor 404 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), or the like.

[0327] The memory 406 can include volatile memory, such as random access memory (RAM), and non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid-state drive (SSD).

[0328] The executable program code is stored in the memory 406, and the processor 404 executes the executable program code to realize the functions of the aforementioned first receiving module, the encoding module, and the generating module, respectively, so as to realize the video generation method shown in FIG. 18. That is, the instructions for executing the video generation method shown in FIG. 18 are stored in the memory 406.

[0329] The communication interface 408 uses a transceiving module such as, but not limited to, a network interface card and a transceiver to realize the communication between the computing device 400 and other devices or communication networks.

[0330] The embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a notebook computer, or a smart phone.

[0331] As shown in FIG. 21, the computing device cluster includes at least one computing device 400. The memory 406 in one or more computing devices 400 in the computing device cluster can store the same instructions for executing the video generation method shown in FIG. 18.

[0332] In some possible implementations, the memory 406 of one or more computing devices 400 in the computing device cluster can also respectively store partial instructions for executing the video generation method shown in FIG. 18. In other words, the combination of one or more computing devices 400 can collectively execute the instructions for executing the video generation method shown in FIG. 18.

[0333] It should be noted that the memories 406 in different computing devices 400 in the computing device cluster can store different instructions, respectively, for executing partial functions of the video generation apparatus. That is, the instructions stored in the memories 406 in different computing devices 400 can realize the functions of one or more of the first receiving module, the encoding module, and the generating module.

[0334] In some possible implementation, one or more of the computing devices in the cluster of computing devices can be connected through a network. In some possible implementation, the network can be a wide area network, a local area network, or the like. FIG. 22 illustrates one possible implementation. As shown in FIG. 22, two computing devices 400A and 400B are connected through a network. Specifically, the computing devices are connected to the network through a communication interface in each of the computing devices. In this type of possible implementation, the memory 406 in the computing device 400A stores instructions for performing the functions of the first receiving module. Meanwhile, the memory 406 in the computing device 400B stores instructions for performing the functions of the encoding module and the generating module.

[0335] The connection between the computing devices in the cluster of computing devices shown in FIG. 22 can be such that the functions performed by the encoding module and the generating module are performed by the computing device 400B, considering that the method provided in the present application requires a large amount of data storage and the like. For example, the function performed by the encoding module can be encoding at least one control condition into at least one latent vector, and the function performed by the generating module can be generating a target video in which the first material interacts with the second material along a motion trajectory. The target video includes the first material and the second material, the at least one attribute includes a structure attribute and a motion attribute, the structure attribute includes a material structure of the first material and the second material, and the motion attribute includes a motion trajectory of the first material.

[0336] It should be understood that the functions of the computing device 400A shown in FIG. 22 can also be performed by a plurality of computing devices 400. Similarly, the functions of the computing device 400B can also be performed by a plurality of computing devices 400.

[0337] The embodiments of the present application also provide another cluster of computing devices. The connection between the computing devices in the cluster of computing devices can be similar to the connection between the computing devices in the cluster of computing devices described with reference to FIG. 21 and FIG. 22. The difference is that the memory 406 in one or more of the computing devices 400 in the cluster of computing devices can store the same instructions for performing the video generation method in the corresponding embodiments of FIG. 10 or FIG. 11.

[0338] In some possible implementation, the memory 406 in one or more of the computing devices 400 in the cluster of computing devices can also respectively store partial instructions for performing the video generation method in the corresponding embodiments of FIG. 10 or FIG. 11. In other words, the combination of one or more of the computing devices 400 can collectively execute the instructions for performing the video generation method in the corresponding embodiments of FIG. 10 or FIG. 11.

[0339] It should be noted that the memories 406 in different computing devices 400 in the computing device cluster can store different instructions for performing part of the functions of the video generation apparatus. That is, the instructions stored in the memories 406 in different computing devices 400 can implement the functions of the video generation apparatus.

[0340] The embodiments of the present application further provide a computer program product containing instructions. The computer program product can be a software or program product containing instructions, which can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device is caused to perform the video generation method.

[0341] The embodiments of the present application further provide a computer readable storage medium. The computer readable storage medium can be any available medium that the computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk), etc. The computer readable storage medium contains instructions, which instruct the computing device to perform the video generation method, or instruct the computing device to perform the video generation method.

[0342] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. A video generation method, characterized in that, The method includes: Receive at least one control condition, the at least one control condition being used to control at least one attribute of the target video to be generated, the at least one attribute including one or more of the following: structural attribute, motion attribute, physical attribute; Encode the at least one control condition into at least one hidden vector; The at least one latent vector is input into a pre-trained video generation model, and the target video that meets the at least one control condition is generated by the video generation model.

2. The method as described in claim 1, characterized in that, The target video includes at least one piece of material. The structural attributes include the material structure and / or camera trajectory in the target video, wherein the material structure includes the location or category of at least one material. The motion attributes include the motion trajectory of at least one element in the target video; The physical properties include the physical phenomena of at least one material in the target video during motion.

3. The method as described in claim 2, characterized in that, The target video includes a first material and a second material. The at least one attribute includes a structural attribute and a motion attribute. The structural attribute includes the material structure of the first material and the second material. The motion attribute includes the motion trajectory of the first material. The target video includes a video in which the first material interacts with the second material along the motion trajectory.

4. The method as described in claim 3, characterized in that, The at least one attribute also includes physical attributes, and the target video includes a video in which the first material interacts with the second material along the motion trajectory using the physical phenomenon.

5. The method according to any one of claims 1-4, characterized in that, The at least one attribute further includes an environmental attribute, which includes environmental parameter information of the target video, and the environmental parameter includes at least one of weather and lighting.

6. The method according to any one of claims 1-5, characterized in that, Receive at least one control condition, including: A first control is displayed, through which control conditions sent by the user are received; wherein, the control conditions indicate control over the structural attribute, and the control conditions include one or more of text and a file with a predetermined format; or, the control conditions indicate control over the motion attribute, and the control conditions include one or more of text and a file with a predetermined format; or, the control conditions indicate control over the physical attribute, and the control conditions include one or more of text and a file with a predetermined format.

7. The method according to any one of claims 1-6, characterized in that, The method further includes: Display a second control; the second control is used to receive prompt information input by the user, the prompt information being used to guide the generation of the target video, the prompt information including text and / or images.

8. The method according to any one of claims 1-7, characterized in that, The video generation model is a diffusion model. The diffusion model includes a cross-attention layer; the cross-attention layer includes at least one first variable; Inputting the at least one latent vector into a pre-trained video generation model includes: The at least one latent vector is input into the cross-attention layer of the pre-trained video generation model, and the value of the at least one latent vector is assigned to the at least one first variable; And / or, The input layer of the diffusion model includes a noise vector corresponding to random noise; at least one second variable is embedded in the noise vector; Inputting the at least one latent vector into a pre-trained video generation model includes: The value of at least one latent vector is assigned to the second variable, and the noise vector embedded with the second variable is input into the pre-trained video generation model.

9. The method as described in claim 8, characterized in that, The noise vector embeds at least one second variable, including: The noise vector is connected to the at least one second variable; or, the at least one second variable serves as the noise vector.

10. The method according to any one of claims 1-9, characterized in that, Before receiving at least one control condition, the method further includes: Obtain labeled training data; wherein the label is at least one first hidden vector obtained by at least one control condition encoding; In one iteration of training the video generation model, the method further includes: Extract at least one control condition from the target video output by the video generation model; Obtain at least one second hidden vector corresponding to at least one control condition extracted from the target video; The value of the loss function is calculated based on the at least one second hidden vector and the at least one first hidden vector; Based on the value of the loss function, the trainable parameters in the video generation model are optimized.

11. The method as described in claim 10, characterized in that, Extract at least one control condition from the target video output by the video generation model, including performing one or more of the following steps: Call the semantic segmentation model and / or the 3D detection model, and extract the control conditions corresponding to the structure from the target video output by the video generation model through the semantic segmentation model and / or the 3D detection model; The 3D tracking model is invoked, and the control conditions corresponding to the motion are extracted from the target video output by the video generation model through the 3D tracking model. The multimodal understanding model is invoked, and the control conditions corresponding to the environment are extracted from the target video output by the video generation model through the multimodal understanding model. The multimodal understanding model is invoked, and the physical corresponding control conditions are extracted from the target video output by the video generation model through the multimodal understanding model.

12. The method as described in claim 10 or 11, characterized in that, Obtain labeled training data, including: Determine the source material and background to be used to generate the sample video; Run the simulation model and import the materials and background into the simulation model; If at least one control condition sent by the user is detected, a sample video conforming to the at least one control condition is generated through the simulation model. The encoder is invoked to encode the at least one control condition, thereby obtaining at least one first hidden vector; The at least one first hidden vector is used as the label of the sample video to obtain labeled training data.

13. The method according to any one of claims 1-12, characterized in that, After generating a target video that meets at least one of the control conditions, the method further includes: The quality feedback module is invoked to obtain the quality score of the target video. The quality score is obtained based on the evaluation results of subjective indicators and / or objective indicators. The evaluation results of the subjective indicators are obtained based on user operations. The objective indicators include the video quality assessment indicator FVD and / or the multimodal model similarity indicator CLIP Similarity. If the quality score is below a threshold, incremental training of the video generation model is triggered.

14. A video generation apparatus, characterized in that, The device includes: A first receiving module is configured to receive at least one control condition, wherein the at least one control condition is configured to control at least one attribute of the target video to be generated, wherein the at least one attribute includes one or more of the following: structural attributes, motion attributes, and physical attributes; An encoding module is used to encode the at least one control condition into at least one hidden vector; A generation module is used to input the at least one latent vector into a pre-trained video generation model, and generate the target video that meets the at least one control condition through the video generation model.

15. The apparatus as claimed in claim 14, characterized in that, The target video includes at least one piece of material. The structural attributes include the material structure and / or camera trajectory in the target video, wherein the material structure includes the location or category of at least one material. The motion attributes include the motion trajectory of at least one element in the target video; The physical properties include the physical phenomena of at least one material in the target video during motion.

16. The apparatus as claimed in claim 15, characterized in that, The target video includes a first material and a second material. The at least one attribute includes a structural attribute and a motion attribute. The structural attribute includes the material structure of the first material and the second material. The motion attribute includes the motion trajectory of the first material. The target video includes a video in which the first material interacts with the second material along the motion trajectory.

17. The apparatus as claimed in claim 16, characterized in that, The at least one attribute also includes physical attributes, and the target video includes a video in which the first material interacts with the second material along the motion trajectory using the physical phenomenon.

18. The apparatus as claimed in any one of claims 14-17, characterized in that, The at least one attribute further includes an environmental attribute, which includes environmental parameter information of the target video, and the environmental parameter includes at least one of weather and lighting.

19. The apparatus as claimed in any one of claims 14-18, characterized in that, When receiving at least one control condition, the first receiving module is specifically used for: A first control is displayed, through which control conditions sent by the user are received; wherein, the control conditions indicate control over the structural attribute, and the control conditions include one or more of text and a file with a predetermined format; or, the control conditions indicate control over the motion attribute, and the control conditions include one or more of text and a file with a predetermined format; or, the control conditions indicate control over the physical attribute, and the control conditions include one or more of text and a file with a predetermined format.

20. The apparatus according to any one of claims 14-19, characterized in that, The device further includes a second receiving module, the second receiving module being used for: Display a second control; the second control is used to receive prompt information input by the user, the prompt information being used to guide the generation of the target video, the prompt information including text and / or images.

21. The apparatus according to any one of claims 14-20, characterized in that, The video generation model is a diffusion model. The diffusion model includes a cross-attention layer; the cross-attention layer includes at least one first variable; When the at least one latent vector is input into a pre-trained video generation model, the generation module is specifically used for: The at least one latent vector is input into the cross-attention layer of the pre-trained video generation model, and the value of the at least one latent vector is assigned to the at least one first variable; And / or, The input layer of the diffusion model includes a noise vector corresponding to random noise; at least one second variable is embedded in the noise vector; When the at least one latent vector is input into a pre-trained video generation model, the generation module is specifically used for: The value of at least one latent vector is assigned to the second variable, and the noise vector embedded with the second variable is input into the pre-trained video generation model.

22. The apparatus as claimed in claim 21, characterized in that, When at least one second variable is embedded in the noise vector, the generation module is specifically used for: The noise vector is connected to the at least one second variable; or, the at least one second variable serves as the noise vector.

23. The apparatus as claimed in any one of claims 14-22, characterized in that, The device further includes a training module, which performs the following steps before receiving at least one control condition: Obtain labeled training data; wherein the label is at least one first hidden vector obtained by at least one control condition encoding; In one iteration of training the video generation model, the training module is further configured to: Extract at least one control condition from the target video output by the video generation model; Obtain at least one second hidden vector corresponding to at least one control condition extracted from the target video; The value of the loss function is calculated based on the at least one second hidden vector and the at least one first hidden vector; Based on the value of the loss function, the trainable parameters in the video generation model are optimized.

24. The apparatus as claimed in claim 23, characterized in that, When extracting at least one control condition from the target video output by the video generation model, the training module is specifically used to perform one or more of the following steps: Call the semantic segmentation model and / or the 3D detection model, and extract the control conditions corresponding to the structure from the target video output by the video generation model through the semantic segmentation model and / or the 3D detection model; The 3D tracking model is invoked, and the control conditions corresponding to the motion are extracted from the target video output by the video generation model through the 3D tracking model. The multimodal understanding model is invoked, and the control conditions corresponding to the environment are extracted from the target video output by the video generation model through the multimodal understanding model. The multimodal understanding model is invoked, and the physical corresponding control conditions are extracted from the target video output by the video generation model through the multimodal understanding model.

25. The apparatus as claimed in claim 23 or 24, characterized in that, When acquiring labeled training data, the training module is specifically used for: Determine the source material and background to be used to generate the sample video; Run the simulation model and import the materials and background into the simulation model; If at least one control condition sent by the user is detected, a sample video conforming to the at least one control condition is generated through the simulation model. The encoder is invoked to encode the at least one control condition, thereby obtaining at least one first hidden vector; The at least one first hidden vector is used as the label of the sample video to obtain labeled training data.

26. The apparatus as claimed in any one of claims 14-25, characterized in that, The device further includes a triggering module and a quality feedback module. The triggering module is configured to perform the following steps after generating a target video that meets at least one control condition: The quality feedback module is invoked to obtain a quality score for the target video. The quality score is obtained based on the evaluation results of subjective and / or objective indicators. The evaluation results of the subjective indicators are obtained based on user operations. The objective indicators include the video quality assessment indicator FVD and / or the multimodal model similarity indicator CLIP Similarity. If the quality score is below a threshold, incremental training of the video generation model is triggered.

27. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1-13.

28. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster performs the method as described in any one of claims 1-13.

29. A computer-readable storage medium, characterized in that, Includes computer program instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1-13.

Citation Information

Patent Citations

  • Video description statement generation method and related equipment

    CN111988673A

  • Image generation method, model training method and corresponding device

    CN117593400A

  • Video synthesis method and device, equipment and storage medium

    CN117750125A

  • Image generation method and device

    CN117893641A