Video generation method and related device
By embedding latent vectors of control conditions into the video generation model, and utilizing a diffusion model and a cross-attention layer for precise control of video attributes, this solves the problem that existing video generation models cannot meet the user's precise control requirements, and achieves fine-grained control and improved controllability of videos.
Patent Information
- Application Number
- CN202411375301.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-06-19
- Filing Date
- 2024-09-29
- Publication Date
- 2025-12-19
AI Technical Summary
Existing video generation models cannot meet users' needs for precise control over videos, especially in autonomous driving scenarios where precise control of video attributes is required. It is difficult to achieve fine-grained control over the structure, motion, and physical properties of videos.
By embedding latent vectors corresponding to control conditions into the video generation model, precise control of video attributes is achieved using a diffusion model and a cross-attention layer. This supports independent or joint control of structural, motion, and physical attributes, and training and optimization are performed in conjunction with a simulation model and a quality feedback module.
It enables precise control over video, meets users' fine-grained needs for video, improves the controllability and adaptability of video generation, and can generate videos that conform to actual application scenarios.
Smart Images

Figure CN121174014A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the field of AI technology, and in particular to a video generation method and related apparatus. BACKGROUND
[0002] With the development of artificial intelligence generated content (AIGC) and other technologies, some video generation models can support text-to-video, image-to-video and video continuation functions, and can generate high realistic videos according to user input text prompts.
[0003] Although the video generation model in the related art can generate high realistic videos, it only supports input of text and / or image prompt information and cannot meet the needs of controllable video generation under more control conditions in actual scenarios. For example, a user wants to generate a large amount of video data as training samples in an automatic driving scenario, which may require accurate regulation and control of various parameters in video attributes. However, the video generation model in the related art only supports generation of videos according to simple prompt information and does not support more accurate control of videos. SUMMARY
[0004] Embodiments of the present application provide a video generation method and an electronic device. Various attribute settings of the video are controlled accurately by setting corresponding control conditions, and by embedding the hidden vectors corresponding to the control conditions in the video generation model, a video meeting the control conditions can be generated to meet the actual needs of users for accurate control of the video.
[0005] In a first aspect, embodiments of the present application provide a video generation method. The method is applied to an electronic device. The method receives at least one control condition, the at least one control condition being used to control at least one attribute of a target video to be generated, the at least one attribute including one or more of the following: a structural attribute, a motion attribute, and a physical attribute. The method encodes the at least one control condition into at least one hidden vector. The method inputs the at least one hidden vector into a pre-trained video generation model, and generates the target video meeting the at least one control condition through the video generation model.
[0006] The at least one attribute of the target video to be generated can be at least one attribute of an image sequence included in the video to be generated. The target video, i.e., the video format data output by the video generation model. The target video to be generated can be understood as the control object of the control condition. The attribute of the video (which can also be referred to as the controllable attribute) can be understood as various controllable objects in the video. By classifying all controllable objects, at least one attribute can be obtained. The control condition is used to control the corresponding attribute, and the control condition and the attribute can be in a one-to-one correspondence. Different control conditions can control different attributes in the video.
[0007] In this way, by using the above method, the user can control at least one of the structural attribute, the motion attribute, and the physical attribute of the video by inputting the control condition. For example, for the structural attribute, the user can specify which materials are included in the video and the distribution position of the materials by inputting the corresponding control condition. For the motion attribute, the user can specify the motion trajectory of the material in the foreground (or the material in the background) by inputting the corresponding control condition. For the physical attribute, the user can control the physical phenomenon expected by the user in the video, such as non-rigid motion, by inputting the corresponding control condition. In this way, the user can control the generated video more accurately by inputting the control condition, thereby meeting the more fine-grained needs of the user and improving the controllability of the target video to be generated. In addition, the user can also achieve more complex control targets by jointly controlling the above different attributes. For example, the user can control the material to move along the set trajectory and then generate a specified physical phenomenon with one of the other materials.
[0008] The at least one control condition can be received by providing a visual interface, receiving the control condition set by the user through the visual interface, such as receiving the text input or the file uploaded by the user. Alternatively, in other embodiments, the visual interface can not be provided, and an API is provided for the user to call. The at least one control condition set by the user is received in the calling request.
[0009] In a possible implementation, the structural attribute includes the material structure and / or the camera trajectory in the target video, and the material structure includes the position or category of at least one material. The motion attribute includes the motion trajectory of at least one material in the target video. The physical attribute includes the physical phenomenon when at least one material moves in the target video.
[0010] The structural attribute, the motion attribute, and the physical attribute can be independent of each other, and the target video can be controlled independently. Alternatively, two or more of the three attributes can be combined to control the target video.
[0011] In a possible implementation, the target video includes a first material and a second material, the at least one attribute includes a structure attribute and a motion attribute, the structure attribute includes a material structure of the first material and the second material, and the motion attribute includes a motion trajectory of the first material, and the target video includes a video in which the first material interacts with the second material along the motion trajectory.
[0012] In this implementation, the motion attribute and the structure attribute can be controlled in combination, for example, the material structure of the first material and the second material can be controlled, and the motion trajectory of the first material can be controlled, so that the first material interacts with the second material in the video. The interaction can be collision, or other forms of interaction such as friction and scratching, wherein the collision can cause the motion trajectory to change and / or cause deformation.
[0013] In a possible implementation, the at least one attribute further includes a physical attribute, and the target video includes a video in which the first material interacts with the second material in a physical phenomenon along the motion trajectory.
[0014] In this implementation, the motion attribute, the structure attribute, and the physical attribute can be jointly controlled, so that the first material and the second material in the video interact in a manner consistent with the specified physical phenomenon. The interaction in a physical phenomenon can be an interaction that conforms to the characteristics or rules of a specified physical phenomenon, for example, the physical phenomenon can be non-rigid body motion, and the interaction can be collision. Collision that conforms to the characteristics of non-rigid body motion can be understood as collision that causes deformation during the collision process. Collision that does not cause deformation is not collision that conforms to the characteristics of non-rigid body motion.
[0015] In a possible implementation, the at least one attribute further includes an environment attribute, and the environment attribute includes environment parameter information of the target video, and the environment parameter includes at least one of weather and illumination.
[0016] In this implementation, the attributes in the video content can be divided into four attributes, namely, structure, motion, environment, and physical, and the four attributes can be controlled respectively. In other embodiments, the video attributes can be divided in other ways, for example, the video attributes can be divided into two controllable attributes, namely, intra-frame attribute and inter-frame attribute. The intra-frame attribute includes various controllable objects in an image, such as material structure. The inter-frame attribute includes relative position relationships between multiple images, such as position change trajectories of the same material in multiple images. Alternatively, at least one controllable attribute can be obtained by using other classification methods, for example, the controllable attribute can be divided into static attribute or dynamic attribute, and the like. The attribute division method provided in this implementation enables the video to be controlled accurately, and is compatible with 3D simulation tools and the like in related technologies, which is beneficial to the generation of training data.
[0017] In a possible implementation, determining the at least one control condition comprises: displaying a first control, and receiving, through the first control, a control condition sent by the user; the control condition indicates control over the structure attribute, and the control condition comprises one or more of a text and a predetermined format file; or the control condition indicates control over the motion attribute, and the control condition comprises one or more of a text and a predetermined format file; or the control condition indicates control over the physical attribute, and the control condition comprises one or more of a text and a predetermined format file.
[0018] In this implementation, the control condition can be a text instruction input by the user and / or a predetermined format file uploaded by the user, and the predetermined format file can be an image, a table, a document, or other supported file formats.
[0019] In a possible implementation, a second control can also be displayed. The second control is configured to receive prompt information input by the user, the prompt information is used to guide the generation of the target video, and the prompt information comprises a text and / or an image.
[0020] This implementation can be compatible with existing text and / or image-based control methods, and is equivalent to adding an attribute-based control method on the basis of existing text and / or image prompt information for controlling video generation.
[0021] In a possible implementation, the video generation model is a diffusion model Diffusion; the diffusion model comprises a cross-attention layer; the cross-attention layer comprises at least one first variable; and inputting the at least one latent vector into the pre-trained video generation model comprises: inputting the at least one latent vector into the cross-attention layer in the pre-trained video generation model, and assigning a value of the at least one latent vector to the at least one first variable; and / or, the input layer of the diffusion model comprises a noise vector corresponding to random noise; the noise vector is embedded with at least one second variable; and inputting the at least one latent vector into the pre-trained video generation model comprises: assigning a value of the at least one latent vector to the second variable, and inputting the noise vector embedded with the second variable into the pre-trained video generation model.
[0022] In this implementation, the video generation model can be Diffusion. In other embodiments, the video generation model can be other autoregressive video generation models containing cross-attention layers other than Diffusion. The cross-attention layer includes a first variable, and the value of the first variable is equal to the value of the hidden vector corresponding to the control condition. It can be understood that the hidden vector corresponding to the control condition is embedded in the cross-attention layer. In some embodiments, at least one hidden vector can be embedded only in the cross-attention layer. In other embodiments, at least one hidden vector can be embedded only through the noise vector corresponding to the random noise. In yet other embodiments, at least one hidden vector can be embedded through the cross-attention layer, and at least one hidden vector can be embedded through the noise vector corresponding to the random noise.
[0023] By embedding the hidden vector in the video generation model, the influence of the control condition on the video can be fully learned by the learning mechanism of the diffusion model itself, so that the trained video generation model can generate videos consistent with the control condition during runtime.
[0024] In a possible implementation, embedding at least one second variable in the noise vector can be connecting the noise vector with the at least one second variable; or the at least one second variable can be the noise vector.
[0025] In this implementation, embedding the second variable in the noise vector can be understood as embedding at least one hidden vector in the noise vector. The specific embedding can be using the concat function to connect the noise vector and the at least one hidden vector. Alternatively, in other embodiments, embedding the at least one hidden vector in the noise vector can be feature fusion between the at least one hidden vector and the noise vector, such as weighted summation or dot product operation between the at least one hidden vector and the noise vector, or other operations that can fuse the features of the at least one hidden vector and the noise vector. Alternatively, the embedding method can also be using the at least one hidden vector to replace the original random noise vector, i.e., using the at least one hidden vector as the noise vector to participate in the subsequent calculation of the Diffusion model.
[0026] The noise vector of the random noise is part of the input data of the Diffusion model. By embedding the hidden vector corresponding to the control condition in the noise vector, the influence of the control condition on the video can be learned by the diffusion mechanism of the diffusion model itself, and a video consistent with the control condition can be generated.
[0027] In a possible implementation, an encoder can be called to encode the at least one control condition to obtain at least one hidden vector corresponding to the at least one control condition.
[0028] The control conditions are encoded to obtain a latent vector that can be input to the video generation model and is compatible with the video generation model.
[0029] In a possible implementation, before the first control is displayed, the method further includes obtaining training data with labels, wherein the labels are at least one first latent vector corresponding to at least one control condition; in one iteration of training the video generation model, the method further includes: extracting the at least one control condition from a target video output by the video generation model; obtaining at least one second latent vector corresponding to the at least one control condition extracted from the target video; calculating a value of a loss function according to the at least one second latent vector and the at least one first latent vector, and optimizing trainable parameters in the video generation model according to the value of the loss function.
[0030] In the training phase, through effective training of the video generation model, the video generation model with the ability to generate videos meeting the control conditions can be obtained.
[0031] In a possible implementation, the at least one control condition is extracted from the target video output by the video generation model, including one or more of the following steps: calling a semantic segmentation model and / or a 3D detection model, and extracting a structure corresponding control condition from the target video output by the video generation model through the semantic segmentation model and / or the 3D detection model; calling a 3D tracking model, and extracting a motion corresponding control condition from the target video output by the video generation model through the 3D tracking model; calling a multi-modal understanding model, and extracting an environment corresponding control condition from the target video output by the video generation model through the multi-modal understanding model; calling the multi-modal understanding model, and extracting a physics corresponding control condition from the target video output by the video generation model through the multi-modal understanding model.
[0032] In the training phase, it is necessary to evaluate whether the video output by the video generation model in the training meets the preset control condition, and the prerequisite for the evaluation is to extract the control condition from the video output by the video generation model. In the extraction, the extraction of the control condition can be implemented by means of the above models.
[0033] In a possible implementation, the training data with labels is obtained, including: determining materials and backgrounds used for generating sample videos; running a simulation model, and importing the materials and backgrounds into the simulation model; detecting at least one control condition set by a user, and generating a sample video meeting the at least one control condition through the simulation model; calling an encoder, and encoding the at least one control condition through the encoder to obtain at least one first latent vector; and taking the at least one first latent vector as a label of the sample video to obtain the training data with labels.
[0034] The simulation model can be a 3D simulation engine, so that the training data can be obtained through 3D simulation, solving the problem of the source of the training data.
[0035] In a possible implementation, after the target video meeting the at least one control condition is generated, a quality feedback module can be invoked to obtain a quality score of the target video through the quality feedback module; the quality score is obtained based on evaluation results of subjective indexes and / or objective indexes; the evaluation results of the subjective indexes are obtained according to user operations; the objective indexes include a video quality evaluation index FVD and / or a multi-modal model similarity index CLIP Similarity; in a case where the quality score is lower than a threshold, incremental training of the video generation model is triggered.
[0036] The incremental training can continue to improve the accuracy of the video generation model in actual application, ensure that the prediction performance of the video generation model does not decrease, or in other words, according to actual data in actual application as a sample for incremental training, the video generation model can better fit the specific characteristics of the actual application scene, output a video that better meets the current application environment, and improve the ability to adapt to different environments.
[0037] It should be noted that in the incremental training process, the video generation model needs to be iteratively trained until the quality score meets the threshold requirement, for example, is equal to or greater than the threshold.
[0038] In a second aspect, the embodiments of the present application further provide a video generation device, which can include: a first receiving module configured to receive at least one control condition, the at least one control condition being used to control at least one attribute of a target video to be generated, the at least one attribute including one or more of the following: a structural attribute, a motion attribute, and a physical attribute; an encoding module configured to encode the at least one control condition into at least one latent vector; and a generation module configured to input the at least one latent vector into a pre-trained video generation model, and generate a target video meeting the at least one control condition through the video generation model.
[0039] In a possible implementation, the target video includes at least one material, the structural attribute includes a material structure and / or a camera track in the target video, the material structure includes a position or a category of the at least one material; the motion attribute includes a motion track of the at least one material in the target video; and the physical attribute includes a physical phenomenon when the at least one material moves in the target video.
[0040] In a possible implementation, the target video includes a first material and a second material, the at least one attribute includes a structure attribute and a motion attribute, the structure attribute includes a material structure of the first material and the second material, and the motion attribute includes a motion track of the first material, and the target video includes a video in which the first material interacts with the second material along the motion track.
[0041] In a possible implementation, the at least one attribute further includes a physical attribute, and the target video includes a video in which the first material interacts with the second material in a physical phenomenon along the motion track.
[0042] In a possible implementation, the at least one attribute further includes an environment attribute, and the environment attribute includes environment parameter information of the target video, and the environment parameter includes at least one of weather and illumination.
[0043] In a possible implementation, when the at least one control condition is received, the first receiving module is specifically configured to display a first control, and receive a control condition sent by a user through the first control, the control condition indicates control on the structure attribute, the control condition includes one or more of a text and a file in a predetermined format, or the control condition indicates control on the motion attribute, the control condition includes one or more of a text and a file in a predetermined format, or the control condition indicates control on the physical attribute, and the control condition includes one or more of a text and a file in a predetermined format.
[0044] In a possible implementation, the apparatus further includes a second receiving module, and the second receiving module is configured to display a second control, and the second control is configured to receive prompt information input by a user, the prompt information is used to guide generation of the target video, and the prompt information includes a text and / or an image.
[0045] In a possible implementation, the video generation model is a diffusion model Diffusion, the diffusion model includes a cross-attention layer, the cross-attention layer includes at least one first variable, and when the at least one latent vector is input into the pre-trained video generation model, the generation module is specifically configured to input the at least one latent vector into the cross-attention layer in the pre-trained video generation model, and assign a value of the at least one latent vector to the at least one first variable, and / or, an input layer of the diffusion model includes a noise vector corresponding to random noise, the noise vector is embedded with at least one second variable, and when the at least one latent vector is input into the pre-trained video generation model, the generation module is specifically configured to assign a value of the at least one latent vector to the second variable, and input the noise vector embedded with the second variable into the pre-trained video generation model.
[0046] In a possible implementation, when the at least one second variable is embedded in the noise vector, the generating module is specifically configured to: connect the noise vector with the at least one second variable; or, take the at least one second variable as the noise vector.
[0047] In a possible implementation, the apparatus further includes a training module configured to, before receiving the at least one control condition, perform the following steps: obtain training data with labels; wherein the labels are at least one first hidden vector coded by the at least one control condition; in one iteration of training the video generation model, the training module is further configured to: extract the at least one control condition from a target video output by the video generation model; obtain at least one second hidden vector corresponding to the at least one control condition extracted from the target video; calculate a value of a loss function according to the at least one second hidden vector and the at least one first hidden vector; and optimize trainable parameters in the video generation model according to the value of the loss function.
[0048] In a possible implementation, when the at least one control condition is extracted from the target video output by the video generation model, the training module is specifically configured to perform one or more of the following steps: call a semantic segmentation model and / or a 3D detection model, and extract a structure corresponding control condition from the target video output by the video generation model through the semantic segmentation model and / or the 3D detection model; call a 3D tracking model, and extract a motion corresponding control condition from the target video output by the video generation model through the 3D tracking model; call a multi-modal understanding model, and extract an environment corresponding control condition from the target video output by the video generation model through the multi-modal understanding model; call the multi-modal understanding model, and extract a physics corresponding control condition from the target video output by the video generation model through the multi-modal understanding model.
[0049] In a possible implementation, when the training data with labels is obtained, the training module is specifically configured to: determine materials and backgrounds used for generating sample videos; run a simulation model, and import the materials and the backgrounds into the simulation model; detect at least one control condition sent by a user, and generate a sample video conforming to the at least one control condition through the simulation model; call an encoder, and encode the at least one control condition through the encoder to obtain at least one first hidden vector; and take the at least one first hidden vector as a label of the sample video to obtain the training data with labels.
[0050] In a possible implementation, the apparatus further includes a triggering module, configured to, after generating the target video meeting the at least one control condition, perform the following steps: calling the quality feedback module, and obtaining a quality score of the target video through the quality feedback module; the quality score is obtained based on evaluation results of subjective indexes and / or objective indexes; the evaluation results of the subjective indexes are obtained according to user operations; the objective indexes include a video quality evaluation index FVD and / or a multi-modal model similarity index CLIP Similarity; and triggering incremental training of the video generation model in a case where the quality score is lower than a threshold.
[0051] In a third aspect, the embodiments of the present application further provide a computing device cluster, including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the method according to any one of the first aspect.
[0052] In a fourth aspect, the embodiments of the present application further provide a computer program product including instructions, which, when executed by a computing device cluster, cause the computing device cluster to perform the method according to any one of the first aspect.
[0053] In a fifth aspect, the embodiments of the present application further provide a computer-readable storage medium including computer program instructions, which, when executed by a computing device cluster, cause the computing device cluster to perform the method according to any one of the first aspect.
[0054] In other aspects, the embodiments of the present application further provide an electronic device, including: a processor configured to execute computer programs or instructions in a memory to implement the method according to any one of the above.
[0055] In other aspects, the embodiments of the present application further provide a chip system, including: a communication interface configured to input and / or output data; and a processor configured to execute computer executable programs, so that a device installed with the chip system performs the method according to any one of the above. BRIEF DESCRIPTION OF DRAWINGS
[0056] Figure 1 An example of a process for generating a video based on a traditional UE simulation engine in the related art;
[0057] Figure 2 An example of generating a video for a general video generation model in the related art;
[0058] Figure 3 A structural schematic diagram of an electronic device (terminal device, such as a smart phone);
[0059] Figure 4 A structural schematic diagram of an electronic device (server);
[0060] Figure 5 A schematic diagram of an application scenario of a video generation method provided by an embodiment of the present application;
[0061] Figure 6 A system architecture example diagram of a video generation method provided by an embodiment of the present application;
[0062] Figure 7 A schematic diagram of foreground structure and background structure in some embodiments of a video generation method provided by an embodiment of the present application;
[0063] Figure 8 A schematic diagram of controlling motion trajectory in some embodiments of a video generation method provided by an embodiment of the present application;
[0064] Figure 9 A schematic diagram of controlling environment in some embodiments of a video generation method provided by an embodiment of the present application;
[0065] Figure 10 A software module architecture example diagram in some embodiments of a video generation method provided by an embodiment of the present application;
[0066] Figure 11 A flow processing example diagram in some embodiments of a video generation method provided by an embodiment of the present application, which is divided into a training stage, a use stage and an incremental training stage;
[0067] Figure 12 An example diagram of cross attention layer embedding hidden vector in some embodiments of a video generation method provided by an embodiment of the present application;
[0068] Figure 13 Another example diagram of cross attention layer embedding hidden vector in some embodiments of a video generation method provided by an embodiment of the present application;
[0069] Figure 14 An example diagram of noise vector embedding hidden vector in some embodiments of a video generation method provided by an embodiment of the present application;
[0070] Figure 15 Another example diagram of noise vector embedding hidden vector in some embodiments of a video generation method provided by an embodiment of the present application;
[0071] Figure 16a An interface example diagram in some embodiments of a video generation method provided by an embodiment of the present application;
[0072] Figure 16b Interface example diagrams in some other embodiments of the video generation method provided in this application;
[0073] Figure 17 A schematic diagram illustrating the embedding of latent vectors in the video generation model at runtime in some embodiments of the video generation method provided in this application;
[0074] Figure 18 A flowchart illustrating the video generation method provided in the embodiments of this application from the perspective of a terminal device;
[0075] Figure 19 A flowchart illustrating the video generation method proposed in the embodiments of this application from the server's perspective;
[0076] Figure 20 A schematic diagram of a computing device provided in an embodiment of this application;
[0077] Figure 21 A schematic diagram of a computing device cluster provided in an embodiment of this application;
[0078] Figure 22 This is a schematic diagram illustrating how one or more computing devices in a computing device cluster provided in an embodiment of this application can be connected via a network. Detailed Implementation
[0079] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of, and not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0080] In the description of the embodiments of this application, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The word "and / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more.
[0081] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0082] In the embodiments of the present application, the word "exemplary" or "for example" is used to mean "an example of" or "an example, only. Any embodiment or design described herein as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs.
[0083] The traditional video generation method is to obtain a simulation video by using a simulation tool such as a virtual engine (Unreal Engine, UE), Carla, Gazebo, and the like, which requires a large amount of manpower for scene production. For example, as shown in Figure 1 Based on a simulation engine such as UE, visual simulation can generally use the following process:
[0084] Material modeling: manually modeling or obtaining a plurality of material models through 3D scanning;
[0085] Static scene editing: a static scene is built by editing the positions of a plurality of materials, adding lighting, weather, and the like, and the like;
[0086] Dynamic scene editing: the scene is made dynamic by editing the motion of the materials, the motion of the camera, and the like;
[0087] Rendering: rendering the edited dynamic scene into a video.
[0088] Such a visual simulation scheme based on a simulation engine such as UE requires a large amount of manual work for each step of scene building, and the final rendering result still has a domain gap (a gap between a simulation system and a real world) from a real scene.
[0089] At present, with the development of artificial intelligence (AI) technology, especially the rapid development of AIGC technology, some AI-based video generation tools have been able to support text-to-video, image-to-video, and video style conversion, and the like, and generate high-fidelity videos. For example, in an animation video production scenario, a user can quickly produce an animation video with a complete plot by inputting a prompt word and an animation style; in a creative video generation scenario, a user can upload an image and input a prompt word describing the image to generate a video with rich facial expressions and various motion gestures; and in a virtual character production scenario, a still image and audio can be uploaded to generate a lifelike dynamic video.
[0090] As shown in Figure 2 It is a text-to-video generation model that can generate high-quality video content according to a descriptive text prompt. The video generation model can control video generation through text and image, or perform operations such as video continuation and style conversion, to directly produce a high-fidelity video.
[0091] Such as Figure 2 As shown in the general video generation model, when a user wants to generate a video, the user generally needs to input a text prompt or upload an image, and the video generation model will generate a corresponding video according to the text and / or image provided by the user. Although a high realistic video can be directly generated, it can only be controlled by text and / or image, and there is a certain limitation for the control of the video.
[0092] In actual application, the user may have more fine-grained needs for the controllability of the video, for example, the user may need to accurately control various attributes of the video such as 3D structure and motion, wherein the 3D structure can include materials in the foreground and material position information, and the motion can refer to the motion trajectory of the material in the foreground. Some of these attributes are difficult to describe in natural language, for example, the position of the material can be a set of coordinate data, which cannot be described by a simple text prompt; or most of the current video generation models have certain requirements for the length of the input text prompt, and a too complex text prompt will cause the video generation model to be unable to recognize, and thus unable to generate a corresponding video.
[0093] Due to the limitations of the text and image control methods, the generated video has a certain randomness, which may deviate from the actual needs of the user, resulting in poor video usability and failing to meet the user's needs for more fine-grained or more accurate control of video generation.
[0094] For example, the user inputs the text: "a dog is flying in the sky", and the video generation model can generate a high realistic video of a dog flying in the sky, but if the user wants to accurately control the dog to fly according to a specified irregular trajectory, or control the intensity of the light to change according to a preset value during the flying process of the dog, or wants to accurately control the distribution position of the clouds in the sky, or wants to know whether a certain specified physical phenomenon occurs at a specified position during the flying process of the dog, etc., some current video generation models may not be able to achieve this.
[0095] Or, in actual application, the user needs to accurately control the parameters of the video attributes, for example, the user needs to accurately control the specific coordinate position of the material, the coordinate of the motion trajectory of the material, and the light parameters, and most of the current video generation models do not support parameter-level control. In the automatic driving scene, a large number of training videos can be obtained by accurately controlling various parameters.
[0096] In view of this, the embodiments of the present application propose a video generation method and an electronic device, which split a video required by a user into multiple attributes for separate control, support the user to input corresponding control conditions for the multiple attributes respectively, and generate a video meeting specific control conditions, so as to more accurately match the user demand, realize more accurate control (parameter level) of the video, and improve the controllability of the video.
[0097] For example, the method proposed in the embodiments of the present application can control the position of the dog in each frame of image in the video, so as to obtain an accurate and controllable flight trajectory.
[0098] The electronic device proposed in the embodiments of the present application can be a terminal device on the user side, for example, can be a smart phone, a Tablet PC, a laptop, a Desktop computer, a wearable device, an AR / VR device, an UMPC, a netbook, or a PDA, etc. The embodiments of the present application do not make any limitation on the specific type of the electronic device.
[0099] For example, Figure 3 The structural schematic diagram of the electronic device provided for an embodiment of the present application, for example, the electronic device can be a smart phone, such as Figure 3As shown, the electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headset jack 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0100] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than shown, or combine certain components, or split certain components, or different arrangement of components. The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0101] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices, or can be integrated in one or more processors.
[0102] The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions.
[0103] The processor 110 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can hold instructions or data that the processor 110 has just used or cycled through. If the processor 110 needs to use the instructions or data again, it can be called directly from the memory. This avoids repeated access and reduces the latency of the processor 110, thus improving the efficiency of the system.
[0104] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0105] The USB interface 130 is an interface that conforms to the USB standard specification, and can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device 100, and can also be used to transmit data between the electronic device 100 and a peripheral device. It can also be used to connect earphones to play audio through the earphones. The interface can also be used to connect other electronic devices, such as AR devices, etc.
[0106] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a structural limitation on the electronic device 100. In some other embodiments of the present application, the electronic device 100 can also use different interface connection methods or combinations of multiple interface connection methods in the above embodiments.
[0107] The charging management module 140 is configured to receive charging input from a charger. The charger can be a wireless charger or a wired charger. In some embodiments with wired charging, the charging management module 140 can receive charging input from a wired charger through the USB interface 130. In some embodiments with wireless charging, the charging management module 140 can receive wireless charging input through a wireless charging coil of the electronic device 100. The charging management module 140 can charge the battery 142 and power the electronic device 100 through the power management module 141.
[0108] The power management module 141 is configured to connect the battery 142 and the charging management module 140 to the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, the internal memory 121, the display 194, the camera 193, and the wireless communication module 160, etc. The power management module 141 can also be configured to monitor parameters such as battery capacity, battery cycle count, battery health (leakage, impedance), etc. In some other embodiments, the power management module 141 can also be disposed in the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 can also be disposed in the same device.
[0109] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, etc.
[0110] The antenna 1 and the antenna 2 are configured to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be configured to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in combination with a tuning switch.
[0111] The mobile communication module 150 can provide a solution for wireless communication including 2G / 3G / 4G / 5G, etc. applied to the electronic device 100. The mobile communication module 150 can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves by the antenna 1, and perform filtering, amplification, etc. on the received electromagnetic waves, and transfer the same to the modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor, and radiate the same as electromagnetic waves through the antenna 1. In some embodiments, at least part of the functional modules of the mobile communication module 150 can be disposed in the processor 110. In some embodiments, at least part of the functional modules of the mobile communication module 150 can be disposed in the same device as at least part of the modules of the processor 110.
[0112] The modem processor can include a modulator and a demodulator. The modulator is configured to modulate a low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is configured to demodulate a received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. The low-frequency baseband signal processed by the baseband processor is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the microphone 170B, etc.), or displays an image or a video through the display screen 194. In some embodiments, the modem processor can be a separate device. In other embodiments, the modem processor can be independent of the processor 110, and disposed in the same device as the mobile communication module 150 or other functional modules.
[0113] The wireless communication module 160 can provide a solution for wireless communication including wireless local area networks (WLAN) (e.g., wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc. applied to the electronic device 100. The wireless communication module 160 can be one or more devices that integrate at least one communication processing module. The wireless communication module 160 receives an electromagnetic wave via the antenna 2, frequency-modulates and filters the electromagnetic wave signal, and transmits the processed signal to the processor 110. The wireless communication module 160 can also receive a signal to be transmitted from the processor 110, frequency-modulate it, amplify it, and radiate it as an electromagnetic wave via the antenna 2.
[0114] In some embodiments, the antenna 1 and the mobile communication module 150 of the electronic device 100 are coupled, and the antenna 2 and the wireless communication module 160 are coupled, so that the electronic device 100 can communicate with a network and other devices through wireless communication technology. The wireless communication technology can include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-CDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS can include global positioning system (GPS), global navigation satellite system (GLONASS), beidu navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS). That is, the electronic device 100 has a positioning function and a wireless communication function.
[0115] The electronic device 100 implements a display function through a GPU, a display 194, and an application processor, etc. The GPU is a microprocessor for image processing, which is connected to the display 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs, which execute program instructions to generate or change display information.
[0116] The display screen 194 is configured to display images, videos, and the like. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diodes (QLED), or the like. In some embodiments, the electronic device 100 can include one or N display screens 194, where N is a positive integer greater than 1.
[0117] The electronic device 100 can implement the photographing function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor.
[0118] The ISP is configured to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing to convert it into an image visible to the naked eye. The ISP can also optimize the noise, brightness, and skin color of the image. The ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene. In some embodiments, the ISP can be disposed in the camera 193.
[0119] The camera 193 is configured to capture still images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV, or the like format. In some embodiments, the electronic device 100 can include one or N cameras 193, where N is a positive integer greater than 1.
[0120] The digital signal processor is used to process digital signals, in addition to being able to process digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0121] The video codec is used to compress or decompress digital video. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple encoding formats, such as: moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.
[0122] The NPU is a neural-network (NN) calculation processor, which can quickly process input information by drawing on the structure of a biological neural network, such as drawing on the transmission mode between human brain neurons, and can also constantly self-learn. Through the NPU, the electronic device 100 can realize intelligent cognition applications such as image recognition, face recognition, voice recognition, text understanding, etc.
[0123] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to realize data storage functions. For example, music, video, etc. Files are saved in the external memory card.
[0124] The internal memory 121 can be used to store computer executable program codes, which include instructions. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), etc. The data storage area can store data created during the use of the electronic device 100 (such as audio data, a phonebook, etc.), etc. In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various function applications and data processing of the electronic device 100 by running instructions stored in the internal memory 121 and / or instructions stored in the memory disposed in the processor.
[0125] The electronic device 100 can realize audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the earphone interface 170D, and the application processor, etc. For example, music playing, recording, etc.
[0126] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.
[0127] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can correspond to touch operations performed on different applications (such as taking photos, playing audio, etc.). Motor 191 can also correspond to different vibration feedback effects for touch operations performed on different areas of the display screen 194. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.
[0128] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.
[0129] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and separate from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to realize functions such as calls and data communication. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.
[0130] The electronic device proposed in the embodiments of this application can also be a server on the service side. For example, Figure 4 This is a schematic diagram of the server structure in one embodiment of this application. Figure 4 As shown, server 200 may include: one or more processors 210, communication interface 220, memory 230, and communication bus 240 connecting different components (including memory 230, communication interface 220 and processor 210).
[0131] The communications bus 240 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration bus, or a local bus using any of a variety of bus architectures. By way of example, and not limitation, the communications bus 240 can include an industry standard architecture (ISA) bus, a micro channel architecture (MCA) bus, an enhanced ISA bus, a video electronics standards association (VESA) local bus, and a peripheral component interconnect (PCI) bus.
[0132] The electronic device typically includes a variety of computer system readable media. These media can be any available media that is accessible by the electronic device and includes both volatile and non-volatile media, removable and non-removable media.
[0133] The memory 230 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The memory 230 can include, without limitation, at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions and / or techniques of embodiments of the application, including the video generation method.
[0134] The program / utility, having a set (at least one) of program modules, can be stored in the memory 230 by way of example, and not limitation, an operating system, one or more application programs, other program modules, and program data, each or some combination thereof, can include implementation of a network environment. The program modules are generally carried in the memory 230 and executed by the processor 210 to implement the functions / techniques described herein.
[0135] The processor 210 executes the program modules stored in the memory 230 to carry out various functions and / or methods of embodiments of the application, such as the video generation method.
[0136] It should be appreciated that Figure 4 The processor 210 in the server 200 shown can be a system on a chip (SOC), which can include a central processing unit (CPU), and can further include other types of processors, such as a graphics processing unit (GPU), etc.
[0137] As Figure 5 indicated, the video generation method proposed in the embodiments of the present application can be applied in various application scenarios such as automatic driving, embodied intelligence, city twinning, video editing, etc., wherein the video editing may, for example, be animation video production, creative video generation, virtual character production, etc.
[0138] Exemplarily, in the field of automatic driving, the method proposed in the embodiments of the present application supports user input of weather, illumination, driving scene, background semantic map, foreground motion trajectory, camera parameter, etc. control conditions, and generates a special driving scene conforming to the control conditions, such as generating a video of a vehicle rapidly cutting in, a "ghost probe" phenomenon, etc. The generated video can be used as a training sample to improve the performance of the automatic driving system by simulating different driving scenes.
[0139] For example, as Figure 5 indicated, the user inputs a text prompt "pedestrian walking in a non-zebra crossing area", and inputs control conditions corresponding to various attributes such as structure, motion, environment, physics, etc. respectively, to generate a corresponding video, which can be used as a sample video to optimize the recognition and response performance of the automatic driving system for the scene of a pedestrian walking in a non-zebra crossing area. Among them, the control conditions such as driving scene, background semantic map, camera parameter (camera trajectory), etc. can be input in the control of the structure, information such as foreground motion trajectory can be input in the control of the motion, and information such as weather, illumination, etc. can be input in the control of the environment.
[0140] Exemplarily, the video generation method provided in the embodiments of the present application can be implemented based on the system architecture as Figure 6 indicated.
[0141] As Figure 6 indicated, the video generation model can be deployed in a terminal device or in a cloud server. According to the deployment location of the video generation model, the method proposed in the embodiments of the present application can be implemented based on a system architecture in which the video generation model is deployed in the cloud, or a system architecture in which the video generation model is deployed in the terminal device.
[0142] Exemplarily, a system architecture in which the video generation model is deployed in the cloud can include a server 301 and one or more terminal devices 302 in remote communication with the server, for example, the terminal devices can be smartphones, desktop computers, tablet computers, etc. In this system architecture example, the server 301 can be a standalone server or a server cluster, for example, a highly available (HA) cluster, a load balancing (LB) cluster, or a high-performance computing (HPC) cluster. The video generation model is deployed in the server 301, and the terminal device can display controls for user input of control conditions and prompts (text and / or images). After detecting the user input of the control conditions and prompts, the terminal device can send a call request to the server 301, the call request carrying the control conditions and the prompts, etc. The server 301 responds to the call request of the terminal device, inputs the control conditions, prompts, etc. uploaded by the terminal device into the video generation model, runs the video generation model, generates and outputs a video, and distributes the generated video to the terminal device. The terminal device can display or play the video.
[0143] In a system architecture example in which the video generation model is deployed in the terminal device, the terminal device 304 has a lightweight video generation model (e.g., Diffusion) deployed therein, and other models are deployed in the server, for example, a multi-modal large model such as Video LLava. The terminal device 304 can locally call the video generation model to generate a video. It should be noted that when incrementally training the video generation model, other models may need to be called to participate in the training process. Due to the limitations of the computing power and storage resources of the terminal device 304 such as a smartphone, the multi-modal large model and other models are generally deployed in the cloud server 303, and the terminal device 304 needs to remotely call.
[0144] The video generation method proposed in the embodiments of the present application adopts a video generation mechanism for separately controlling at least one attribute of a video, supports a user to set at least one control condition, and each control condition is used to control a corresponding attribute in the video. The attribute of the video refers to an object that can be adjusted and controlled in the video generation or editing process. For example, the at least one attribute can include at least one of the following attributes: structure, motion, environment, and physics. The user can input a corresponding control condition for each attribute to control a corresponding attribute, for example, inputting weather and lighting (which can be text) as a control condition to control the environment attribute, or inputting a background semantic graph (which can be in image format) to control the structure attribute, etc.
[0145] Specifically, structural attributes may include material structure and / or camera trajectory.
[0146] Camera trajectory refers to the changing position of the camera across each frame of a video. Camera position, in this context, is the position of the camera corresponding to each frame when the image is captured by the camera. Multiple camera positions across multiple frames constitute the camera trajectory. Footage structure defines the information within a single frame. Camera trajectory defines the relative positional relationships (camera viewpoint) between multiple frames.
[0147] The material structure refers to the various information used to describe the material in a frame of an image. For example, the material structure may include material location and / or material category. Material location sets the position information of each material in a frame of the video, such as the boundary coordinates of each material. Material category refers to the category to which each material belongs, such as people, vehicles, animals, etc. Each category can be further divided into multiple subcategories; for example, people can be various types like beautiful women, children, and businessmen, and vehicles can be divided into large trucks, cars, tankers, etc. Alternatively, in some embodiments, the material category may specifically be information such as material name or material ID, allowing the corresponding material to be found in a material library based on the material name or material ID. The material category may also include other material-related information, such as the storage address of the material, or additional information related to the material, such as information about other items used to decorate the material.
[0148] For example, the material structure can be divided into a foreground structure and / or a background structure. The foreground structure can include a foreground location and foreground material categories, used to represent the composition of the materials in the foreground and the distribution of each material in the image. For example, foreground information can indicate which materials are included in the foreground and the boundary coordinates of each material. The background structure is used to represent the composition of the background materials and the distribution of each material in the image. For example, background information can indicate which materials are included in the background and the boundary coordinates of each material.
[0149] For example, such as Figure 7 As shown, Figure 7 The image shows a frame depicting a dog wearing sunglasses flying in the sky. In the foreground, the subject is a dog, and its location could be the dog's boundary coordinates or other location data. In the background, the subjects are blue sky and white clouds, and their locations could be the boundary coordinates of the blue sky and white clouds, or other location data.
[0150] Exemplarily, when the structure property is controlled, a user input image format and / or text format control condition is supported. The image format can be a single image or multiple images (image sequence).
[0151] For example, in the present embodiment, for the control of the foreground structure, a three-dimensional target detection technology can be used to generate image data containing a two-dimensional bounding box (2D Bounding Box, 2D Bbox) or a three-dimensional bounding box (3D Bounding Box, 3D Bbox) as a control condition input or upload for controlling the foreground structure. The 2D Bbox or 3D Bbox can represent the position of the material in the foreground structure. For example, as shown in FIG. 7, the specific position of the dog in the foreground structure can be indicated by an image with a 3D Bbox.
[0152] In the present embodiment, the background structure can be represented by a background semantic map, which can contain the category and coordinate information of each material, for example, the material category can be a vehicle or a human, etc. Exemplarily, as shown in FIG. 8, when a video of a dog flying in the sky needs to be generated, a background semantic map containing the material categories of blue sky and white clouds and the coordinate information of the blue sky and white clouds can be generated first. Then, the background semantic map is uploaded as a control condition for the background structure in the structure property. Figure 7
[0153] Exemplarily, the background semantic map can be generated by using a semantic segmentation model or other semantic map model, for example, a semantic segmentation model is used to extract the semantic map corresponding to a sample image or a reference image, and the semantic map is uploaded as a control condition for the background structure.
[0154] In other embodiments, a bird's eye view (Bird's Eye View map, BEV map) semantic map can be used to represent the background structure, for example, when a drone aerial video needs to be generated, a BEV format background semantic map can be used.
[0155] It should be noted that in other embodiments, the foreground structure can also be represented by a foreground semantic map.
[0156] In addition, the method proposed in the embodiments of the present application is applicable to both 3D video generation and 2D video generation. For the sake of brevity of description, the following description will be given mainly with respect to 3D, which should not be understood as being applicable only to 3D.
[0157] Motion attributes are used to set the motion trajectory of at least one target object in the foreground of the target video. The target object can be one or more foreground elements. Motion attributes, also known as motion trajectories, represent the positional change trajectory of foreground elements within a video segment. The position of the foreground element in one frame can be obtained as a set of coordinate data; the coordinate data corresponding to multiple frames constitute a trajectory. The method proposed in this application supports setting the position of foreground elements in each frame, or supports setting the position of the target object in the foreground every N frames. The value of N can be controlled according to the video playback frame rate. For example, if the frame rate is 24 frames / s, N can be set every 3, 5, or 11 frames.
[0158] For example, such as Figure 8 As shown, the foreground material is a young woman walking. Assuming the generated video requires controlling the motion trajectory of this material from... Figure 8 The movement of point P1 to point P5, along the trajectory formed by points P2, P3, and P4, generates multi-frame images containing 3D bounding boxes. The 3D bounding box in each frame indicates the specific location of the material. For example, the position information of the foreground material is set every N frames, generating the 3D bounding boxes corresponding to frames 1, 1+N, 1+2N, 1+3N, and 1+4N. Based on the boundary coordinates of the 3D bounding boxes in each frame, the precise position information of the foreground material in each corresponding frame can be determined, resulting in a video of a young woman walking along the trajectory P1→P2→P3→P4→P5. Thus, using the method proposed in this application, precise control over the motion trajectory of a specified target object (one or more materials) in an image can be achieved, specifically, the position information can be set precisely every frame or every N frames.
[0159] It should be noted that, for the motion trajectory, the input control conditions can be an image sequence, such as an image sequence with 3D bounding boxes. Alternatively, in some embodiments, the control conditions for motion attribute input can be only the boundary coordinate information corresponding to the 3D bounding boxes in each frame of the image. When setting the control conditions, the user can upload a file of a specified format, which stores arrays, vectors, matrices, and other data used to record the boundary coordinate information of the 3D bounding boxes in each frame of the image.
[0160] Environment is used to set environmental parameters in the target video. For example, environmental parameters can be environmental parameters such as weather and light intensity in each frame of the video, or they can be parameters such as video style.
[0161] Taking environmental parameters, including weather, as an example, Figure 9As shown, when a video of a dog flying in the sky needs to be generated, the user can set the weather to change from cloudy to rainy and then to sunny, for example, the user can specify that the weather of the first to M frames is cloudy, the weather of the M+1 frame to the M+Q frame is rainy, and the weather of the M+Q+1 frame to the X frame is sunny. Wherein, M, Q, X are integers, and M+QX.
[0162] It should be noted that in the related art, a corresponding video can be generated by inputting the text prompt "a dog is flying in the sky, and the weather changes from rainy to sunny", but the position (i.e. which frame) of the weather change in the video cannot be accurately controlled. In some scenarios that require accurate control of environmental parameters, the user's demand cannot be accurately matched. The method proposed in the embodiments of the present application can accurately control the correspondence between the weather and each frame of image, and can control the environmental parameters of each frame of image in the video or control the environmental parameters every N frames.
[0163] Physical property, used to set the occurrence of a specified physical phenomenon in the target video. A physical phenomenon is a kind of phenomenon that occurs based on a certain physical law, and a video generation model can generate a phenomenon that conforms to this physical law by learning the physical law.
[0164] For example, the physical phenomenon can be rigid body motion or non-rigid body motion. A rigid body is an object whose shape and size do not change when subjected to force or motion, and the relative positions of its internal points do not change. A non-rigid body is an object whose shape and size change when subjected to force or motion, and the relative positions of its internal points change. Rigid body motion has no deformation or deformation velocity. Non-rigid body motion, i.e. deformation or deformation velocity, such as linear deformation rate and angular deformation rate. The method proposed in the embodiments of the present application supports user setting whether non-rigid body motion or rigid body motion needs to occur in the video.
[0165] For example, in the automatic driving scene, in order to improve the safety performance of the automatic driving system, some video data needs to be used to train the automatic driving system, and it may be necessary to generate a video of a vehicle collision accident, which can be understood as a non-rigid body motion occurring in the video.
[0166] The physical phenomenon can also be other phenomena that satisfy the physical law, for example, whether an optical phenomenon (shadow of an object, rainbow, mirror emission, etc.) occurs, whether a mechanical phenomenon (motion and stillness of an object, deformation and vibration of an elastic object, flow of a fluid) occurs, whether an acoustic phenomenon (such as echo, tone change and reflection of sound waves) occurs, whether a thermal phenomenon (thermal expansion and contraction, etc.) occurs. Heating causes changes in the shape of ice, metal and other objects, etc.
[0167] According to the above description, it can be seen that the control condition can be text, a single image or an image sequence, or a file in other formats that can be uploaded from the file system.
[0168] It should be noted that for the control of multiple attributes, each attribute can be independently controlled, or the control conditions can be associated to control multiple attributes in combination.
[0169] For example, in the automatic driving scenario, to train an automatic driving system, it can be necessary to generate a high-fidelity video simulating a real accident scene. For example, a traffic accident such as a collision, friction, or scratching needs to occur in the video.
[0170] For example, in the automatic driving scenario, to train an automatic driving system, it can be necessary to generate a high-fidelity video simulating a real accident scene. For example, a traffic accident such as a collision, friction, or scratching needs to occur in the video.
[0171] Through the first control condition, the material categories in the material structure include a car and a tree beside the road, and various materials have corresponding IDs, and the material positions are the specific positions of the car and the tree beside the road in the image. The second control condition can be associated with the first control condition, and according to each material category in the material structure, a car is selected as a control object, and the motion trajectory of the car is controlled, which is a motion trajectory that can interact with the tree. For example, the interaction can be a collision. In addition, through the third control condition, the physical phenomenon to be generated in the video is controlled to be a non-rigid body motion, for example, a collision that causes deformation.
[0172] Through the joint control of the above-mentioned three attributes, the car in the generated video can interact with the tree beside the road when moving along the set motion trajectory, and the car deforms in the interaction.
[0173] For example, the first control condition is also used to control the material categories to include multiple pedestrians, and the second control condition is also used to control the motion trajectories of the multiple pedestrians. Through the association setting of the motion trajectories of the multiple pedestrians and the car, the generated video can make a car along a specified motion trajectory interact with multiple pedestrians moving along another motion trajectory, for example, the interaction can be a collision.
[0174] To facilitate understanding of the video generation method proposed in the embodiments of the present application, the electronic device 100 with the structure shown in Figure 3 or the server 200 with the structure shown in Figure 4 will be taken as examples, combined with the application scenarios shown in Figure 5 and the system architecture shown in Figure 6 , the video generation method provided by the embodiments of the present application will be exemplarily described.
[0175] As shown in Figure 10 ,Figure 10 A software module architecture is shown, which can be built based on one or more servers 200, or can be built based on Figure 6 the system architecture shown.
[0176] The embodiment of the present application proposes a software module architecture for implementing the video generation method proposed by the embodiment of the present application. From the perspective of software implementation, the software module architecture at least includes a controllable video generation module (referred to as a video generation module), and in some embodiments, can further include a quality feedback module. Alternatively, a simulation data generation module can be further included.
[0177] In the embodiment, the video attributes are divided into structure, motion, environment and physics, which are controlled respectively.
[0178] Video generation module: used for encoding (Encoder) the control conditions corresponding to the four attributes of structure (such as 3D structure), motion, environment and physical law, inputting the control conditions into the encoder, embedding the video generation model according to the hidden vector obtained by encoding, and training the processed video generation model. The training data can be simulation video data obtained by the simulation data generation module, or other data sets used for training.
[0179] The physical feedback module can be deployed in the video generation module, which is used to evaluate the video generated in the training process to determine whether it meets the preset control condition, give the corresponding scoring order, and optimize the trainable parameters in the controllable video generation model according to the score. For example, the training method is supervised learning, the training data is a video with control label (i.e. control condition as label), and the scoring method can be to calculate the value of the loss function, and to optimize the trainable parameters in the video generation model according to the value of the loss function. After training, the video generation model can accept user control in 3D structure, motion, environment, physics and other aspects, and generate videos meeting the control condition.
[0180] Quality feedback module: the quality of the video generated by the video generation model is evaluated by objective indicators and / or subjective indicators, and if the application requirement is not met, incremental training (Incremental Training) is triggered, which can be specifically triggered by simulation data generation, and the generated simulation video data is used to perform incremental training on the video generation model until the quality of the controllable video generation module meets the application requirement.
[0181] The simulation data generation module: the simulation data generation module, through the 3D reconstruction algorithm to obtain the material or the material in the existing material library, constructs the 3D simulation scene related to the application scene, and then renders the video data with control tags related to the application scene under the control of the control conditions corresponding to the 3D structure, motion, environment, physics and other attributes.
[0182] In this way, a video generation model capable of generating videos conforming to various control conditions can be obtained, and the controllability of video generation is enhanced.
[0183] In the above description based on the software module architecture, three stages are involved: the training stage, the use stage and the incremental training stage. To prevent confusion, as shown in Figure 11 The whole process is divided into three stages: training stage, use stage and incremental training stage.
[0184] Training stage:
[0185] As shown in Figure 10 and Figure 11 First, based on the training data, the training of the video generation model is performed. The process of the training stage can be implemented on the service side, for example, implemented by one or more servers 200 in the system architecture shown in Figure 4 .
[0186] The training data includes input data and control conditions as labels (referred to as control labels). For example, the hidden vector obtained after encoding the control condition can be used as the label. Alternatively, in other embodiments, the control condition can be directly used as the label.
[0187] Exemplarily, the training data therein can be the simulation video data obtained by the simulation data generation module shown in Figure 10 , which can be obtained in the following way:
[0188] According to the user operation, the video material and the 3D simulation scene are determined. For example, by calling a 3D reconstruction algorithm model, the material required for the video is generated; or according to the user's instruction, the user's required material is searched in the specified material library, or the material is determined according to the user's selection operation. Exemplarily, the 3D reconstruction algorithm model therein can be a neural radiation field (NeRF) or a 3D Gaussian splatting model.
[0189] After saving the video material, the 3D simulation scene required for the target video can be constructed based on the existing material and combined with the actual application scenario. It should be noted that the 3D simulation scene in the training data can be generated in response to user operations and / or instructions corresponding to the user operations and / or instructions, which can be understood as constructing a 3D simulation scene manually built by the user based on user operations. Alternatively, the 3D simulation scene can be automatically constructed by a 3D-GPT or other 3D simulation model. 3D-GPT, i.e., a 3D modeling model based on a large language model (Procedural 3D Modeling With Large Language Models).
[0190] The determined material and 3D simulation scene are saved or cached, and the 3D scene containing various 3D materials is imported into the 3D simulation model. The 3D simulation model receives the control conditions set by the user, such as receiving 3D structure, motion, environment, physics, and other control conditions, and receiving text, images, and other conventional input data that the user may input. According to the input data and control conditions, the video data related to the application scenario is rendered, the video data is taken as a target video sample, and the corresponding control conditions are taken as labels to obtain a target video sample with control labels.
[0191] For example, the 3D simulation model can use a rendering, physics engine such as UE and / or Blender, or a domain simulator such as Carla and Gazebo. The specific embodiments of the present application are not listed one by one.
[0192] Next, the control conditions are encoded, that is, the control conditions corresponding to the structure, motion, environment, and physics are respectively encoded by an encoder to obtain the hidden vectors corresponding to the various control conditions.
[0193] For example, the encoder can use one or more of the following encoders: STC-encoder, SparseCondition Encoder, multilayer perceptron (MLP, Multilayer Perceptron, MLP), or other encoders.
[0194] For different control conditions, the same encoder can be used, or different encoders can be used for encoding. The same encoder can encode different control conditions at different times; or multiple encoders can be used in parallel to encode multiple control conditions.
[0195] The various control conditions are encoded into a latent space vector and embedded into the video generation model. The embedding method can be to add a first variable that participates in cross-attention calculation and subsequent calculation of the model in the cross-attention layer. In the training stage or the inference stage, after obtaining the latent vector corresponding to the control condition, the value of the latent vector is assigned to the first variable. And / or, the embedding method can be to embed a second variable in the noise vector. In the training stage or the inference stage, after obtaining the latent vector corresponding to the control condition, the value of the latent vector is assigned to the second variable. The first variable or the second variable can be a vector with the same dimension as the latent vector, and the elements in the vector are variables.
[0196] In this embodiment, the video generation model can be a diffusion model or an improved model based on diffusion, for example, it can be a Video Diffusion Model.
[0197] Alternatively, in other embodiments, the video generation model can also be other video generation models containing cross-attention layers, or a combination of diffusion and other video generation models. For example, the video generation model can be an autoregressive model containing a cross-attention layer, and exemplarily, it can be one or more combinations of Temporal Generative Adversarial Net (TGANv2), Video Generation using VQ-VAE and Transformers (VideoGPT), Dual Variational Generation (DVG), and A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2 (StyleGAN-V) based on a generative adversarial network.
[0198] Exemplarily, the embedding of the latent vector can be implemented by one or more of the following embedding methods:
[0199] Embedding method one: embedding through the cross-attention layer.
[0200] The embodiments of the present application propose a computer mechanism for embedding control conditions in the cross-attention layer.
[0201] The video generation model can be a Diffusion model including cross attention. The cross attention layer can be used to process the association between multiple different modal sequences, for example, can be used to perform cross-modal attention calculation between the image modal and the text modal. Specifically, in Diffusion, the cross-modal attention calculation can be performed on the intermediate features (or called latent vectors) corresponding to the generated object Image and the embedded vectors Context Embedding corresponding to the text.
[0202] In the embodiments of the present application, the cross attention layer is used to perform cross attention calculation (cross-modal attention calculation) between the latent vectors corresponding to the generated object and the latent vectors corresponding to the control condition. Wherein, the generated object can be a single image or an image sequence (i.e. video) including multiple images, and the control condition can be text, a single image or an image sequence.
[0203] Specifically, in the embodiments, the following two ways are provided to realize the cross attention calculation between the latent vectors of the generated object and the latent vectors corresponding to the control condition:
[0204] Cross attention method one:
[0205] If the control condition contains an image or a video (i.e. an image sequence), the latent vector corresponding to the image or the video in the control condition can be embedded into the latent vector corresponding to the image or the video of the generated object. For example, the embedding can be to do Concat between the image or the video in the control condition and the image or the video of the generated object, and then perform cross attention calculation.
[0206] For example, as shown in Figure 12 The encoder Encoder is used to encode the control condition provided by the user to obtain the latent vector corresponding to the control condition. Assuming that the control condition contains two modalities of text and image, after encoding, the latent vector corresponding to the text in the control condition Control Context Embedding and the latent vector corresponding to the image in the control condition Control Image Embedding can be obtained. The latent vector corresponding to the image in the control condition Control Image Embedding is embedded into the latent vector of the image Image, and Image is the generated object. The embedding can be to do Concat between Control Image Embedding and Image to obtain the Image embedded with the control condition (hereinafter referred to as the embedded Image).
[0207] It should be noted that the Concat can be to call the Concat function to connect or combine two objects. The Concat function is only an example, and other functions or other ways can be used instead, such as using the concatenate() function, etc. Alternatively, embedding can be understood as the fusion of Control Image Embedding and image features of Image. Other image feature fusion methods can also be used.
[0208] Next, the cross attention calculation can be performed on the Control Context Embedding corresponding to the text in the control condition and the embedded Image.
[0209] Specifically, the embedded Image is processed by the reshaping layer, input into the Q (query) corresponding linear (Linear) layer, and the Q matrix is generated based on the embedded Image.
[0210] The Control Context Embedding corresponding to the text in the control condition is directly input into the K (key) and V (value) corresponding linear (Linear) layer based on the text modal data, and the K and V matrices are generated based on the text modal data.
[0211] Next, as shown in Figure 12 , for each element in Q, the correlation between it and the corresponding position element in K is calculated to obtain a similarity matrix, and then as described in Figure 12 , the calculation formula is used to normalize the similarity matrix to obtain a weight matrix representing the correlation between image Q and text K elements, and then multiplied by V to perform weighted summation to obtain new image data. After linear layer processing, the Q matrix, K matrix and V matrix have the same dimension, and d represents the dimension of the matrix after linear layer processing.
[0212] Among them, the role of the linear layer (Linear Layer) is to perform linear transformation on the input data of the corresponding modal. The linear layer (Linear Layer) can also be a fully connected layer (Fully Connected Layer) or a dense layer (Dense Layer).
[0213] cross attention way two:
[0214] As shown in Figure 13As shown, another way to embed the control condition hidden vector through the cross attention layer can be to set multiple cross attention layers, and embed multiple hidden vectors corresponding to multiple control conditions layer by layer.
[0215] As shown, for example, assuming that the user sets 3 control conditions, including 2 text format control conditions and 1 image format control file, 3 hidden vectors are obtained, which are: Control Context Embedding1, Control Context Embedding2, and Control Image Embedding. Figure 13
[0216] At least 3 cross attention layers can be designed in Diffusion, which are: Cross Attention Layer1, Cross Attention Layer2, and Cross Attention Layer3.
[0217] In Cross Attention Layer1, cross-modal attention calculation is performed between the generated object Image I0 and Control Context Embedding1 to obtain the associated image feature Image I1. For details of the calculation process, see cross attention method one.
[0218] In Cross Attention Layer2, Image I1 obtained in Cross Attention Layer1 is taken as one of the modes participating in the cross-modal calculation of this layer, and cross-modal attention calculation is performed between Image I1 and Control Image Embedding corresponding to the image in the control condition to obtain the associated image feature Image I2.
[0219] In Cross Attention Layer3, Image I2 obtained in Cross Attention Layer2 is taken as one of the modes participating in the cross-modal calculation of this layer, and cross-modal attention calculation is performed between Image I2 and Control Context Embedding2 to obtain the associated image feature Image I3.
[0220] It should be noted that the above Figure 12 or Figure 13 For the sake of clear view, only the case that the generated object Image or the image in the control condition is a single image is shown, in fact, the single image can be a sequence of images, and it should not be understood as being limited to processing a single image.
[0221] Embedding mode two: concatenated into the noise vector, or replace the noise vector.
[0222] The noise vector, i.e., the embedding of random noise, can also be referred to as the embedding vector corresponding to random noise, or simply referred to as the noise vector.
[0223] As shown in Figure 14 , one implementation is to use the Concat function to concatenate the random noise embedding vector Noise embedding and the control condition hidden vector Control embedding, and the concatenated embedding vector is used as Noise embedding for subsequent model calculation.
[0224] Similarly, the Concat function here is only an example, and other functions or other connection methods can be used for replacement, such as using the concatenate() function, or using other methods that can realize vector splicing or combination.
[0225] As shown in Figure 15 , another implementation is that the control condition hidden vector Controlembedding obtained by encoding replaces the random noise embedding vector Noise embedding, that is, Control embedding is used as Noise embedding to participate in subsequent calculation.
[0226] The above-mentioned embedding mode one and embedding mode two can be implemented at the same time, or only one of the embedding modes can be used. In this embodiment, both of the above-mentioned embedding modes are used.
[0227] It should be noted that the above-mentioned embedding mode is described by taking diffusion as an example, and for other video generation models, the embedding mode can also be to embed the hidden vector corresponding to the control condition in the cross-attention layer, and / or to embed the hidden vector corresponding to the control condition in the input layer. The specific embedding mode can be designed according to the above-mentioned exemplary description and in combination with the hierarchical architecture of the video generation model, and this specification does not list them one by one.
[0228] The video generation model can be trained after the embedding process. It should be noted that embedding has two meanings. One meaning is to modify the hierarchical architecture of the model, which can be achieved by modifying the code of the model. The other meaning is to input the specific value of the control condition input by the user during the running of the model. In the training stage, the value of the control condition in the training sample is input in one iteration.
[0229] The input data in the training data is input into the video generation model, and the target video is output by the video generation model.
[0230] As shown in Figure 10 and Figure 11 In the training stage, the physical feedback module is used to evaluate whether the target video generated in the training process meets the control condition and give the corresponding score ranking. The video generation model is optimized according to the score.
[0231] The physical feedback module can extract the control condition corresponding to the generated video by calling a specified model, compare the extracted control condition with the control condition label, and calculate the value of the loss function. Then, the trainable parameters in Diffusion are optimized according to the value of the loss function.
[0232] As described above, the background structure or foreground structure in the structural attribute can be represented by a background semantic graph or a foreground semantic graph. The latter does not distinguish between foreground and background, and represents the entire image structure as a semantic segmentation graph. Therefore, when the control condition contains a semantic graph, the semantic segmentation model can be used to extract the semantic segmentation graph (referred to as semantic graph) in the target video, such as the background semantic graph or the foreground semantic graph.
[0233] For the foreground structure, a 3D detection model can also be used to detect the material position of each frame image in the target video.
[0234] Then, the extracted semantic graph is compared with the semantic graph in the control label, and the detected material position, such as the bounding box coordinate information detected by the 3D Bbox, is compared with the bounding box information of the 3D Bbox in the control label, to obtain the index score value corresponding to the structural attribute.
[0235] Regarding the motion attribute, the 3D tracking model can be used to track the motion trajectory of the material in the foreground of each frame image in the video, and compare it with the motion trajectory in the control label to obtain the index score value corresponding to the motion.
[0236] As for the environmental attribute, the environmental parameters in the target video can be detected by a multi-modal understanding model, and compared with the environmental parameters in the control label to obtain the index score value corresponding to the environment. The multi-modal understanding model can include but is not limited to one or more of VideoLLaVA, GPT4Video, etc.
[0237] As for the physical attribute, the multi-modal understanding model can be used to detect whether the target video has a predetermined physical phenomenon, such as whether there is a non-rigid body motion, and compare the detection result with the physical control condition in the control label to obtain the index score value corresponding to the physical attribute.
[0238] Alternatively, for both motion and physical attributes, 3D reconstruction can be performed on the target video, and then physical simulation can be used to extract the corresponding control condition. For example, 3D reconstruction is performed on the target video to obtain a 3D model, and physical simulation is used to predict the deformation of the object in the video or the motion trajectory of the object. For example, the Lagrangian mathematical model can be used to obtain a predicted differential equation, and the differential equation can be used to predict whether deformation occurs.
[0239] The value of the loss function can be obtained according to the index score values corresponding to the above-mentioned multiple control conditions, for example, the multiple index score values corresponding to the multiple control conditions are weighted and summed. By setting the weights of various control conditions, the influence of each control condition on the video relative to other control conditions can be adjusted. The index score value can be obtained by using a mean squared error loss function (MSE) or a cross-entropy loss function (CE) to calculate the two objects compared. The two objects compared can be the hidden vectors obtained after encoding, i.e., the hidden vectors of the control label after encoding. The control condition extracted from the target video generated by the video generation model can be encoded, and the obtained hidden vector can be compared with the hidden vector of the control label to calculate the loss function.
[0240] In this way, according to the above description, the video generation model with expected performance can be obtained by multiple iterations of training. The expected performance refers to the ability to generate a target video that meets the control condition.
[0241] Usage stage:
[0242] The trained video generation model can be deployed on a user-side electronic device, such as a server, a mobile phone, a tablet computer, a smart speaker, a smart display, a smart television, a smart watch, a smart glasses, a smart car, a smart home, a smart city, etc. Figure 6In the terminal device 304 on the user side shown in the figure, the video generation model can be run in a local call manner to obtain the target video that the user wants. Alternatively, the trained video generation model can be deployed in a server on the service side, for example, deployed in the server 301 shown in the figure. Figure 6 In the terminal device 302 on the user side shown in the figure, the terminal device can remotely obtain the target video that the user wants through the video generation model deployed in the service 300 in a remote call manner. Figure 6
[0243] Exemplarily, as shown in FIG. 16, in the terminal device on the user side, an interface example shown in FIG. 16 can be displayed.
[0244] In the interface, a first control 401 for inputting control conditions and a second control 402 for inputting prompt information are displayed.
[0245] Through the first control 401, input boxes for inputting control conditions corresponding to multiple attributes such as structure, motion, environment, and physics are displayed, and the input boxes support multi-modal input, for example, text input and uploading files. The files can be images in various image formats, for example, semantic segmentation images, 3D Bbox images, and the like, and other formats of data or various formats of files in the file system can also be supported.
[0246] For example, Figure 8 The coordinate data corresponding to the 3D Bbox image shown in the figure can be stored in a file in a specified format, and the user only needs to click the right side "+" in the input box corresponding to the structure attribute shown in FIG. 16, select the file for saving the coordinate data, and upload it.
[0247] For another example, the motion trajectory can be represented in the form of an array, a tensor, or a matrix, and the input box corresponding to the motion can upload a file containing array or tensor, matrix, or the like.
[0248] In this embodiment, the existing video generation method can be compatible. In the existing video generation method, the user can input text and / or images. For example, in the interface example shown in FIG. 16, the second control 402 for inputting text and / or uploading images is displayed, and the text prompt input by the user in the prompt input box is "a man and a woman cross the road without walking on the zebra crossing".
[0249] It should be noted that, in the solutions proposed in this application, users can choose not to set control conditions, or select one or more control conditions from a variety of options. For example, in the interface example shown in Figure 16, the user only entered the text information "non-rigid body motion" in the input box corresponding to the physical attribute. In other embodiments provided in this application, the user needs to input at least two of the control conditions corresponding to various attributes such as structural attributes, motion attributes, environmental attributes, and physical attributes. For example, the user needs to input the control conditions corresponding to the structural attributes and the control conditions corresponding to the motion attributes, thereby controlling at least one material in the target video to move along the motion trajectory defined by the motion attributes according to the initial position defined by the structural attributes; or, the user needs to input the control conditions corresponding to the motion attributes and the control conditions corresponding to the physical attributes, thereby controlling at least one material in the target video to move according to the motion trajectory defined by the motion attributes, using the physical phenomena defined by the physical attributes. And so on. This application does not limit the number of control conditions that the user needs to input.
[0250] Based on the user-input control conditions and text prompts, the encoder is invoked to encode the user-input control condition "non-rigid body motion" to obtain a latent vector. Additionally, the user-input text prompt "a man and a woman crossed the road without using the crosswalk" is processed to obtain an embedded vector.
[0251] Next, the user-side electronic device 303 (desktop computer) calls the local video generation model or remotely calls the video generation model on the server side to obtain the target video output by the video generation model.
[0252] When running the video generation model locally, the latent vectors corresponding to the control conditions and the embedding vectors corresponding to the text prompts can be input into the video generation model. When making a remote call, the latent vectors corresponding to the control conditions and the embedding vectors corresponding to the text prompts can be included in the call request.
[0253] In response to the call request, the video generation model begins to run, such as Figure 17 As shown, during operation, latent vectors corresponding to control conditions are embedded in the input layer and the cross-attention layer. Specifically, in the Diffusion input layer, the latent vectors of the control conditions and the noise vector are concatenated using the concat function to obtain the embedded noise vector; and in the Diffusion cross-attention layer, the value of the latent vector control contextembedding obtained after encoding the control conditions is embedded into the context embedding.
[0254] The video generation model outputs the target video desired by the user.
[0255] Next, the target video is scored using a quality feedback module.
[0256] Specifically, the quality feedback module can comprehensively evaluate the video quality generated by the controllable video generation model and whether the control condition meets the application requirements through subjective indicators and objective indicators. If the application requirements are met, the process ends; if the application requirements are not met, such as Figure 11 as shown, the incremental training of the video generation model is triggered. Specifically, subsequent simulation data generation can be triggered for incremental training of the controllable video generation module.
[0257] Exemplarily, the subjective evaluation indicator can calculate the score value of the subjective indicator based on user operations through an interactive interface. For example, multiple videos generated under controllable conditions are displayed on the interactive interface, and the user selects the available videos, and the available proportion is counted. If it is higher than a certain threshold, it is considered to meet the application requirements.
[0258] The objective evaluation indicator can include video quality evaluation indicators (Frechet Video Distance, FVD), CLIP (Contrastive Language-Image Pre-training) Similarity, and other control condition compliance indicators. Among them, FVD can use Inflated-3D Convnets (I3D) pre-trained on Kinetics to extract features from video clips, and calculate the FVD score by calculating the combination of the mean and covariance matrix.
[0259] CLIP Similarity can be the cosine similarity between each frame image in the target video and the text input by the user calculated using the CLIP model, realizing a cross-modal similarity measurement method between the text input by the user and the output video.
[0260] The total score value obtained by the quality feedback module can be the comprehensive calculation result of the subjective evaluation indicator score value and the objective evaluation indicator score value, for example, the weighted sum of the subjective evaluation score value and the objective evaluation score value.
[0261] It should be noted that in some embodiments, the objective indicator can also include the evaluation indicator of the control condition compliance of the physical feedback module used in the training stage, that is, in the running of the quality feedback module, the physical feedback module can be called to extract the control condition corresponding to the current output target video, compare it with the user input control condition, and then evaluate whether the generated video can meet the control condition set by the user.
[0262] If the quality score evaluated by the quality feedback module does not reach the threshold value, the process of incremental training is triggered, and the video generation model is iteratively trained in the incremental training stage until the quality score is equal to or greater than the threshold value.
[0263] Incremental training stage:
[0264] In the incremental training stage, according to the user's real input control conditions and text, etc., new training data is obtained through the 3D simulation module, the user's actual input control conditions are taken as labels, and the embedding vector corresponding to the text, the noise vector including the hidden vector corresponding to the embedding control condition, and the vector corresponding to the time step are input into the video generation model to output the target video.
[0265] In the incremental training stage, the training process of the video generation model can refer to the training stage, which will not be described here.
[0266] It should be noted that in the training stage, the value of the loss function is calculated by the physical feedback module, and the trainable parameters in the video generation model are optimized according to the value of the loss function until the model converges. In the incremental training stage, a quality score is evaluated by the quality feedback module, and if the quality score is lower than the threshold value, the video generation model is iteratively trained until the quality score is equal to or greater than the threshold value, and the incremental training is stopped. When the next obtained quality score fails to reach the threshold value, the incremental training process is triggered again.
[0267] As can be seen, the embodiments of the present application propose a method capable of stimulating the controllability of the video generation model, so that a controllable video generation model can accept various combinations of control conditions at the same time, generate videos that better match user needs, and match user needs at a finer granularity.
[0268] In addition, the iterative optimization system of controllable video generation and visual simulation proposed by the embodiments of the present application provides effective guarantee for the controllability compliance of the controllable video generation model, and guarantees the performance of the video generation model and the usability of the generated video.
[0269] According to the above exemplary description, from the perspective of the electronic device on the user side, the video generation method provided by the embodiments of the present application can include the following processes as shown in Figure 18
[0270] S10: receiving at least one control condition.
[0271] Exemplarily, in some embodiments, receiving at least one control condition can be through a visual interface, for example, the electronic device displays a first control, and receives the user's input text or uploaded file through the first control.
[0272] The electronic device can beFigure 3 The electronic device 100 shown in the figure, or Figure 6 The terminal device 302 or the terminal device 304 shown in the figure.
[0273] As shown in the figure, Figure 16a The first control 401 is configured to receive at least one control condition set by a user. The user operation includes inputting text and uploading files. For example, various attributes such as structure attribute correspond to input boxes respectively, and the input boxes display "Please enter or upload files here". The user can select a control condition in the form of input text, or select a file and upload it by clicking the "+" button on the right.
[0274] It should be noted that in other embodiments, the visualization interface can not be set, but an application programming interface (API) can be set. In the process of calling the API, the user can indicate at least one control condition. The visualization interface such as the one shown in the figures Figure 16a Or 16b can not be displayed.
[0275] S11: The electronic device encodes the at least one control condition into at least one latent vector.
[0276] After determining the at least one control condition set by the user through the interface such as the one shown in the figure Figure 16a Or by other means, an encoder can be called to encode various control conditions respectively to obtain at least one latent vector. The encoder can be deployed on the server side or in the electronic device on the user side. The electronic device can call the encoder locally or remotely. For the encoder that can be called, please refer to the above exemplary description, which will not be repeated here.
[0277] The at least one control condition is used to control at least one attribute of the target video.
[0278] S12: Input the at least one latent vector into a pre-trained video generation model, and generate a target video that meets the at least one control condition through the video generation model.
[0279] As mentioned in the above exemplary description, the video generation model can be deployed in the server or in the electronic device on the user side. Therefore, the electronic device on the user side can call the pre-trained video generation model locally or remotely, and input the at least one latent vector into the video generation model.
[0280] According to the above exemplary description, the video generation model can be Diffusion, at least one latent vector is embedded in the video generation model, at least one latent vector can be embedded in the cross attention layer, and / or at least one latent vector is embedded through a noise vector corresponding to random noise. The input data of Diffusion can include embedding vectors corresponding to text and / or images, noise vectors, and time step vectors. Among them, the noise vector, i.e. the noise vector corresponding to the random noise, at least one latent vector is embedded in the noise vector, which can be to realize the connection of at least one latent vector and the noise vector by using functions such as concat, or at least one latent vector is used as a noise vector, i.e. the latent vector replaces the original random noise noise vector.
[0281] In some embodiments, at least one attribute in the target video can be one or more of the following attributes: structural attribute, motion attribute, environmental attribute, and physical attribute.
[0282] Among them, the structural attribute can be the material structure and / or camera trajectory in the target video, for example, Figure 7 The background semantic graph and 3D Bbox in the above can be used to control the material structure.
[0283] The motion attribute can be the motion trajectory of at least one material in the foreground in the target video, for example Figure 8 The coordinate data corresponding to the 3D Bbox of each frame from the 1st frame to the 1+4Nth frame in the above can be used as a control condition to control Figure 8 The motion trajectory of the female character material shown in the above.
[0284] The environmental attribute can be the environmental parameter information in the target video, for example, the environmental parameter can include weather, Figure 9 As shown in the above, the weather is controlled to be rainy in the M+1th to M+Qth frames, and the weather is controlled to be sunny in the M+Q+1th to Xth frames.
[0285] The physical attribute can be a physical phenomenon appearing in the target video. For example, by inputting the text of the physical phenomenon, the physical phenomenon corresponding to the text is controlled to appear in the video. For example, Figure 16a In the above, the text "non-rigid body motion" is input in the input box corresponding to the physical attribute, and the non-rigid body motion is controlled to appear in the video, for example, in the generated video, when a man and a woman cross the road without passing through the zebra crossing, an accident occurs, and the vehicle is deformed.
[0286] It should be noted that the above-mentioned multiple controllable attributes are only examples, and other multiple controllable attributes can also be obtained by other division manners. For example, it can be divided into two controllable attributes of intra-frame and inter-frame. The intra-frame attribute is mainly used to set the characteristics of various materials in a picture frame, such as material structure, environment, etc. Inter-frame is mainly used to set the relative position relationship between multiple frames and other characteristics, such as motion trajectory, camera trajectory, etc. which can be used as inter-frame attributes. Or, it can also be divided into static attributes or dynamic attributes, etc. The division manner of the attribute is not listed one by one in this specification.
[0287] In addition, the video generation method provided in the embodiments of the present application can be compatible with the existing text / image control mode in some embodiments, that is, it supports text / image control and also supports setting control conditions. For example, as shown in Figure 16a , the first control 4021 and the second control 402 are displayed on the interface. The second control is used to determine the input data of the video generation model according to the user operation, and the input data includes the text prompt input in the “prompt” input box and / or the image uploaded through the “image” button.
[0288] In other embodiments, as shown in Figure 16b , only the user setting control condition can be supported, and the user setting text / image is not supported, and only the first control 401 is displayed. In the interface example shown in FIG. 16, for the input box of the structure attribute, the user can input the text “foreground: two pedestrians, one male and one female; background: street with zebra crossing, trees on both sides of the street”, and in the input box corresponding to the motion attribute, the user can input “two pedestrians crossing the street” or upload coordinate data or images for accurately controlling the motion trajectory of the two pedestrians.
[0289] Therefore, the method provided in the embodiments of the present application displays at least the first control 401 on the electronic device.
[0290] It should be noted that the interface shown in Figure 16a or Figure 16b is only an example, and other multiple interface layout styles can be set in actual applications, which are not limited to the examples shown in Figure 16a or Figure 16b . It should be noted that in actual applications, an interactive interface can also not be set, but a data interface for user calling can be provided, for example, an API calling interface is provided, and the user can obtain a video meeting the control condition by calling the API after obtaining the control condition.
[0291] In the embodiments of the present application, the setting of the control condition can involve 3D space and motion, and therefore a tool with a 3D editing interface can be used to generate the control condition, for example, the motion trajectory can be edited through blender.
[0292] Further, it needs to be noted that the method proposed in the embodiments of the present application can be used as a plug-in compatible with some existing tools. For example, the software program product obtained based on the method proposed in the embodiments of the present application can be a plug-in, which can be installed into an existing 3D simulation tool with a 3D editing interface or a video generation tool with a video editing interface. When a user uses these tools, the user sets control conditions through the editing interface of the tool itself, and then runs the plug-in. The plug-in has the permission to read these control conditions. After the plug-in reads the control conditions, the plug-in generates video data meeting the control conditions, and the video data can be played in the 3D simulation tool or the video generation tool.
[0293] Figure 18 In the method flow shown, the video generation model is a trained model. In some embodiments, the video generation model needs to be trained first. Specifically, before the first control is displayed, the video generation model can be trained in the following manner:
[0294] In combination with Figure 10 and Figure 11 As shown, first, training data with labels is obtained. The labels can be at least one first hidden vector corresponding to at least one control condition.
[0295] After the training data is obtained, in one iteration of training the video generation model, the value of the loss function corresponding to the target video output by the video generation model can be obtained by calling the physical feedback module, and the trainable parameters in the video generation model are optimized according to the value of the loss function. Specifically, at least one control condition can be extracted from the target video output by the video generation model, and at least one second hidden vector corresponding to the at least one control condition extracted from the target video can be obtained, for example, the at least one control condition extracted from the target video can be encoded to obtain the at least one second hidden vector.
[0296] According to the at least one second hidden vector and the at least one first hidden vector, the value of the loss function is calculated, for example, the loss function can be a cross entropy loss or other loss function, which is not listed one by one here. According to the value of the loss function, the trainable parameters in the video generation model are optimized. Such iteration is performed until the model converges.
[0297] Exemplarily, at least one control condition is extracted from the target video output by the video generation model. The control condition can be extracted by calling a specified model. For example, a semantic segmentation model and / or a 3D detection model can be called to extract a structure corresponding control condition from the target video output by the video generation model. For example, a 3D tracking model can be called to extract a motion corresponding control condition from the target video output by the video generation model.
[0298] For example, a multi-modal understanding model can be called to extract an environment corresponding control condition from the target video output by the video generation model.
[0299] For another example, a multi-modal understanding model can be called to extract a physics corresponding control condition from the target video output by the video generation model.
[0300] It should be noted that, in the training phase or the incremental training phase shown in Figure 11 In the training phase or the incremental training phase shown in
[0301] In addition, in the use phase shown in Figure 11 In addition, in the use phase shown in
[0302] As shown in Figure 19 As shown in Figure 4The server 200 shown, or the server 301 or the server 303 shown in the middle. The method can include: Figure 6
[0303] S20: The server receives the calling request sent by the terminal device.
[0304] The calling request can include at least one control condition.
[0305] S21: In response to the calling request, a video generation model is run to generate a target video that meets the at least one control condition.
[0306] Among them, at least one hidden vector is embedded in the video generation model, and the at least one hidden vector can be obtained based on the at least one control condition, for example, is obtained by encoding the control condition. The at least one control condition is used to control at least one attribute in the target video.
[0307] S22: The target video is sent to the terminal device.
[0308] The application also provides a video generation device, including:
[0309] A first receiving module is configured to receive at least one control condition. The at least one control condition is used to control at least one attribute of a target video to be generated, and the at least one attribute includes one or more of the following: structural attribute, motion attribute, and physical attribute.
[0310] An encoding module is configured to encode the at least one control condition into at least one hidden vector.
[0311] A generation module is configured to input the at least one hidden vector into a pre-trained video generation model, and generate a target video that meets the at least one control condition through the video generation model.
[0312] Among them, the first receiving module, the encoding module and the generation module can be realized by software or by hardware. For example, the implementation of the first receiving module is introduced as follows. Similarly, the implementation of the encoding module and the generation module can refer to the implementation of the first receiving module.
[0313] As an example of a software functional unit, the first receiving module can include code running on a compute instance. The compute instance can include at least one of a physical host (computing device), a virtual machine, a container. Further, the compute instance can be one or more. For example, the first receiving module can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code can be distributed in the same region, or in different regions. Further, the multiple hosts / virtual machines / containers for running the code can be distributed in the same availability zone (AZ), or in different AZs, each AZ including one data center or multiple data centers in close geographical proximity. Generally, one region can include multiple AZs.
[0314] Similarly, the multiple hosts / virtual machines / containers for running the code can be distributed in the same virtual private cloud (VPC), or in multiple VPCs. Generally, one VPC is set up in one region, and communication between two VPCs in the same region, or between VPCs in different regions, needs to be set up in each VPC to set up a communication gateway, and the interconnection between VPCs is realized through the communication gateway.
[0315] As an example of a hardware functional unit, the first receiving module can include at least one computing device, such as a server, etc. Alternatively, the first receiving module can also be a device implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), etc. The PLD can be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0316] The multiple computing devices included in the first receiving module can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the first receiving module can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the first receiving module can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0317] It should be noted that, in other embodiments, the first receiving module can be used to execute a video generation method (e.g., ...). Figure 18 The encoding module can be used to perform any step in the method shown, such as Figure 18 The generation module can be used to perform any step in the video generation method shown. Figure 18 In the video generation method shown, any step implemented by the first receiving module, encoding module, and generation module can be specified as needed. The first receiving module, encoding module, and generation module respectively implement the steps as follows: Figure 18 The different steps in the video generation method shown enable the full functionality of the video generation device.
[0318] As an example of a software functional unit, a video generation device may include code running on a computing instance. This computing instance can be at least one of a physical host (computing device), a virtual machine, a container, or other computing devices. Furthermore, the aforementioned computing device may be one or more. For example, the video generation device may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the application can be distributed within the same region or in different regions. The multiple hosts / virtual machines / containers used to run the code can be distributed within the same Availability Zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.
[0319] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same VPC or across multiple VPCs. Typically, a VPC is set up within a single region. Communication between two VPCs within the same region, and between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0320] As an example of the module as a hardware functional unit, the video generation apparatus can include at least one computing device, such as a server or the like. Alternatively, the video generation apparatus can also be a device implemented by an ASIC or a PLD, or the like. The PLD can be a CPLD, an FPGA, a GAL, or any combination thereof.
[0321] The plurality of computing devices included in the video generation apparatus can be distributed in the same region or in different regions. The plurality of computing devices included in the video generation apparatus can be distributed in the same AZ or in different AZs. Similarly, the plurality of computing devices included in the video generation apparatus can be distributed in the same VPC or in multiple VPCs. The plurality of computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0322] The present application also provides a computing device 400. As shown in Figure 20 The computing device 400 includes a bus 402, a processor 404, a memory 406, and a communication interface 408. The processor 404, the memory 406, and the communication interface 408 communicate with each other through the bus 402. The computing device 400 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 400.
[0323] The bus 402 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 3 only one line is used, but it does not mean that there is only one bus or only one type of bus. The bus 402 can include a path for transmitting information between various components (e.g., the memory 406, the processor 404, the communication interface 408) of the computing device 400.
[0324] The processor 404 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), or the like.
[0325] The memory 406 can include volatile memory, such as random access memory (RAM), and non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid-state drive (SSD).
[0326] The executable program code stored in the memory 406 is executed by the processor 404 to implement the functions of the aforementioned first receiving module, encoding module, and generating module, respectively, so as to implement the video generation method as shown in Figure 18 That is, the memory 406 stores instructions for executing the video generation method as shown in Figure 18 That is, the memory 406 stores instructions for executing the video generation method as shown in
[0327] The communication interface 408 uses a transceiving module such as, but not limited to, a network interface card and a transceiver to implement the communication between the computing device 400 and other devices or communication networks.
[0328] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a notebook computer, or a smart phone.
[0329] As shown in Figure 21 The computing device cluster includes at least one computing device 400. The memory 406 in one or more computing devices 400 in the computing device cluster can store the same instructions for executing the video generation method as shown in Figure 18 That is, the memory 406 in one or more computing devices 400 in the computing device cluster can store the same instructions for executing the video generation method as shown in
[0330] In some possible implementations, the memory 406 in one or more computing devices 400 in the computing device cluster can also respectively store partial instructions for executing the video generation method as shown in Figure 18 That is, the combination of one or more computing devices 400 can collectively execute the instructions for executing the video generation method as shown in Figure 18 That is, the combination of one or more computing devices 400 can collectively execute the instructions for executing the video generation method as shown in
[0331] It should be noted that the memories 406 in different computing devices 400 in the computing device cluster can store different instructions for respectively performing part of the functions of the video generation apparatus. That is, the memories 406 in different computing devices 400 store instructions for implementing the functions of one or more of the first receiving module, the encoding module, and the generating module.
[0332] In some possible implementation manners, one or more computing devices in the computing device cluster can be connected through a network. The network can be a wide area network, a local area network, or the like. Figure 22 A possible implementation manner is shown. As shown in Figure 22 The two computing devices 400A and 400B are connected through a network. Specifically, the computing devices are connected to the network through the communication interfaces in the computing devices. In this type of possible implementation manner, the memory 406 in the computing device 400A stores instructions for performing the functions of the first receiving module. Meanwhile, the memory 406 in the computing device 400B stores instructions for performing the functions of the encoding module and the generating module.
[0333] Figure 22 The connection manner between the computing device cluster shown in the figure can be that the XX method provided in the present application needs to (for example, store a large amount of data) and, therefore, the functions implemented by the encoding module and the generating module are performed by the computing device 400B. For example, the function implemented by the encoding module can be encoding at least one control condition into at least one latent vector, and the function implemented by the generating module can be generating a target video in which the first material interacts with the second material along the motion trajectory. The target video includes the first material and the second material, the at least one attribute includes a structure attribute and a motion attribute, the structure attribute includes a material structure of the first material and the second material, and the motion attribute includes the motion trajectory of the first material.
[0334] It should be understood that Figure 22 The functions of the computing device 400A shown in the figure can also be performed by a plurality of computing devices 400. Similarly, the functions of the computing device 400B can also be performed by a plurality of computing devices 400.
[0335] The present application also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to the connection relationship between the computing devices shown in Figure 21 and Figure 22 The connection manner of the computing device cluster. The difference is that the memory 406 in one or more computing devices 400 in the computing device cluster can store the same instructions for performing Figure 10 or Figure 11 the video generation method in the corresponding embodiment.
[0336] In some possible implementation manners, the memory 406 of one or more computing devices 400 in the computing device cluster can also respectively store part of instructions of the video generation method for performing Figure 10 or Figure 11 the video generation method in the corresponding embodiment. In other words, the combination of one or more computing devices 400 can jointly execute instructions for performing the video generation method in the corresponding embodiment. Figure 10 or Figure 11 the video generation method in the corresponding embodiment.
[0337] It should be noted that the memory 406 in different computing devices 400 in the computing device cluster can store different instructions for performing part of the functions of the video generation apparatus. That is, the instructions stored in the memory 406 in different computing devices 400 can implement the functions of the video generation apparatus.
[0338] Embodiments of the present application also provide a computer program product containing instructions. The computer program product can be a software or program product containing instructions, which can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device is caused to perform the video generation method.
[0339] Embodiments of the present application also provide a computer readable storage medium. The computer readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk), etc. The computer readable storage medium includes instructions instructing the computing device to perform the video generation method, or instructing the computing device to perform the video generation method.
[0340] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. A video generation method, characterized in that, The method includes: Receive at least one control condition, the at least one control condition being used to control at least one attribute of the target video to be generated, the at least one attribute including one or more of the following: structural attribute, motion attribute, physical attribute; Encode the at least one control condition into at least one hidden vector; The at least one latent vector is input into a pre-trained video generation model, and the target video that meets the at least one control condition is generated by the video generation model.
2. The method as described in claim 1, characterized in that, The target video includes at least one piece of material. The structural attributes include the material structure and / or camera trajectory in the target video, wherein the material structure includes the location or category of at least one material. The motion attributes include the motion trajectory of at least one element in the target video; The physical properties include the physical phenomena of at least one material in the target video during motion.
3. The method as described in claim 2, characterized in that, The target video includes a first material and a second material. The at least one attribute includes a structural attribute and a motion attribute. The structural attribute includes the material structure of the first material and the second material. The motion attribute includes the motion trajectory of the first material. The target video includes a video in which the first material interacts with the second material along the motion trajectory.
4. The method as described in claim 3, characterized in that, The at least one attribute also includes physical attributes, and the target video includes a video in which the first material interacts with the second material along the motion trajectory using the physical phenomenon.
5. The method according to any one of claims 1-4, characterized in that, The at least one attribute further includes an environmental attribute, which includes environmental parameter information of the target video, and the environmental parameter includes at least one of weather and lighting.
6. The method according to any one of claims 1-5, characterized in that, Receive at least one control condition, including: A first control is displayed, through which control conditions sent by the user are received; wherein, the control conditions indicate control over the structural attribute, and the control conditions include one or more of text and a file with a predetermined format; or, the control conditions indicate control over the motion attribute, and the control conditions include one or more of text and a file with a predetermined format; or, the control conditions indicate control over the physical attribute, and the control conditions include one or more of text and a file with a predetermined format.
7. The method according to any one of claims 1-6, characterized in that, The method further includes: Display a second control; the second control is used to receive prompt information input by the user, the prompt information being used to guide the generation of the target video, the prompt information including text and / or images.
8. The method according to any one of claims 1-7, characterized in that, The video generation model is a diffusion model. The diffusion model includes a cross-attention layer; the cross-attention layer includes at least one first variable; Inputting the at least one latent vector into a pre-trained video generation model includes: The at least one latent vector is input into the cross-attention layer of the pre-trained video generation model, and the value of the at least one latent vector is assigned to the at least one first variable; And / or, The input layer of the diffusion model includes a noise vector corresponding to random noise; at least one second variable is embedded in the noise vector; Inputting the at least one latent vector into a pre-trained video generation model includes: The value of at least one latent vector is assigned to the second variable, and the noise vector embedded with the second variable is input into the pre-trained video generation model.
9. The method as described in claim 8, characterized in that, The noise vector embeds at least one second variable, including: The noise vector is connected to the at least one second variable; or, the at least one second variable serves as the noise vector.
10. The method according to any one of claims 1-9, characterized in that, Before receiving at least one control condition, the method further includes: Obtain labeled training data; wherein the label is at least one first hidden vector obtained by at least one control condition encoding; In one iteration of training the video generation model, the method further includes: Extract at least one control condition from the target video output by the video generation model; Obtain at least one second hidden vector corresponding to at least one control condition extracted from the target video; The value of the loss function is calculated based on the at least one second hidden vector and the at least one first hidden vector; Based on the value of the loss function, the trainable parameters in the video generation model are optimized.
11. The method as described in claim 10, characterized in that, Extract at least one control condition from the target video output by the video generation model, including performing one or more of the following steps: Call the semantic segmentation model and / or the 3D detection model, and extract the control conditions corresponding to the structure from the target video output by the video generation model through the semantic segmentation model and / or the 3D detection model; The 3D tracking model is invoked, and the control conditions corresponding to the motion are extracted from the target video output by the video generation model through the 3D tracking model. The multimodal understanding model is invoked, and the control conditions corresponding to the environment are extracted from the target video output by the video generation model through the multimodal understanding model. The multimodal understanding model is invoked, and the physical corresponding control conditions are extracted from the target video output by the video generation model through the multimodal understanding model.
12. The method as described in claim 10 or 11, characterized in that, Obtain labeled training data, including: Determine the source material and background to be used to generate the sample video; Run the simulation model and import the materials and background into the simulation model; If at least one control condition sent by the user is detected, a sample video conforming to the at least one control condition is generated through the simulation model. The encoder is invoked to encode the at least one control condition, thereby obtaining at least one first hidden vector; The at least one first hidden vector is used as the label of the sample video to obtain labeled training data.
13. The method according to any one of claims 1-12, characterized in that, After generating a target video that meets at least one of the control conditions, the method further includes: The quality feedback module is invoked to obtain the quality score of the target video. The quality score is obtained based on the evaluation results of subjective indicators and / or objective indicators. The evaluation results of the subjective indicators are obtained based on user operations. The objective indicators include the video quality assessment indicator FVD and / or the multimodal model similarity indicator CLIP Similarity. If the quality score is below a threshold, incremental training of the video generation model is triggered.
14. A video generation apparatus, characterized in that, The device includes: A first receiving module is configured to receive at least one control condition, wherein the at least one control condition is configured to control at least one attribute of the target video to be generated, wherein the at least one attribute includes one or more of the following: structural attributes, motion attributes, and physical attributes; An encoding module is used to encode the at least one control condition into at least one hidden vector; A generation module is used to input the at least one latent vector into a pre-trained video generation model, and generate the target video that meets the at least one control condition through the video generation model.
15. The apparatus as claimed in claim 14, characterized in that, The target video includes at least one piece of material. The structural attributes include the material structure and / or camera trajectory in the target video, wherein the material structure includes the location or category of at least one material. The motion attributes include the motion trajectory of at least one element in the target video; The physical properties include the physical phenomena of at least one material in the target video during motion.
16. The apparatus as claimed in claim 15, characterized in that, The target video includes a first material and a second material. The at least one attribute includes a structural attribute and a motion attribute. The structural attribute includes the material structure of the first material and the second material. The motion attribute includes the motion trajectory of the first material. The target video includes a video in which the first material interacts with the second material along the motion trajectory.
17. The apparatus as claimed in claim 16, characterized in that, The at least one attribute also includes physical attributes, and the target video includes a video in which the first material interacts with the second material along the motion trajectory using the physical phenomenon.
18. The apparatus as claimed in any one of claims 14-17, characterized in that, The at least one attribute further includes an environmental attribute, which includes environmental parameter information of the target video, and the environmental parameter includes at least one of weather and lighting.
19. The apparatus as claimed in any one of claims 14-18, characterized in that, When receiving at least one control condition, the first receiving module is specifically used for: A first control is displayed, through which control conditions sent by the user are received; wherein, the control conditions indicate control over the structural attribute, and the control conditions include one or more of text and a file with a predetermined format; or, the control conditions indicate control over the motion attribute, and the control conditions include one or more of text and a file with a predetermined format; or, the control conditions indicate control over the physical attribute, and the control conditions include one or more of text and a file with a predetermined format.
20. The apparatus according to any one of claims 14-19, characterized in that, The device further includes a second receiving module, the second receiving module being used for: Display a second control; the second control is used to receive prompt information input by the user, the prompt information being used to guide the generation of the target video, the prompt information including text and / or images.
21. The apparatus according to any one of claims 14-20, characterized in that, The video generation model is a diffusion model. The diffusion model includes a cross-attention layer; the cross-attention layer includes at least one first variable; When the at least one latent vector is input into a pre-trained video generation model, the generation module is specifically used for: The at least one latent vector is input into the cross-attention layer of the pre-trained video generation model, and the value of the at least one latent vector is assigned to the at least one first variable; And / or, The input layer of the diffusion model includes a noise vector corresponding to random noise; at least one second variable is embedded in the noise vector; When the at least one latent vector is input into a pre-trained video generation model, the generation module is specifically used for: The value of at least one latent vector is assigned to the second variable, and the noise vector embedded with the second variable is input into the pre-trained video generation model.
22. The apparatus as claimed in claim 21, characterized in that, When at least one second variable is embedded in the noise vector, the generation module is specifically used for: The noise vector is connected to the at least one second variable; or, the at least one second variable serves as the noise vector.
23. The apparatus as claimed in any one of claims 14-22, characterized in that, The device further includes a training module, which performs the following steps before receiving at least one control condition: Obtain labeled training data; wherein the label is at least one first hidden vector obtained by at least one control condition encoding; In one iteration of training the video generation model, the training module is further configured to: Extract at least one control condition from the target video output by the video generation model; Obtain at least one second hidden vector corresponding to at least one control condition extracted from the target video; The value of the loss function is calculated based on the at least one second hidden vector and the at least one first hidden vector; Based on the value of the loss function, the trainable parameters in the video generation model are optimized.
24. The apparatus as claimed in claim 23, characterized in that, When extracting at least one control condition from the target video output by the video generation model, the training module is specifically used to perform one or more of the following steps: Call the semantic segmentation model and / or the 3D detection model, and extract the control conditions corresponding to the structure from the target video output by the video generation model through the semantic segmentation model and / or the 3D detection model; The 3D tracking model is invoked, and the control conditions corresponding to the motion are extracted from the target video output by the video generation model through the 3D tracking model. The multimodal understanding model is invoked, and the control conditions corresponding to the environment are extracted from the target video output by the video generation model through the multimodal understanding model. The multimodal understanding model is invoked, and the physical corresponding control conditions are extracted from the target video output by the video generation model through the multimodal understanding model.
25. The apparatus as claimed in claim 23 or 24, characterized in that, When acquiring labeled training data, the training module is specifically used for: Determine the source material and background to be used to generate the sample video; Run the simulation model and import the materials and background into the simulation model; If at least one control condition sent by the user is detected, a sample video conforming to the at least one control condition is generated through the simulation model. The encoder is invoked to encode the at least one control condition, thereby obtaining at least one first hidden vector; The at least one first hidden vector is used as the label of the sample video to obtain labeled training data.
26. The apparatus as claimed in any one of claims 14-25, characterized in that, The device further includes a triggering module and a quality feedback module. The triggering module is configured to perform the following steps after generating a target video that meets at least one control condition: The quality feedback module is invoked to obtain a quality score for the target video. The quality score is obtained based on the evaluation results of subjective and / or objective indicators. The evaluation results of the subjective indicators are obtained based on user operations. The objective indicators include the video quality assessment indicator FVD and / or the multimodal model similarity indicator CLIP Similarity. If the quality score is below a threshold, incremental training of the video generation model is triggered.
27. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1-13.
28. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster performs the method as described in any one of claims 1-13.
29. A computer-readable storage medium, characterized in that, Includes computer program instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1-13.