Video generation method, motion video generation method for virtual object, video editing method, video generation model training method, and video generation model-based information processing method

By using temporal attention and parameter cross-attention units in the video generation model, controllability of video camera movement is achieved, generating realistic and natural target videos, thus solving the problem of complex camera movement control in video generation.

WO2026001197A1PCT designated stage Publication Date: 2026-01-02ALIBABA (CHINA) CO LTD

Patent Information

Application Number
PCT/CN2025/088428
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-25
Filing Date
2025-04-11
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

In existing video generation technologies, the controllability of video camera movement is relatively complex, making it difficult to achieve good camera movement control, which affects the presentation of video effects and the expression of content.

Method used

By employing a video generation model that combines temporal attention units and parameter cross-attention units, and learning camera parameter control by associating virtual camera parameters with the feature dimensions of video frames, a target video that supports a wide range of camera movement effects and is temporally stable is generated.

Benefits of technology

The generated target videos are more realistic, natural, and have a smooth timeline, meeting the high standards required for video content production and film and television production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025088428_02012026_PF_FP_ABST
    Figure CN2025088428_02012026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a video generation method, a motion video generation method for a virtual object, a video editing method, a video generation model training method, and a video generation model-based information processing method. The video generation method comprises: acquiring video generation data, the video generation data comprising a virtual camera parameter, a reference action sequence of a target object, and an object image of the target object (202); inputting the video generation data into a video generation model to obtain a target video of the target object, the video generation model being obtained by performing parameter adjustment on a parameter encoding unit, a parameter cross-attention unit, and a temporal attention unit in a pre-trained generation model on the basis of a sample video, and a sample object image, sample action sequence and sample camera parameter corresponding to the sample video, and the pre-trained generation model being obtained by performing parameter adjustment on an object encoding unit, an action encoding unit, and a generation unit in an initial generation model, on the basis of the sample video (204). The virtual camera parameter is inputted into a model comprising the temporal attention unit and the parameter cross-attention unit, so that a video that is temporally stable and conforms to a camera motion trajectory is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Video generation, motion video generation of virtual object, video editing, video generation model training and information processing method based on video generation model

[0001] The present disclosure claims priority to Chinese Patent Application No. 202410833166.5, filed on June 25, 2024, with the Chinese Patent Office, entitled “Video generation, motion video generation of virtual object, video editing, video generation model training and information processing method based on video generation model”, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] Embodiments of the present disclosure relate to the field of computer technology, in particular to a video generation method, a motion video generation method of a virtual object, a video editing method, a video generation model training method, and an information processing method based on a video generation model. BACKGROUND

[0003] With the development of computer technology, video generation has gradually become an important research content of multi-modal and vision in generative artificial intelligence (AIGC, AI Generated Content), and has broad prospects and application value in digital human, e-commerce content, film and television production and other scenarios.

[0004] In the video generation scenario, in addition to keeping the role of the character consistent and the action controllable, it is also desired to further achieve controllability of video panning. Good panning is crucial to the presentation of video effects and the expression of content. However, since the controllability of video panning is relatively complex, it is a difficult problem to solve. Therefore, there is an urgent need for a video generation scheme with controllable video panning. SUMMARY

[0005] Therefore, the embodiments of the present disclosure provide a video generation method. One or more embodiments of the present disclosure also relate to a motion video generation method of a virtual object, a video editing method, a video generation model training method, an information processing method based on a video generation model, a task platform, a video generation device, a motion video generation device of a virtual object, a video editing device, a video generation model training device, an information processing device based on a video generation model, a computing device, a computer-readable storage medium, and a computer program product, to solve the technical defects in the prior art.

[0006] According to a first aspect of embodiments of the present disclosure, a video generation method is provided, including: obtaining video generation data, wherein the video generation data includes virtual camera parameters, a reference action sequence of a target object, and an object image of the target object; inputting the video generation data into a video generation model to obtain a target video of the target object, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit, and a timing attention unit in a pre-trained generation model based on a sample video and a sample object image corresponding to the sample video, a sample action sequence, and sample camera parameters, and the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit, and a generation unit in an initial generation model based on the sample video.

[0007] According to a second aspect of embodiments of the present disclosure, a motion video generation method of a virtual object is provided, including: obtaining motion video generation data of the virtual object, wherein the motion video generation data includes virtual camera parameters, a reference action sequence of the virtual object, and a virtual object image; inputting the motion video generation data into a video generation model to obtain a target motion video of the virtual object, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit, and a timing attention unit in a pre-trained generation model based on a sample video and a sample object image corresponding to the sample video, a sample action sequence, and sample camera parameters, and the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit, and a generation unit in an initial generation model based on the sample video.

[0008] According to a third aspect of embodiments of the present disclosure, a video editing method is provided, including: obtaining an original video and target editing data, wherein the target editing data includes at least one of a target object image, target camera parameters, and a target action sequence; analyzing the original video based on the target editing data to determine original video data in the original video; inputting the target editing data and the original video data into a video generation model to obtain a target editing video, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit, and a timing attention unit in a pre-trained generation model based on a sample video and a sample object image corresponding to the sample video, a sample action sequence, and sample camera parameters, and the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit, and a generation unit in an initial generation model based on the sample video.

[0009] According to a fourth aspect of the embodiments of the present disclosure, a video generation model training method is provided, including: obtaining a sample video and sample object images corresponding to the sample video, a sample action sequence, and sample camera parameters; inputting the sample object images, the sample action sequence, and the sample camera parameters into a pre-trained generation model to obtain first predicted features output by a generation unit in the pre-trained generation model, wherein the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit, and a generation unit in an initial generation model based on a sample video; inputting the sample video into an image encoding unit in the pre-trained generation model to obtain first sample features; adjusting unit parameters of a parameter encoding unit, a time sequence attention unit, and a parameter cross attention unit in the pre-trained generation model based on the first predicted features and the first sample features to obtain a trained video generation model.

[0010] According to a fifth aspect of the embodiments of the present disclosure, an information processing method based on a video generation model is provided, including: receiving a task generation request, wherein the task generation request includes request information; obtaining a video generation model based on the request information, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit, and a time sequence attention unit in a pre-trained generation model based on sample videos and sample object images corresponding to the sample videos, a sample action sequence, and sample camera parameters, and the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit, and a generation unit in an initial generation model based on a sample video; generating task information based on the video generation model, wherein the task information is used to perform a target video task.

[0011] According to a sixth aspect of the embodiments of the present disclosure, a task platform is provided, including a request interface and a response unit; the request interface is configured to receive a task generation request, wherein the task generation request includes request information; the response unit is configured to obtain a video generation model based on the request information, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit, and a time sequence attention unit in a pre-trained generation model based on sample videos and sample object images corresponding to the sample videos, a sample action sequence, and sample camera parameters, and the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit, and a generation unit in an initial generation model based on a sample video; and the response unit is configured to generate task information based on the video generation model, wherein the task information is used to perform a target video task.

[0012] According to a seventh aspect of embodiments of the present disclosure, a video generation apparatus is provided, comprising: a first obtaining module configured to obtain video generation data, wherein the video generation data comprises virtual camera parameters, a reference action sequence of a target object, and an object image of the target object; a first input module configured to input the video generation data into a video generation model to obtain a target video of the target object, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit, and a timing attention unit in a pre-trained generation model based on a sample video and a sample object image corresponding to the sample video, a sample action sequence, and sample camera parameters, and the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit, and a generation unit in an initial generation model based on the sample video.

[0013] According to an eighth aspect of embodiments of the present disclosure, a motion video generation apparatus of a virtual object is provided, comprising: a second obtaining module configured to obtain motion video generation data of a virtual object, wherein the motion video generation data comprises virtual camera parameters, a reference action sequence of the virtual object, and a virtual object image; a second input module configured to input the motion video generation data into a video generation model to obtain a target motion video of the virtual object, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit, and a timing attention unit in a pre-trained generation model based on a sample video and a sample object image corresponding to the sample video, a sample action sequence, and sample camera parameters, and the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit, and a generation unit in an initial generation model based on the sample video.

[0014] According to a ninth aspect of embodiments of the present disclosure, a video editing apparatus is provided, comprising: a third obtaining module configured to obtain an original video and target editing data, wherein the target editing data comprises at least one of a target object image, target camera parameters, and a target action sequence; an analysis module configured to analyze the original video based on the target editing data to determine original video data in the original video; a third input module configured to input the target editing data and the original video data into a video generation model to obtain a target edited video, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit, and a timing attention unit in a pre-trained generation model based on a sample video and a sample object image corresponding to the sample video, a sample action sequence, and sample camera parameters, and the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit, and a generation unit in an initial generation model based on the sample video.

[0015] According to a tenth aspect of the embodiments of the present disclosure, a video generation model training apparatus is provided, including: a fourth acquisition module configured to acquire a sample video and sample object images, sample action sequences and sample camera parameters corresponding to the sample video; a fourth input module configured to input the sample object images, the sample action sequences and the sample camera parameters into a pre-trained generation model to obtain first predicted features output by a generation unit in the pre-trained generation model, wherein the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit and a generation unit in an initial generation model based on the sample video; a fifth input module configured to input the sample video into an image encoding unit in the pre-trained generation model to obtain first sample features; and an adjustment module configured to adjust unit parameters of a parameter encoding unit, a temporal attention unit and a parameter cross attention unit in the pre-trained generation model based on the first predicted features and the first sample features to obtain a trained video generation model.

[0016] According to an eleventh aspect of the embodiments of the present disclosure, an information processing apparatus based on a video generation model is provided, including: a receiving module configured to receive a task generation request, wherein the task generation request includes request information; a fifth acquisition module configured to acquire a video generation model based on the request information, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit and a temporal attention unit in a pre-trained generation model based on sample videos and sample object images, sample action sequences and sample camera parameters corresponding to the sample videos, and the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit and a generation unit in an initial generation model based on the sample videos; and a generation module configured to generate task information based on the video generation model, wherein the task information is used to perform a target video task.

[0017] According to a twelfth aspect of the embodiments of the present disclosure, a computing device is provided, including: a memory and a processor; the memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which realize the steps of the method provided in the first aspect or the second aspect or the third aspect or the fourth aspect or the fifth aspect when executed by the processor.

[0018] According to a thirteenth aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, which stores computer programs / instructions, which realize the steps of the method provided in the first aspect or the second aspect or the third aspect or the fourth aspect or the fifth aspect when executed by a processor.

[0019] According to a fourteenth aspect of an embodiment of the present disclosure, a computer program product is provided, including computer programs / instructions that, when executed by a processor, implement the steps of the method provided in the first aspect or the second aspect or the third aspect or the fourth aspect or the fifth aspect.

[0020] The video generation method provided by one embodiment of the present disclosure includes: obtaining video generation data, wherein the video generation data includes virtual camera parameters, a reference action sequence of a target object, and an object image of the target object; inputting the video generation data into a video generation model to obtain a target video of the target object, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit, and a timing attention unit in a pre-trained generation model based on a sample video and a sample object image corresponding to the sample video, a sample action sequence, and sample camera parameters, and the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit, and a generation unit in an initial generation model based on the sample video. By inputting the virtual camera parameters into the video generation model, the camera control of the target video is realized. In addition, because the video generation model includes the timing attention unit and the parameter cross attention unit, the video generation model can learn the timing information and the camera parameters, thereby generating a target video that supports a large range of camera effects and is stable in timing, so that the target video is more realistic and natural. BRIEF DESCRIPTION OF DRAWINGS

[0021] FIG. 1 is an architecture diagram of a video generation system according to one embodiment of the present disclosure;

[0022] FIG. 2 is a flowchart of a video generation method according to one embodiment of the present disclosure;

[0023] FIG. 3 is a flowchart of a processing process of a video generation model training method according to one embodiment of the present disclosure;

[0024] FIG. 4 is a flowchart of a processing process of a pre-trained generation model training method according to one embodiment of the present disclosure;

[0025] FIG. 5 is a flowchart of a motion video generation method of a virtual object according to one embodiment of the present disclosure;

[0026] FIG. 6 is a flowchart of a video editing method according to one embodiment of the present disclosure;

[0027] FIG. 7 is a flowchart of a video generation model training method according to one embodiment of the present disclosure;

[0028] FIG. 8 is a flowchart of an information processing method based on a video generation model according to one embodiment of the present disclosure;

[0029] FIG. 9 is a structural schematic diagram of a task platform according to an embodiment of the present disclosure;

[0030] FIG. 10 is a structural schematic diagram of a video generation apparatus according to an embodiment of the present disclosure;

[0031] FIG. 11 is a structural schematic diagram of a motion video generation apparatus of a virtual object according to an embodiment of the present disclosure;

[0032] FIG. 12 is a structural schematic diagram of a video editing apparatus according to an embodiment of the present disclosure;

[0033] FIG. 13 is a structural schematic diagram of a video generation model training apparatus according to an embodiment of the present disclosure;

[0034] FIG. 14 is a structural schematic diagram of an information processing apparatus based on a video generation model according to an embodiment of the present disclosure;

[0035] FIG. 15 is a structural block diagram of a computing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0036] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, the present disclosure can be practiced without the specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to obscure the present disclosure. Some portions of the detailed description are presented in terms of algorithms, symbolic representations of operations on data bits or binary digital signals stored within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art.

[0037] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of one or more embodiments of the present disclosure. As used in one or more embodiments of the present disclosure and the accompanying claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in one or more embodiments of the present disclosure, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0038] It will be understood that, although the terms first, second, etc. can be used herein to describe various information, these terms are not intended to denote a temporal or chronological order. Rather, these terms are used only to distinguish one from another. For example, without departing from the scope of one or more embodiments of the present disclosure, a first can be termed a second, and, similarly, a second can be termed a first. The term "if' can be interpreted to mean "when" or "upon" or "in response to determining" depending on the context.

[0039] In addition, it should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or rejection.

[0040] In one or more embodiments of the present disclosure, a large model refers to a deep learning model with a large number of model parameters, usually containing hundreds of millions, tens of billions, hundreds of billions, thousands of billions or even tens of billions of model parameters. The large model can also be called a foundation model. Through large-scale unlabeled corpus pre-training, a pre-trained model with hundreds of millions of parameters is output. Such a model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLM, Large Language Model), multi-modal pre-training models, etc.

[0041] In actual application, the large model only needs a small amount of samples to fine-tune the pre-trained model and can be applied to different tasks. The large model can be widely applied to natural language processing (NLP, Natural Language Processing) and computer vision fields. Specifically, it can be applied to computer vision field tasks such as visual question answering (VQA, Visual Question Answering), image captioning (IC, Image Caption), image generation, and natural language processing field tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of the large model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.

[0042] First, the nomenclature involved in one or more embodiments of the present disclosure is explained.

[0043] Cross Attention: A type of attention mechanism used to handle dependencies between two different data sources or sequences, such as injecting data from different modalities into an image / video generation network. Cross Attention allows the model to selectively focus on relevant parts of the input data when generating output, thereby more effectively utilizing contextual information.

[0044] Deep Self-Attention (Transformer) Model: A network structure based on multi-head self-attention mechanism module, mainly used for processing sequence data. The Transformer model includes a stack of repeated encoding units (Encoder) and decoding units (Decoder). This design allows the Transformer to efficiently learn long-term dependencies, making it suitable for a variety of natural language processing tasks, including machine translation, text summarization, and question answering systems.

[0045] Stable Diffusion Model: A generative model based on generative adversarial networks, which gradually increases the noise disturbance to make the generated images gradually clear and realistic.

[0046] Low-Rank Adaptation (Lora): A low-rank adaptive expression method in large models, which can use smaller parameter quantities to increase the additional parameters and functions of large models.

[0047] Text-to-Video (T2V): Automatically converting written text content into video content containing visual and auditory elements.

[0048] Convolutional Neural Network (CNN): A deep learning model suitable for processing data with grid structure, such as images, videos and speech signals.

[0049] Camera Trajectory: The path of the camera's position and orientation changes in three-dimensional space relative to time. The camera trajectory contains a series of camera poses that change over time, each pose usually consists of camera parameters such as position (translation) and rotation (rotation), which can be used to describe the camera's movement and viewing direction changes in the scene.

[0050] In recent years, thanks to the model framework of Transformer, Stable Diffusion, etc., the rapid development and sedimentation of GPU (Graphics Processing Unit) hardware and big data content, AIGC has made great progress in generation effect. For example, a text-to-video model can generate corresponding videos based on text input, and can achieve relatively long-time, stable, and realistic results. However, the current text-to-video model has poor controllability, and many text-to-video models often need to run multiple times to get a relatively satisfactory result. Ensuring the effect and strong controllability are the problems that need to be solved in the actual landing of the generation model. Therefore, controllable video generation has gradually become a research focus, especially for downstream content production, film and television production, and other high-standard scenarios. The controllability of video content is a very important hard requirement. Among various controllable input conditions, the camera motion trajectory is a very important parameter. In the video production process, the motion control of the camera is a very important part, and good camera operation is crucial to the presentation of video effects and the expression of content.

[0051] Based on this, the embodiment of the disclosure provides a video generation scheme using a video generation model. The video generation model includes a time sequence attention unit and a parameter cross attention unit. Through the parameter cross attention unit, the model can associate the virtual camera parameters and the video frames in the feature dimension, learn the control and generation of the camera parameters, and support a large range of camera operation effects. Through the time sequence attention unit, the video generation model can learn the time sequence information, so that the generated video time sequences smoothly transition, and are more realistic and natural.

[0052] In the disclosure, a video generation method is provided. The disclosure also relates to a motion video generation method of a virtual object, a video editing method, a video generation model training method, an information processing method based on a video generation model, a task platform, a video generation device, a motion video generation device of a virtual object, a video editing device, a video generation model training device, an information processing device based on a video generation model, a computing device, a computer-readable storage medium, and a computer program product. Each of the embodiments is described in detail below.

[0053] Referring to FIG. 1, FIG. 1 shows an architecture diagram of a video generation system according to an embodiment of the disclosure. The video generation system can include a client 100 and a server 200.

[0054] The client 100 is configured to send video generation data to the server 200, wherein the video generation data includes virtual camera parameters, a reference action sequence of a target object, and an object image of the target object.

[0055] The server 200 is configured to input the video generation data into a video generation model to obtain a target video of a target object, where the video generation model is obtained by adjusting parameters of parameter encoding units, parameter cross attention units and time sequence attention units in a pre-trained generation model based on sample videos and sample object images, sample action sequences and sample camera parameters corresponding to the sample videos, and the pre-trained generation model is obtained by adjusting parameters of object encoding units, action encoding units and generation units in an initial generation model based on the sample videos; and the server 200 is configured to send the target video to the client 100.

[0056] The client 100 is further configured to receive the target video sent by the server 200.

[0057] By inputting the virtual camera parameters into the video generation model, the application of the scheme of the embodiments of the present disclosure realizes the control of the camera operation of the target video, and since the video generation model includes the time sequence attention units and the parameter cross attention units, the video generation model can learn the time sequence information and the camera parameters, thereby generating the target video that supports a large range of camera operation effects and is stable in time sequence, so that the target video is more real and natural.

[0058] In actual application, the video generation system can include a plurality of clients 100. The plurality of clients 100 can establish a communication connection through the server 200, and in the video generation scenario, the server 200 is used to provide video generation services between the plurality of clients 100, and the plurality of clients 100 can be respectively used as a sending end or a receiving end to realize communication through the server 200. A user can interact with the server 200 through the client 100 to receive data sent by other clients 100 or send data to other clients 100, and the like. In the video generation scenario, the user can publish a data stream to the server 200 through the client 100, the server 200 generates a target video according to the data stream, and pushes the target video to other clients that establish a communication. The client 100 and the server 200 establish a connection through a network. The network provides a medium for a communication link between the client 100 and the server 200. The network can include various connection types, such as wired, wireless communication links or optical fiber cables, and the like. The data transmitted by the client 100 can need to be processed through encoding, transcoding, compression and the like before being published to the server 200. It should be noted that the video generation method provided in the embodiments of the present disclosure is generally executed by the server, but in other embodiments of the present disclosure, the client can also have similar functions as the server, so as to execute the video generation method provided in the embodiments of the present disclosure. In other embodiments, the video generation method provided in the embodiments of the present disclosure can also be executed by the client and the server together. Next, the video generation method proposed in the embodiments of the present disclosure is taken as an example to be executed by the server, and the video generation method is described in detail.

[0059] Referring to FIG. 2, FIG. 2 shows a flowchart of a video generation method according to an embodiment of the present disclosure, which specifically includes the following steps:

[0060] Step 202: Obtain video generation data, wherein the video generation data includes virtual camera parameters, a reference action sequence of a target object, and an object image of the target object.

[0061] In one or more embodiments of the present disclosure, when generating a video, video generation data can be obtained, and a target video meeting requirements can be generated based on the video generation data.

[0062] Specifically, the video generation data refers to a data set used to generate a new target video. The video generation data can be data of different video tasks, including but not limited to a video generation task, a video panning adjustment task, a video object replacement task, and a video action replacement task. The video task can be a task in a motion video scene, a task in a film and television video editing scene, a task in a multimedia content production scene, and the like. The target object refers to an object appearing in the target video, which can be a person, an animal, a virtual character, and the like. The object image refers to an image of the target object, which can also be understood as a video foreground image. For example, the object image can be an image of a person standing. The number of object images can be one or multiple, for example, the object images include a front image and a side image of the target object. The reference action sequence includes multiple reference action images, each of which contains multiple action keypoints. Through the reference action sequence, the action or posture (such as standing or lying) to be demonstrated by the target object in the target video can be defined. The virtual camera parameter can also be referred to as a virtual camera trajectory parameter. The virtual camera parameter is a series of data used to describe the motion state of the virtual camera itself when the virtual camera is moving or shooting a dynamic scene. Through the virtual camera parameter, the motion state of the virtual camera itself can be controlled, thereby controlling the panning trajectory of the generated video. Specifically, the virtual camera parameter includes the camera extrinsic parameter at each time t in the entire T time period, such as the position and posture parameters of the virtual camera in the world coordinate system, i.e., the rotation matrix and the translation vector. The rotation matrix is used to describe the rotation direction and angle of the camera coordinate system relative to the world coordinate system. It maintains the distance and direction relationship in space, ensuring that no scale distortion occurs during rotation. The translation vector represents the position of the origin of the camera coordinate system in the world coordinate system. Through the translation vector and the rotation matrix, any point in the world coordinate system can be converted to the camera coordinate system through translation and rotation.

[0063] In actual applications, there are various ways to obtain the video generation data, which are selected according to actual conditions, and the embodiments of the present disclosure do not make any limitation on this. In the first possible implementation manner of the present disclosure, the video generation data sent by the user through the client can be received. In the second possible implementation manner of the present disclosure, since it is difficult for the user to directly construct the reference action sequence and the virtual camera parameter, the user can be provided with multiple candidate action sequences and candidate camera parameters, the reference action sequence selected by the user from the multiple candidate action sequences is determined, and the virtual camera parameter selected by the user from the multiple candidate camera parameters is determined. In the third possible implementation manner of the present disclosure, the user can send the reference video and the object image of the target object through the client, and the server parses the reference action sequence and / or the virtual camera parameter from the reference video.

[0064] In an optional embodiment of the present disclosure, the above obtaining the video generation data can include the following steps:

[0065] receiving the reference video and the object image of the target object sent by the client;

[0066] inputting the reference video into the action extraction model to obtain the reference action sequence;

[0067] inputting the reference video into the parameter extraction model to obtain the virtual camera parameter.

[0068] Specifically, the action extraction model is used to extract the action sequence in the input video. The action extraction model can be a deep learning model trained based on a CNN model. The parameter extraction model is used to extract the camera parameter in the input video. The action extraction model and the parameter extraction model can be deep learning models trained based on a CNN model. For example, the action extraction model can be trained based on a sample video and a sample action sequence carried by the sample video. The parameter extraction model can be trained based on a sample video and a sample camera parameter carried by the sample video.

[0069] It should be noted that the reference action sequence and / or the virtual camera parameter can be parsed from the reference video, and therefore, the target object can be included in the reference video or can not be included in the reference video. For example, the target object is object A, the reference video can be a dance video of object A, the reference video can also be a dance video of object B, and the reference video can also be a landscape video without any object.

[0070] By applying the scheme of the embodiments of the present disclosure, the user only needs to provide the reference video and the object image of the target object, the server determines the reference action sequence and the virtual camera parameter based on the reference video, reduces the difficulty of data provision of the user, and reduces the threshold of video generation.

[0071] In step 204, the video generation data is input into the video generation model to obtain a target video of the target object, where the video generation model is obtained by adjusting parameters of a parameter coding unit, a parameter cross attention unit and a time sequence attention unit in a pre-trained generation model based on a sample video and sample object images, sample action sequences and sample camera parameters corresponding to the sample video, and the pre-trained generation model is obtained by adjusting parameters of an object coding unit, an action coding unit and a generation unit in an initial generation model based on the sample video.

[0072] In one or more embodiments of the present disclosure, after obtaining the video generation data, the reference action sequence of the target object, the virtual camera parameters and the object images can be further input into the video generation model to obtain the target video of the target object.

[0073] Specifically, the target video refers to a video in which the action of the target object conforms to the reference action sequence and the video panning effect conforms to the virtual camera parameters. The training methods of the video generation model and the pre-trained generation model are both supervised training. The video generation model includes an image coding unit, an object coding unit, an action coding unit, a parameter coding unit, a generation unit and a decoding unit, and the video generation model is obtained by adjusting parameters of the parameter coding unit, the parameter cross attention unit and the time sequence attention unit based on a sample video and sample object images, sample action sequences and sample camera parameters corresponding to the sample video.

[0074] In an optional embodiment of the present disclosure, the structure of the video generation model is described. That is, the video generation model includes an image coding unit, an object coding unit, an action coding unit, a parameter coding unit, a generation unit and a decoding unit; and the above step of inputting the video generation data into the video generation model to obtain the target video of the target object can include the following steps:

[0075] The object images are input into the image coding unit to obtain image coding features;

[0076] The image coding features are input into the object coding unit to obtain object coding features;

[0077] The reference action sequence is input into the action coding unit to obtain an action coding sequence;

[0078] The virtual camera parameters are input into the parameter coding unit to obtain parameter coding features;

[0079] The object coding features, the action coding sequence and the parameter coding features are input into the generation unit to obtain video coding features;

[0080] The video coding features are input into the decoding unit to obtain the target video of the target object.

[0081] It should be noted that the image encoding unit is connected with the object encoding unit, the object encoding unit is connected with the generation unit, the action encoding unit is connected with the generation unit, the parameter encoding unit is connected with the generation unit, and the generation unit is connected with the decoding unit. The image encoding unit, the action encoding unit and the parameter encoding unit are composed of CNN network layers. The object encoding unit and the generation unit are composed of Transformer network modules. The generation unit can be constructed based on the UNet network of Stable Diffusion or based on Video Stable Diffusion. The image encoding unit is used for feature extraction of the object image to obtain image encoding features. The object encoding unit is used for feature processing of the image encoding features to obtain object encoding features. The action encoding unit is used for feature extraction of the reference action sequence to obtain an action encoding sequence. The parameter encoding unit is used for feature extraction of the virtual camera parameters to obtain parameter encoding features.

[0082] By applying the scheme of the embodiment of the present disclosure, an object encoding unit is connected after the image encoding unit to control the video encoding features based on objects, so that the objects in each frame of the generated target video are consistent with the target objects.

[0083] In an optional embodiment of the present disclosure, the structure of the object encoding unit is described. That is, the object encoding unit includes a plurality of encoding blocks; the step of inputting the image encoding features into the object encoding unit to obtain the object encoding features can include the following steps:

[0084] The image encoding features are input into the first encoding block to obtain first encoding features output by the first encoding block, wherein the first encoding block is the first encoding block in the plurality of encoding blocks;

[0085] The encoding features output by the previous encoding block of the second encoding block are input into the second encoding block to obtain second encoding features output by the second encoding block, wherein the second encoding block is any encoding block in the plurality of encoding blocks except the first encoding block;

[0086] The object encoding features are determined according to the encoding features output by the plurality of encoding blocks respectively.

[0087] Specifically, the encoding block includes a self-attention unit and a cross-attention unit. When the image encoding features are input into the first encoding block to obtain first encoding features output by the first encoding block, the image encoding features can be input into the self-attention unit in the encoding block to obtain image self-attention features, and the image self-attention features are input into the cross-attention unit in the encoding block to obtain the first encoding features.

[0088] Exemplarily, it is assumed that the object coding unit includes three coding blocks connected in sequence, which are coding block 1, coding block 2 and coding block 3 respectively. The process of inputting the image coding features into the object coding unit can include: first, inputting the image coding features into the coding block 1 to obtain coding features 1; then, inputting the coding features 1 into the coding block 2 to obtain coding features 2; finally, inputting the coding features 2 into the coding block 3 to obtain coding features 3. The coding features 1, the coding features 2 and the coding features 3 jointly constitute the object coding features.

[0089] By applying the scheme of the embodiment of the present disclosure, after the image coding unit, an object coding unit is connected to perform object control on the video coding features, so as to ensure that the object in each frame of the generated target video is consistent with the target object.

[0090] In an optional embodiment of the present disclosure, the structure of the generation unit is described. That is, the generation unit includes a plurality of generation blocks, the object coding unit includes a plurality of coding blocks, and the generation blocks and the coding blocks correspond one by one; the above-mentioned inputting the object coding features, the action coding sequence and the parameter coding features into the generation unit to obtain the video coding features can include the following steps:

[0091] According to the correspondence between the generation blocks and the coding blocks, the coding features corresponding to the generation blocks are determined;

[0092] Inputting the action coding sequence, the parameter coding features and the coding features corresponding to the first generation block into the first generation block to obtain the first generation features output by the first generation block, wherein the first generation block is the first generation block in the plurality of generation blocks;

[0093] Inputting the generation features output by the previous generation block of the second generation block, the coding features corresponding to the second generation block and the parameter coding features into the second generation block to obtain the second generation features output by the second generation block, wherein the second generation block is any generation block in the plurality of generation blocks except the first generation block;

[0094] In the case where the second generation block is the last generation block in the plurality of generation blocks, the second generation features output by the second generation block are determined as the video coding features.

[0095] It should be noted that by using the object coding unit to extract the feature expression of the object image of different layers and using the cross-attention mechanism in the same layer of the generation unit to fuse the object features into the generation unit, the object in each frame of the generated target video can be consistent with the target object.

[0096] Exemplarily, it is assumed that the object encoding unit includes three encoding blocks connected in sequence, which are encoding block 1, encoding block 2 and encoding block 3 respectively. The generation unit includes three generation blocks connected in sequence, which are generation block 1, generation block 2 and generation block 3 respectively. When the image encoding features are input into the object encoding unit, the encoding feature 1 output by the encoding block 1, the encoding feature 2 output by the encoding block 2 and the encoding feature 3 output by the encoding block 3 can be obtained. After the encoding block 1 outputs the encoding feature 1, the cross-attention mechanism can be used to fuse the encoding feature 1 into the input of the generation block 1. For the generation block 1, the action encoding sequence, the parameter encoding feature and the encoding feature 1 are input into the generation block 1, and the generation block 1 outputs the generation feature 1. After the encoding block 2 outputs the encoding feature 2, the cross-attention mechanism can be used to fuse the encoding feature 2 into the input of the generation block 2. For the generation block 2, the generation feature 1, the parameter encoding feature and the encoding feature 2 are input into the generation block 2, and the generation block 2 outputs the generation feature 2. After the encoding block 3 outputs the encoding feature 3, the cross-attention mechanism can be used to fuse the encoding feature 3 into the input of the generation block 3. For the generation block 3, the generation feature 2, the parameter encoding feature and the encoding feature 3 are input into the generation block 3, and the generation block 3 outputs the generation feature 3. Since the generation block 3 is the last generation block, the generation feature 3 is determined as the video encoding feature.

[0097] By using the scheme of the embodiment of the present disclosure, the generation process of each generation block is constrained by the encoding features output by the corresponding encoding block by using the cross-attention mechanism, so that the generated video has object consistency.

[0098] In an optional embodiment of the present disclosure, the structure of the generation block in the generation unit is described. Each generation block includes a self-attention unit, a cross-attention unit, a temporal attention unit and a parameter cross-attention unit. Taking the second generation block as an example, the second generation block includes a self-attention unit, a cross-attention unit, a temporal attention unit and a parameter cross-attention unit; the above-mentioned input of the generation feature output by the previous generation block of the second generation block, the encoding feature corresponding to the second generation block and the parameter encoding feature into the second generation block to obtain the second generation feature output by the second generation block can include the following steps:

[0099] The generation feature output by the previous generation block of the second generation block is input into the self-attention unit to obtain a self-attention feature;

[0100] The self-attention feature and the encoding feature corresponding to the second generation block are input into the cross-attention unit to obtain a cross-attention feature;

[0101] The cross-attention feature is input into the temporal attention unit to obtain a temporal attention feature;

[0102] The time sequence attention feature and the parameter encoding feature are input into the parameter cross attention unit to obtain a second generation feature.

[0103] Specifically, the self-attention unit is connected with the cross-attention unit, the cross-attention unit is connected with the time sequence attention unit, and the time sequence attention unit is connected with the parameter cross-attention unit. For example, in the second generation block, the time sequence attention unit is after the cross-attention unit, the self-attention unit is before the cross-attention unit, and the parameter cross-attention unit is after the time sequence attention unit. The self-attention unit is used for data processing by using a self-attention mechanism. The self-attention mechanism allows the self-attention unit to consider how each position in the input data (such as different regions of an image or different parts of a feature map) depends on all other positions. The self-attention mechanism can help the model identify and utilize the correlation between distant pixels. The cross-attention unit is used for data processing by using a cross-attention mechanism. The cross-attention mechanism can help the model guide the generation process of the video according to the conditional information (such as the parameter encoding feature and the action feature sequence). The cross-attention mechanism can enable the model to align the conditional information with the image features, ensuring that the generated image meets the given conditions. The time sequence attention unit is used for data processing by using an attention mechanism in the time dimension, so that the time sequences of the features are related to each other, and the generated video result is more stable. The purpose of the parameter cross-attention unit is to enable the model to learn the camera motion trajectory. Given the virtual camera parameters, the parameter encoding unit is used to process the virtual camera parameters to obtain the parameter encoding feature, and then the parameter encoding feature is injected into the generation unit by the cross-attention mechanism, so that the video encoding feature can learn the corresponding relationship with the camera parameters, and finally the generated video meets the input camera motion trajectory.

[0104] By applying the scheme of the embodiments of the present disclosure, since the time sequence attention unit and the parameter cross-attention unit are included in the video generation model, the video generation model can learn the time sequence information and the camera parameters, thereby generating a target video that supports a large amplitude of camera movement effects and is stable in time sequence, so that the target video is more realistic and natural.

[0105] In an optional embodiment of the present disclosure, after the video generation data is input into the video generation model to obtain the target video of the target object, the following steps can be further included:

[0106] The target video is sent to the client.

[0107] The result feedback information sent by the client is received, wherein the result feedback information is the information fed back by the client to the target video.

[0108] The model optimization data is constructed according to the result feedback information.

[0109] The video generation model is adjusted in parameters by using the model optimization data.

[0110] Specifically, the result feedback information can be information for feeding back the content, quality and completion degree of the target video, reflecting the real feelings and expectations of the client for the target video, and the result feedback information includes but is not limited to quality evaluation information, the corrected accurate target video, and the optimization field of the model. The model optimization data refers to the accurate optimization sample video used for optimizing the video generation model.

[0111] In actual application, there are various ways to construct the model optimization data according to the result feedback information, which are selected according to actual conditions, and the present disclosure does not make any limitation on this. In a possible implementation manner of the present disclosure, the model optimization data can be automatically constructed according to the result feedback information. In another possible implementation manner of the present disclosure, the optimization prompt information can be generated based on the result feedback information, and the model optimization data sent by the client based on the optimization prompt information is received.

[0112] Exemplarily, when the optimization prompt information is generated based on the result feedback information, in a possible implementation manner of the present disclosure, the preset prompt information can be directly obtained, and the result feedback information is added in the preset prompt information to obtain the optimization prompt information. For example, the preset prompt information is "I am very sorry to bring you inaccurate video. Please point out where is not accurate enough or provide accurate video, and I will correct and optimize my answer as soon as possible to better serve you". The result feedback information is "video dolly is poor", and the optimization prompt information is "I am very sorry to bring you poor video according to your feedback on the video dolly. Please point out where is not accurate enough or provide accurate video, and I will correct and optimize my answer as soon as possible to better serve you". In another possible implementation manner of the present disclosure, the result feedback information can be type-identified to determine the information type of the result feedback information, and the information type is further matched with the prompt type of each prompt information in the prompt information library, and the prompt information with the same prompt type as the information type is determined as the optimization prompt information.

[0113] Further, when the model optimization data is constructed directly according to the result feedback information, if the result feedback information is the corrected accurate target video, the corrected accurate target video can be determined as the model optimization data. If the result feedback information is the optimization field of the model, such as the XXX field, a plurality of sample videos in the XXX field can be obtained, and the plurality of sample videos in the XXX field are determined as the model optimization data.

[0114] By applying the scheme of the present disclosure, the performance of the video generation model is continuously optimized by collecting and utilizing the result feedback information, so as to more accurately meet the actual needs of the client and improve the quality and accuracy of the final target video.

[0115] In an optional embodiment of the present disclosure, the training manner of the video generation model is described, that is, before the video generation data is input into the video generation model to obtain the target video of the target object, the following steps can be further included:

[0116] obtaining a sample video and sample object images, sample action sequences and sample camera parameters corresponding to the sample video;

[0117] inputting the sample object images, sample action sequences and sample camera parameters into the pre-trained generation model to obtain first prediction features output by a generation unit in the pre-trained generation model;

[0118] inputting the sample video into an image encoding unit in the pre-trained generation model to obtain first sample features;

[0119] adjusting unit parameters of a parameter encoding unit, a time sequence attention unit and a parameter cross attention unit in the pre-trained generation model according to the first prediction features and the first sample features to obtain the trained video generation model.

[0120] Specifically, the first sample features are generation targets of the generation unit in the pre-trained generation model, and are used to guide the training process of the pre-trained generation model. When the pre-trained generation model is trained, a time sequence attention unit and a parameter cross attention unit are additionally added in each generation block of the generation unit in the pre-trained generation model, that is, the pre-trained generation model includes an image encoding unit, an object encoding unit, an action encoding unit, a parameter encoding unit, a generation unit and a decoding unit; the object encoding unit includes a plurality of encoding blocks; the encoding block includes a self-attention unit and a cross-attention unit. The generation unit includes a plurality of generation blocks, and each generation block includes a self-attention unit, a cross-attention unit, a time sequence attention unit and a parameter cross attention unit.

[0121] It should be noted that there are various ways to obtain the sample video and the sample object images, sample action sequences and sample camera parameters corresponding to the sample video. In a possible implementation manner of the present disclosure, the sample video and the sample object images, sample action sequences and sample camera parameters corresponding to the sample video can be read from other data acquisition devices or databases. In another possible implementation manner of the present disclosure, the sample video can be obtained, and the sample object images, sample action sequences and sample camera parameters can be determined by analyzing the sample video. There are various ways to obtain the sample video. In a possible implementation manner of the present disclosure, a plurality of sample videos sent by a user through a client can be received. In another possible implementation manner of the present disclosure, a plurality of sample videos can be read from other data acquisition devices or databases.

[0122] Further, after obtaining the sample video, for the sample object image: any video frame including the sample object can be randomly selected from the sample video as the sample object image, or a video frame including the sample object with higher quality can be selected as the sample object image. For the sample action sequence: the sample video can be input into an action extraction model to obtain the sample action sequence; or sample video frame extraction can be performed on the sample video, for each sample video frame, the sample video frame is input into a key point identification model to obtain sample action key points, and a sample action image including object bones is drawn. The sample action image corresponding to each sample video frame can constitute the sample action sequence. For the sample camera parameter: the sample video can be input into a parameter extraction model to obtain the sample camera parameter.

[0123] In actual applications, when adjusting the unit parameters of the parameter encoding unit, the time sequence attention unit and the parameter cross attention unit in the pre-training generation model according to the first predicted feature and the first sample feature, a loss value can be calculated according to the first predicted feature and the first sample feature, and the unit parameters of the parameter encoding unit, the time sequence attention unit and the parameter cross attention unit are adjusted according to the loss value in a low-rank adaptive manner until the training process meets a preset stopping condition, and a trained video generation model is obtained. The functions for calculating the loss value include, but are not limited to, cross-entropy loss function, L1 norm loss function, maximum loss function, mean square error loss function, logarithmic loss function, etc. The specific selection is based on actual conditions, and the embodiments of the present disclosure do not make any limitation on this. The preset stopping condition includes, but is not limited to, the loss value being less than or equal to a preset threshold, and the number of iterations reaching a preset iteration number, wherein the preset threshold and the preset iteration number are selected according to actual conditions, and the embodiments of the present disclosure do not make any limitation on this.

[0124] In a possible implementation of the present disclosure, after calculating the loss value, the loss value is compared with a preset threshold. Specifically, if the loss value is greater than the preset threshold, it indicates that the difference between the first predicted feature and the first sample feature is large, and the pre-training generation model has poor prediction ability for the video feature. At this time, the unit parameters of the parameter encoding unit, the time sequence attention unit and the parameter cross attention unit can be adjusted, and the step of inputting the sample object image, the sample action sequence and the sample camera parameter into the pre-training generation model to obtain the first predicted feature output by the generation unit in the pre-training generation model is returned, and the pre-training generation model is continuously trained until the loss value is less than or equal to the preset threshold, indicating that the difference between the target first predicted feature and the first sample feature is small, the preset stopping condition is reached, and the trained video generation model is obtained.

[0125] In another possible implementation of the present disclosure, in addition to comparing the size relationship between the loss value and the preset threshold, the number of iterations can also be combined to determine whether the current pre-training generation model is trained. Specifically, if the loss value is greater than the preset threshold, the unit parameters of the parameter encoding unit, the temporal attention unit and the parameter cross attention unit are adjusted, and the step of inputting the sample object image, the sample action sequence and the sample camera parameter into the pre-training generation model to obtain the first predicted feature output by the generation unit in the pre-training generation model is returned. The pre-training generation model is continuously trained until the preset number of iterations is reached, and the iteration is stopped to obtain the trained video generation model.

[0126] By applying the scheme of the embodiment of the present disclosure, the temporal attention unit and the parameter cross attention unit are additionally added in each generation block of the generation unit in the pre-training generation model, so that the video generation model obtained by training can learn the temporal information and the camera parameter, thereby generating the target video supporting the large-amplitude dolly effect and the stable time sequence, and making the target video more real and natural.

[0127] Referring to FIG. 3, FIG. 3 shows a processing process flowchart of a video generation model training method provided by an embodiment of the present disclosure, and the video generation model is obtained by training a pre-training generation model. The training stage can be referred to as a video generation training stage. In this stage, the model training target is to generate a video segment with stable time sequence and capable of controlling the camera dolly. Therefore, the temporal attention unit and the parameter cross attention unit are introduced in the pre-training generation model, the temporal dimension between different frames can learn the relevant correlation by increasing the temporal attention unit, and the video generation model can further learn the temporal information, such as the floating hem, and the continuous stable time sequence change feature. By increasing the parameter cross attention unit, the model can learn the information of the camera trajectory, and generate the video consistent with the input camera parameter. That is, the pre-training generation model includes an image encoding unit, an object encoding unit, an action encoding unit, a parameter encoding unit, a generation unit and a decoding unit; the object encoding unit includes a plurality of encoding blocks, and each encoding block includes a self-attention unit and a cross-attention unit. The generation unit includes a plurality of generation blocks, and each generation block includes a self-attention unit, a cross-attention unit, a temporal attention unit and a parameter cross attention unit. In the training process, only the unit parameters of the parameter encoding unit, the temporal attention unit and the parameter cross attention unit are adjusted, and the unit parameters of other units except the parameter encoding unit, the temporal attention unit and the parameter cross attention unit are fixed.

[0128] Next, the processing process of the pre-trained generation model is described. As shown in FIG. 3, a sample video is obtained; the sample video is parsed to obtain a sample object image, a sample action sequence and a sample camera parameter; the sample object image is input into an image encoding unit in the pre-trained generation model to obtain a sample image encoding feature, and the image encoding feature is input into an object encoding unit in the pre-trained generation model to obtain a sample object encoding feature; the sample action sequence is input into an action encoding unit in the pre-trained generation model to obtain a sample action encoding sequence; the sample camera parameter is input into a parameter encoding unit in the pre-trained generation model to obtain a sample parameter encoding feature; the sample object encoding feature, the sample action encoding sequence and the sample parameter encoding feature are input into a generation unit in the pre-trained generation model to obtain a first prediction feature; the first prediction feature is input into a decoding unit in the pre-trained generation model to obtain a predicted video. The sample video is input into the image encoding unit in the pre-trained generation model to obtain a first sample feature; the unit parameters of the parameter encoding unit, the time sequence attention unit and the parameter cross attention unit in the pre-trained generation model are adjusted according to the first prediction feature and the first sample feature, and a trained video generation model is obtained.

[0129] By applying the scheme of the embodiments of the present disclosure, when training the pre-trained generation model, the model can also ensure the generation effect when inputting conditions through the sample object image and the sample action sequence; the model is trained by using the sample object image including the sample object, so that the effect of the model on the object (such as a person) video generation task is greatly improved. The model outputs a video of a consistent character, consistent with the input action sequence, and a specified camera track, so that the entire video effect is real and stable, the time sequence is consistent, and a large-scale panning effect can be supported.

[0130] In an optional embodiment of the present disclosure, the training method of the pre-trained generation model is described, that is, the sample video includes a plurality of sample video frames; before the sample object image, the sample action sequence and the sample camera parameter are input into the pre-trained generation model to obtain the first prediction feature output by the generation unit in the pre-trained generation model, the following steps can also be included:

[0131] For the first sample video frame, a first sample object image and a first sample action image corresponding to the first sample video frame are determined, wherein the first sample video frame is sampled from the plurality of sample video frames;

[0132] The first sample object image and the first sample action image are input into the initial generation model to obtain a second prediction feature output by a generation unit in the initial generation model;

[0133] The first sample video frame is input into an image encoding unit in the initial generation model to obtain a second sample feature;

[0134] According to the second prediction feature and the second sample feature, the unit parameters of the object encoding unit, the action encoding unit and the generation unit in the initial generation model are adjusted to obtain a trained pre-training generation model.

[0135] Specifically, the second sample feature is a generation target of the generation unit in the initial generation model, and is used to guide the training process of the initial generation model. Since the initial generation model is trained based on a single video frame and does not include time sequence information and camera motion trajectory, the initial generation model does not include a parameter encoding unit, and the generation unit of the initial generation model does not include a time sequence attention unit and a parameter cross attention unit, that is, the initial generation model includes an image encoding unit, an object encoding unit, an action encoding unit, a generation unit and a decoding unit; the object encoding unit includes a plurality of encoding blocks; the encoding block includes a self-attention unit and a cross-attention unit. The generation unit includes a plurality of generation blocks, and the generation block includes a self-attention unit and a cross-attention unit.

[0136] It should be noted that when the first sample video frame is sampled from the plurality of sample video frames, in one possible implementation, all sample video frames can be determined as the first sample video frame. In another possible implementation, since the similarity between adjacent frames is high, the first sample video frame can be extracted every preset time length.

[0137] In actual application, the implementation of "according to the second prediction feature and the second sample feature, adjusting the unit parameters of the object encoding unit, the action encoding unit and the generation unit in the initial generation model to obtain a trained pre-training generation model" is the same as the implementation of "according to the first prediction feature and the first sample feature, adjusting the unit parameters of the parameter encoding unit, the time sequence attention unit and the parameter cross attention unit in the pre-training generation model to obtain a trained video generation model", and the embodiments of the present disclosure will not be described again.

[0138] Further, after the unit parameters of the object encoding unit, the action encoding unit and the generation unit in the initial generation model are adjusted according to the second prediction feature and the second sample feature to obtain a trained pre-training generation model, the pre-training generation model can be used to process an image generation task, that is, the object image and the action image of the target object are input into the pre-training generation model to obtain a target image, the target image includes the target object, and the action of the target object is consistent with the action in the action image.

[0139] By applying the scheme of the embodiments of the present disclosure, the unit parameters of the object encoding unit, the action encoding unit and the generation unit in the initial generation model are adjusted according to the second prediction feature and the second sample feature, so that the pre-training generation model can generate an image with consistent object and action, and the generation ability of the pre-training generation model is ensured.

[0140] Referring to FIG. 4, FIG. 4 shows a process flow diagram of a pre-training generation model training method provided by an embodiment of the present disclosure, and the pre-training generation model is obtained by training an initial generation model. This training stage can be referred to as an image generation training stage. In this stage, the model training target is to generate images that are consistent in objects and conform to target actions. The initial generation model can reuse the weights of Stable Diffusion and add an action encoding unit and an object encoding unit on the basis of Stable Diffusion. That is, as shown in FIG. 4, the initial generation model includes an image encoding unit, an object encoding unit, an action encoding unit, a generation unit, and a decoding unit; the object encoding unit includes a plurality of encoding blocks; each encoding block includes a self-attention unit and a cross-attention unit. The generation unit includes a plurality of generation blocks, and each generation block includes a self-attention unit and a cross-attention unit. In the training process, only the unit parameters of the object encoding unit, the action encoding unit, and the generation unit are adjusted, and the unit parameters of the image encoding unit and the decoding unit remain unchanged.

[0141] Next, the processing process of the initial generation model is described. As shown in FIG. 4, a sample video is obtained, wherein the sample video includes a plurality of sample video frames; for a first sample video frame, a first sample object image and a first sample action image corresponding to the first sample video frame are determined, wherein the first sample video frame is sampled from the plurality of sample video frames; the first sample object image and the first sample action image are input into the initial generation model to obtain a first predicted video frame output by the initial generation model. The first sample video frame is input into the image encoding unit in the initial generation model to obtain a second sample feature; and the unit parameters of the object encoding unit, the action encoding unit, and the generation unit in the initial generation model are adjusted according to the second sample feature and a second predicted feature output by the generation unit in the initial generation model, to obtain a pre-training generation model trained.

[0142] The video generation method provided by the present disclosure is further described below by taking the application of the video generation method in the virtual object motion video generation scenario as an example with reference to FIG. 5. FIG. 5 shows a flowchart of a virtual object motion video generation method provided by an embodiment of the present disclosure, and specifically includes the following steps:

[0143] Step 502: Obtain motion video generation data of a virtual object, wherein the motion video generation data includes virtual camera parameters, a reference action sequence of the virtual object, and a virtual object image.

[0144] Step 504: input the motion video generation data into the video generation model to obtain a target motion video of the virtual object, wherein the video generation model is obtained by performing parameter adjustment on a parameter coding unit, a parameter cross attention unit and a timing attention unit in a pre-trained generation model based on a sample video and sample object images corresponding to the sample video, a sample action sequence and sample camera parameters, and the pre-trained generation model is obtained by performing parameter adjustment on an object coding unit, an action coding unit and a generation unit in an initial generation model based on the sample video.

[0145] It should be noted that the implementation manners of steps 502 to 504 are the same as those of steps 202 to 204, and the embodiments of the present disclosure will not be described in detail.

[0146] In the virtual object motion video generation scenario, the virtual object image of the virtual object, the action sequence of a certain motion item and the virtual camera parameters can be input into the video generation model to obtain a target motion video of the virtual object performing the motion item.

[0147] By inputting the virtual camera parameters into the video generation model, the application of the scheme of the embodiments of the present disclosure realizes the camera operation control of the target motion video, and since the video generation model includes the timing attention unit and the parameter cross attention unit, the video generation model can learn the timing information and the camera parameters, thereby generating the target motion video with a large amplitude of camera operation effect and stable timing, making the target motion video more real and natural.

[0148] The video generation method provided by the embodiments of the present disclosure is further described below in combination with FIG. 6, taking the application of the video generation method in the video editing (such as replacement and adjustment) scenario as an example. FIG. 6 is a flowchart of a video editing method according to an embodiment of the present disclosure, which specifically includes the following steps:

[0149] Step 602: obtain an original video and target editing data, wherein the target editing data includes at least one of a target object image, a target camera parameter and a target action sequence.

[0150] Step 604: analyze the original video based on the target editing data to determine original video data in the original video.

[0151] Step 606: input the target editing data and the original video data into a video generation model to obtain a target editing video, wherein the video generation model is obtained by performing parameter adjustment on a parameter coding unit, a parameter cross attention unit and a timing attention unit in a pre-trained generation model based on a sample video and sample object images corresponding to the sample video, a sample action sequence and sample camera parameters, and the pre-trained generation model is obtained by performing parameter adjustment on an object coding unit, an action coding unit and a generation unit in an initial generation model based on the sample video.

[0152] It should be noted that the implementation of steps 602 to 606 is the same as that of steps 202 to 204, and the embodiments of the present disclosure will not be described again.

[0153] In actual application, if the video editing task is a video camera adjustment task, the target editing data is a target camera parameter, the original video is parsed based on the target editing data, and it is determined that the original video data in the original video includes an original action sequence and an original object image. If the video editing task is a video object editing task, the target editing data is a target object image, the original video is parsed based on the target editing data, and it is determined that the original video data in the original video includes an original camera parameter and an original action sequence. If the video editing task is a video action editing task, the target editing data is a target action sequence, the original video is parsed based on the target editing data, and it is determined that the original video data in the original video includes an original camera parameter and an original object image.

[0154] By inputting the virtual camera parameter into the video generation model, the camera control of the target edited video is realized, and because the video generation model includes the time sequence attention unit and the parameter cross attention unit, the video generation model can learn the time sequence information and the camera parameter, thereby generating the target edited video with a large amplitude of camera effect and stable time sequence, so that the target edited video is more real and natural.

[0155] Referring to FIG. 7, FIG. 7 shows a flowchart of a video generation model training method according to an embodiment of the present disclosure, which specifically includes the following steps:

[0156] Step 702: Obtain a sample video and sample object images, sample action sequences and sample camera parameters corresponding to the sample video.

[0157] Step 704: Input the sample object images, sample action sequences and sample camera parameters into the pre-trained generation model to obtain first predicted features output by a generation unit in the pre-trained generation model, wherein the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit and a generation unit in an initial generation model based on the sample video.

[0158] Step 706: Input the sample video into an image encoding unit in the pre-trained generation model to obtain first sample features.

[0159] Step 708: Adjust unit parameters of a parameter encoding unit, a time sequence attention unit and a parameter cross attention unit in the pre-trained generation model according to the first predicted features and the first sample features to obtain a trained video generation model.

[0160] It should be noted that the implementation of steps 702 to 708 is the same as the training manner of the video generation model in the video generation method provided in FIG. 2, and the embodiments of the present disclosure will not be described again.

[0161] By applying the scheme of the embodiments of the present disclosure, when the pre-training generation model is trained, the time sequence attention unit and the parameter cross attention unit are additionally added in each generation block included in the generation unit of the pre-training generation model, so that the video generation model obtained by training can learn the time sequence information and the camera parameter, thereby generating the target video supporting a large amplitude of lens movement effect and stable time sequence, so that the target video is more real and natural.

[0162] Referring to FIG. 8, FIG. 8 shows a flowchart of an information processing method based on a video generation model according to an embodiment of the present disclosure, which specifically includes the following steps:

[0163] Step 802: receiving a task generation request, wherein the task generation request includes request information.

[0164] Specifically, the information processing method based on the video generation model can be applied to a task platform, which can be deployed on an end-side device or a cloud-side device. The task generation request is used to request to generate task information of a target video task. The task generation request usually includes a task type, an expected output format, and request information. The request information refers to the parameters or description information related to the target video task carried in the task generation request. The request information includes but is not limited to a task scene identifier of the target video task, a task model identifier, and a plurality of sample videos corresponding to the target video task.

[0165] Step 804: obtaining a video generation model based on the request information, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit, and a time sequence attention unit in a pre-training generation model based on sample videos, sample object images corresponding to the sample videos, sample action sequences, and sample camera parameters, and the pre-training generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit, and a generation unit in an initial generation model based on the sample videos.

[0166] In actual application, there are various ways to obtain the video generation model based on the request information, which are selected according to actual conditions, and the embodiments of the present disclosure do not make any limitation in this regard.

[0167] In an optional embodiment of the present disclosure, the above obtaining the video generation model based on the request information can include the following steps:

[0168] The target scene template is determined from a plurality of preset scene templates based on the task scene identification, and the video generation model is searched from the model library based on the target scene template, wherein the model library stores a plurality of video generation models, and the request information includes the task scene identification of the target video task; or

[0169] The video generation model is searched from the model library based on the task model identification, wherein the request information includes the task model identification of the target video task.

[0170] Specifically, the task scene identification refers to a unique or specific label used to distinguish different task application scenarios. In the embodiments of the present disclosure, the task scene identification is part of the request information. Through the task scene identification, a target scene template matching the request information can be selected from a series of preset scene templates to generate task information. For example, the task scene identification is "video panning control", which means that the user wants to adjust the panning effect of the uploaded original video, and then a video panning control scene template can be selected from a plurality of preset scene templates according to the task scene identification.

[0171] The preset scene template is a standard configuration scene template defined in advance for different target video task application scenarios, and each template contains model information and task processing flow information matching the task application scenario. Through a series of preset scene templates, different scene task generation requests can be quickly responded. Different preset scene templates correspond to different task types, model information and processing flow. For example, there may be a video action control scene template in the preset scene template that is specifically for video action control tasks, which contains model information and processing flow of the trained video action control model.

[0172] The target scene template refers to the scene template matching the task scene identification. When analyzing the task generation request, the corresponding target scene template can be located based on the task scene identification, and the corresponding video generation model and other related configuration information can be selected from the model library according to the model information included in the target scene template. For example, when the task scene identification is "video object control", the target scene template is the template containing the model information and related configuration parameters of the video object control model.

[0173] The model library is a repository that centrally stores video generation models, which are trained and optimized to solve different target video tasks. Moreover, the video generation models in the model library can be divided into different versions according to different applicable tasks, for example, there may be a video panning control model for video panning control tasks, a video object control model for video object control tasks, and a video action control model for video action control tasks in the model library.

[0174] The task model identifier refers to a unique or specific label used to distinguish different models applicable to different tasks. For example, the task model identifier can be "video dolly control". Based on the task model identifier, a video dolly control model applicable to the video dolly control task can be found from the model library.

[0175] By using the scheme of the embodiments of the present disclosure, the pre-defined task scenario template, the task model identifier, and the model library resource make the acquisition process of the video generation model more flexible, efficient, and standard.

[0176] In another optional embodiment of the present disclosure, in addition to selecting a video generation model from the model library, a pre-trained generation model in the model library can also be trained according to the sample video in the request information and the sample object image, sample action sequence, and sample camera parameter corresponding to the sample video, to obtain a video generation model, that is, the request information includes a sample video of a target video task and a sample object image, a sample action sequence, and a sample camera parameter corresponding to the sample video; the above step of acquiring a video generation model based on the request information can include the following steps:

[0177] The pre-trained generation model corresponding to the target video task is trained based on the sample video and the sample object image, sample action sequence, and sample camera parameter corresponding to the sample video, to obtain a trained video generation model.

[0178] It should be noted that the implementation of "training the pre-trained generation model corresponding to the target video task based on the sample video and the sample object image, sample action sequence, and sample camera parameter corresponding to the sample video, to obtain a trained video generation model" is the same as the training method of the above-mentioned "video generation model", and the embodiments of the present disclosure will not be described in detail.

[0179] By using the scheme of the embodiments of the present disclosure, the pre-trained generation model corresponding to the target video task is trained based on the sample video, to obtain a trained video generation model, which makes the video generation model more consistent with the task scenario of the target video task on the basis of ensuring the accuracy of the video generation model.

[0180] Step 806: generating task information based on the video generation model, wherein the task information is used to execute the target video task.

[0181] Specifically, the task information includes model configuration and processing flow required for executing the target video task. The terminal device or other server components can correctly use the video generation model to process the target video task based on the task information. The target video task includes but is not limited to video dolly control, video action control task, video object control task, and video generation task.

[0182] It should be noted that based on the video generation model, the model parameters of the video generation model can be directly packaged to obtain the task information when generating the task information. Other model information of the video generation model can also be obtained, and the task information is constructed based on the other model information, wherein the other model information is, for example, a processing manner of model input data, a specification of an expected output result, and possible intermediate steps and other auxiliary information.

[0183] By applying the scheme of the embodiments of the present disclosure, the task information can be generated to ensure the processing quality and efficiency of the target video task while reducing the system deployment and operation and maintenance costs, thereby providing the user with convenient and efficient target video task processing services.

[0184] Referring to FIG. 9, FIG. 9 shows a structural schematic diagram of a task platform according to an embodiment of the present disclosure. The task platform includes a request interface 902 and a response unit 904.

[0185] The request interface 902 is configured to receive a task generation request, wherein the task generation request includes request information.

[0186] The response unit 904 is configured to obtain a video generation model based on the request information, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit and a time sequence attention unit in a pre-trained generation model based on a sample video, a sample object image corresponding to the sample video, a sample action sequence and a sample camera parameter, the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit and a generation unit in an initial generation model based on the sample video; and generate task information based on the video generation model, wherein the task information is used to execute a target video task.

[0187] In an optional embodiment of the present disclosure, the task platform further includes a model library, wherein the model library stores a plurality of video generation models.

[0188] The response unit is specifically configured to determine a target scene template from a plurality of preset scene templates based on a task scene identifier, and find a video generation model from the model library based on the target scene template, wherein the request information includes a task scene identifier of the target video task; or find a video generation model from the model library based on a task model identifier, wherein the request information includes a task model identifier of the target video task.

[0189] By applying the scheme of the embodiments of the present disclosure, the task information of the target video task can be generated to ensure the processing quality and efficiency of the target video task while reducing the system deployment and operation and maintenance costs, thereby providing the user with convenient and efficient target video task processing services.

[0190] The above is a schematic solution of the task platform of the embodiment. The technical solution of the task platform and the technical solution of the information processing method based on the video generation model belong to the same concept. For details of the technical solution of the task platform that are not described in detail, please refer to the description of the technical solution of the information processing method based on the video generation model.

[0191] Corresponding to the video generation method embodiment, the disclosure also provides a video generation device embodiment. FIG. 10 shows a structural schematic diagram of a video generation device according to an embodiment of the disclosure. As shown in FIG. 10, the device comprises:

[0192] The first acquisition module 1002 is configured to acquire video generation data, wherein the video generation data comprises virtual camera parameters, a reference action sequence of a target object, and an object image of the target object;

[0193] The first input module 1004 is configured to input the video generation data into a video generation model to obtain a target video of the target object, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit, and a time sequence attention unit in a pre-training generation model based on a sample video and a sample object image corresponding to the sample video, a sample action sequence, and sample camera parameters, and the pre-training generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit, and a generation unit in an initial generation model based on the sample video.

[0194] Optionally, the video generation model comprises an image encoding unit, an object encoding unit, an action encoding unit, a parameter encoding unit, a generation unit, and a decoding unit. The first input module 1004 is further configured to input the object image into the image encoding unit to obtain image encoding features; input the image encoding features into the object encoding unit to obtain object encoding features; input the reference action sequence into the action encoding unit to obtain an action encoding sequence; input the virtual camera parameters into the parameter encoding unit to obtain parameter encoding features; input the object encoding features, the action encoding sequence, and the parameter encoding features into the generation unit to obtain video encoding features; and input the video encoding features into the decoding unit to obtain the target video of the target object.

[0195] Optionally, the object encoding unit comprises a plurality of encoding blocks. The first input module 1004 is further configured to input the image encoding features into a first encoding block to obtain first encoding features output by the first encoding block, wherein the first encoding block is a first encoding block in the plurality of encoding blocks; input encoding features output by a previous encoding block of a second encoding block into the second encoding block to obtain second encoding features output by the second encoding block, wherein the second encoding block is any encoding block other than the first encoding block in the plurality of encoding blocks; and determine the object encoding features according to the encoding features output by the plurality of encoding blocks respectively.

[0196] Optionally, the generation unit comprises a plurality of generation blocks, the object encoding unit comprises a plurality of encoding blocks, and the generation blocks and the encoding blocks correspond one by one; the first input module 1004 is further configured to determine, according to the correspondence between the generation blocks and the encoding blocks, an encoding feature corresponding to a generation block; input the action encoding sequence, the parameter encoding feature, and the encoding feature corresponding to the first generation block into the first generation block to obtain a first generation feature output by the first generation block, wherein the first generation block is a first generation block in the plurality of generation blocks; input the generation feature output by the previous generation block of the second generation block, the encoding feature corresponding to the second generation block, and the parameter encoding feature into the second generation block to obtain a second generation feature output by the second generation block, wherein the second generation block is any generation block other than the first generation block in the plurality of generation blocks; in the case where the second generation block is the last generation block in the plurality of generation blocks, determine the second generation feature output by the second generation block as the video encoding feature.

[0197] Optionally, the second generation block comprises a self-attention unit, a cross-attention unit, a temporal attention unit, and a parameter cross-attention unit; the first input module 1004 is further configured to input the generation feature output by the previous generation block of the second generation block into the self-attention unit to obtain a self-attention feature; input the self-attention feature and the encoding feature corresponding to the second generation block into the cross-attention unit to obtain a cross-attention feature; input the cross-attention feature into the temporal attention unit to obtain a temporal attention feature; and input the temporal attention feature and the parameter encoding feature into the parameter cross-attention unit to obtain the second generation feature.

[0198] Optionally, the apparatus further comprises a first training module configured to obtain a sample video and a sample object image corresponding to the sample video, a sample action sequence, and a sample camera parameter; input the sample object image, the sample action sequence, and the sample camera parameter into the pre-trained generation model to obtain a first prediction feature output by the generation unit in the pre-trained generation model; input the sample video into the image encoding unit in the pre-trained generation model to obtain a first sample feature; and adjust the unit parameters of the parameter encoding unit, the temporal attention unit, and the parameter cross-attention unit in the pre-trained generation model according to the first prediction feature and the first sample feature to obtain a trained video generation model.

[0199] Optionally, the sample video includes a plurality of sample video frames; the apparatus further includes a second training module configured to, for a first sample video frame, determine a first sample object image and a first sample action image corresponding to the first sample video frame, where the first sample video frame is sampled from the plurality of sample video frames; input the first sample object image and the first sample action image into the initial generation model to obtain a second predicted feature output by a generation unit in the initial generation model; input the first sample video frame into an image encoding unit in the initial generation model to obtain a second sample feature; and adjust unit parameters of an object encoding unit, an action encoding unit, and the generation unit in the initial generation model according to the second predicted feature and the second sample feature, to obtain the pre-trained generation model trained.

[0200] Optionally, the first obtaining module 1002 is further configured to receive the object image of the target object and the reference video sent by the client; input the reference video into the action extraction model to obtain a reference action sequence; and input the reference video into the parameter extraction model to obtain the virtual camera parameter.

[0201] Optionally, the apparatus further includes a sending module configured to send the target video to the client; and receive result feedback information sent by the client, where the result feedback information is information fed back by the client on the target video; construct model optimization data according to the result feedback information; and adjust parameters of the video generation model by using the model optimization data.

[0202] By inputting the virtual camera parameter into the video generation model, the scheme of the embodiment of the present disclosure realizes the control of the camera operation of the target video, and because the video generation model includes the time sequence attention unit and the parameter cross attention unit, the video generation model can learn the time sequence information and the camera parameter, thereby generating the target video that supports a large range of camera operation effects and is stable in time sequence, so that the target video is more real and natural.

[0203] The above is a schematic scheme of the video generation apparatus of the embodiment. It should be noted that the technical scheme of the video generation apparatus and the technical scheme of the video generation method described above belong to the same concept, and the details of the technical scheme of the video generation apparatus that are not described in detail can be referred to the description of the technical scheme of the video generation method.

[0204] Corresponding to the motion video generation method of the virtual object described above, the present disclosure further provides a motion video generation apparatus of a virtual object, and FIG. 11 shows a structural schematic diagram of a motion video generation apparatus of a virtual object according to an embodiment of the present disclosure. As shown in FIG. 11, the apparatus includes:

[0205] The second obtaining module 1102 is configured to obtain motion video generation data of the virtual object, where the motion video generation data includes virtual camera parameters, a reference action sequence of the virtual object, and a virtual object image;

[0206] The second input module 1104 is configured to input the motion video generation data into a video generation model to obtain a target motion video of the virtual object, where the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit, and a timing attention unit in a pre-trained generation model based on a sample video and sample object images, sample action sequences, and sample camera parameters corresponding to the sample video, and the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit, and a generation unit in an initial generation model based on the sample video.

[0207] By inputting the virtual camera parameters into the video generation model, the application of the scheme of the embodiment of the present disclosure achieves the control of the camera operation of the target motion video. Moreover, because the video generation model includes the timing attention unit and the parameter cross attention unit, the video generation model can learn the timing information and the camera parameters, thereby generating the target motion video that supports a large range of camera operation effects and is stable in timing, so that the target motion video is more real and natural.

[0208] The above is a schematic scheme of the virtual object motion video generation device of the embodiment. It should be noted that the technical scheme of the virtual object motion video generation device is of the same concept as the technical scheme of the virtual object motion video generation method described above, and the details of the technical scheme of the virtual object motion video generation device that are not described in detail can be referred to the description of the technical scheme of the virtual object motion video generation method.

[0209] Corresponding to the video editing method embodiment described above, the present disclosure also provides a video editing device embodiment. FIG. 12 shows a structural schematic diagram of a video editing device according to an embodiment of the present disclosure. As shown in FIG. 12, the device includes:

[0210] The third obtaining module 1202 is configured to obtain an original video and target editing data, where the target editing data includes at least one of a target object image, a target camera parameter, and a target action sequence;

[0211] The analysis module 1204 is configured to analyze the original video based on the target editing data to determine original video data in the original video;

[0212] The third input module 1206 is configured to input the target editing data and the original video data into the video generation model to obtain a target editing video, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit and a timing attention unit in a pre-training generation model based on a sample video and sample object images, sample action sequences and sample camera parameters corresponding to the sample video, and the pre-training generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit and a generation unit in an initial generation model based on the sample video.

[0213] By inputting the virtual camera parameters into the video generation model, the operation control of the target editing video is realized, and because the video generation model includes the timing attention unit and the parameter cross attention unit, the video generation model can learn the timing information and the camera parameters, so that the target editing video with a large operation effect and stable timing is generated, and the target editing video is more real and natural.

[0214] The above is a schematic scheme of the video editing device according to the embodiment. It should be noted that the technical scheme of the video editing device belongs to the same concept as the technical scheme of the video editing method described above, and the details of the technical scheme of the video editing device that are not described in detail can be referred to the description of the technical scheme of the video editing method.

[0215] Corresponding to the video generation model training method embodiment described above, the disclosure further provides a video generation model training device embodiment. FIG. 13 shows a structural schematic diagram of a video generation model training device according to an embodiment of the disclosure. As shown in FIG. 13, the device includes:

[0216] The fourth acquisition module 1302 is configured to acquire a sample video and sample object images, sample action sequences and sample camera parameters corresponding to the sample video;

[0217] The fourth input module 1304 is configured to input the sample object images, the sample action sequences and the sample camera parameters into a pre-training generation model to obtain first predicted features output by a generation unit in the pre-training generation model, wherein the pre-training generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit and a generation unit in an initial generation model based on a sample video;

[0218] The fifth input module 1306 is configured to input the sample video into an image encoding unit in the pre-training generation model to obtain first sample features;

[0219] The adjusting module 1308 is configured to adjust the unit parameters of the parameter encoding unit, the timing attention unit and the parameter cross attention unit in the pre-training generation model according to the first predicted feature and the first sample feature, to obtain a trained video generation model.

[0220] By applying the scheme of the embodiment of the present disclosure, when the pre-training generation model is trained, the timing attention unit and the parameter cross attention unit are additionally added in each generation block included in the generation unit of the pre-training generation model, so that the video generation model obtained by training can learn the timing information and the camera parameter, thereby generating a target video supporting a large-amplitude dolly effect and stable timing, so that the target video is more real and natural.

[0221] The above is a schematic scheme of the video generation model training device of the embodiment. It should be noted that the technical scheme of the video generation model training device belongs to the same concept as the technical scheme of the video generation model training method described above, and the details of the technical scheme of the video generation model training device that are not described in detail can be referred to the description of the technical scheme of the video generation model training method.

[0222] Corresponding to the above-mentioned embodiment of the information processing method based on the video generation model, the present disclosure also provides an embodiment of an information processing device based on the video generation model. FIG. 14 shows a structural schematic diagram of an information processing device based on a video generation model according to an embodiment of the present disclosure. As shown in FIG. 14, the device includes:

[0223] The receiving module 1402 is configured to receive a task generation request, wherein the task generation request includes request information;

[0224] The fifth obtaining module 1404 is configured to obtain a video generation model based on the request information, wherein the video generation model is obtained by adjusting the parameters of the parameter encoding unit, the parameter cross attention unit and the timing attention unit in the pre-training generation model based on the sample video, the sample object image corresponding to the sample video, the sample action sequence and the sample camera parameter, and the pre-training generation model is obtained by adjusting the parameters of the object encoding unit, the action encoding unit and the generation unit in the initial generation model based on the sample video;

[0225] The generating module 1406 is configured to generate task information based on the video generation model, wherein the task information is used to execute a target video task.

[0226] Optionally, the fifth obtaining module 1404 is further configured to determine a target scene template from a plurality of preset scene templates based on the task scene identifier, and find the video generation model from a model library based on the target scene template, wherein the model library stores a plurality of video generation models, and the request information includes a task scene identifier of the target video task; or find the video generation model from the model library based on a task model identifier, wherein the request information includes a task model identifier of the target video task.

[0227] Optionally, the request information includes a sample video of the target video task and a sample object image, a sample action sequence and a sample camera parameter corresponding to the sample video; and the fifth obtaining module 1404 is further configured to train the pre-trained generation model corresponding to the target video task based on the sample video and the sample object image, the sample action sequence and the sample camera parameter corresponding to the sample video, to obtain the trained video generation model.

[0228] By applying the scheme of the embodiments of the present disclosure, the task information of the target video task is generated, the processing quality and efficiency of the target video task are ensured, the system deployment and operation and maintenance costs are reduced, and convenient and efficient task processing services are provided for users.

[0229] The above is a schematic scheme of the information processing device based on a video generation model according to the present embodiment. It should be noted that the technical scheme of the information processing device based on a video generation model belongs to the same concept as the technical scheme of the information processing method based on a video generation model described above, and the details of the technical scheme of the information processing device based on a video generation model that are not described in detail can be seen from the description of the technical scheme of the information processing method based on a video generation model described above.

[0230] FIG. 15 shows a structural block diagram of a computing device according to an embodiment of the present disclosure. The components of the computing device 1500 include, but are not limited to, a memory 1510 and a processor 1520. The processor 1520 is connected to the memory 1510 through a bus 1530, and a database 1550 is used to save data.

[0231] The computing device 1500 also includes an access device 1540 that enables the computing device 1500 to communicate via one or more networks 1560. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of such networks, such as the Internet. The access device 1540 can include one or more of any type of network interface (for example, a network interface card (NIC)) such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, or the like.

[0232] In one embodiment of the present disclosure, the above-mentioned components of the computing device 1500 and other components not shown in FIG. 15 can also be connected to each other, for example, through a bus. It should be understood that the computing device structure block diagram shown in FIG. 15 is only for the purpose of example, and is not a limitation on the scope of the present disclosure. Those skilled in the art can add or replace other components as needed.

[0233] The computing device 1500 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (for example, a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, and the like), a mobile phone (for example, a smartphone), a wearable computing device (for example, a smartwatch, smart glasses, and the like), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 1500 can also be a mobile or stationary server.

[0234] The processor 1520 is configured to execute computer program / instructions, which, when executed by the processor, implement the steps of the above-mentioned video generation method or the motion video generation method of a virtual object or the video editing method or the video generation model training method or the information processing method based on a video generation model.

[0235] The above is a schematic scheme of the computing device of the embodiment. It should be noted that the technical scheme of the computing device and the technical schemes of the video generation method, the motion video generation method of a virtual object, the video editing method, the video generation model training method, and the information processing method based on the video generation model belong to the same concept. For details of the technical scheme of the computing device that are not described in detail, refer to the description of the technical scheme of the video generation method or the motion video generation method of a virtual object or the video editing method or the video generation model training method or the information processing method based on the video generation model.

[0236] An embodiment of the present disclosure further provides a computer readable storage medium storing computer programs / instructions, which, when executed by a processor, implement the steps of the video generation method or the motion video generation method of a virtual object or the video editing method or the video generation model training method or the information processing method based on the video generation model.

[0237] The above is a schematic scheme of the computer readable storage medium of the embodiment. It should be noted that the technical scheme of the storage medium and the technical schemes of the video generation method, the motion video generation method of a virtual object, the video editing method, the video generation model training method, and the information processing method based on the video generation model belong to the same concept. For details of the technical scheme of the storage medium that are not described in detail, refer to the description of the technical scheme of the video generation method or the motion video generation method of a virtual object or the video editing method or the video generation model training method or the information processing method based on the video generation model.

[0238] An embodiment of the present disclosure further provides a computer program product, comprising computer programs / instructions, which, when executed by a processor, implement the steps of the video generation method or the motion video generation method of a virtual object or the video editing method or the video generation model training method or the information processing method based on the video generation model.

[0239] The above is a schematic scheme of the computer program product of the embodiment. It should be noted that the technical scheme of the computer program product and the technical schemes of the video generation method, the motion video generation method of a virtual object, the video editing method, the video generation model training method, and the information processing method based on the video generation model belong to the same concept. For details of the technical scheme of the computer program product that are not described in detail, refer to the description of the technical scheme of the video generation method or the motion video generation method of a virtual object or the video editing method or the video generation model training method or the information processing method based on the video generation model.

[0240] The above describes particular embodiments of the present disclosure. Other embodiments are within the scope of the following claims. In some cases, the actions or steps recited in the claims can be performed in a different order and still accomplish desirable results. Additionally, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order to achieve desirable results. In certain implementations, multitasking and parallel processing can be advantageous or necessary.

[0241] The computer readable medium can include any entity or apparatus capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, software distribution medium, etc. It should be noted that the computer readable medium can include appropriate additions or subtractions according to the requirements of patent practice, for example, in some regions, according to the patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.

[0242] It should be noted that for the foregoing method embodiments, in order to facilitate description, they are all expressed as a combination of a series of actions, but those skilled in the art should know that the embodiments of the present disclosure are not limited by the order of the actions described, because according to the embodiments of the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily necessary for the embodiments of the present disclosure.

[0243] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0244] The preferred embodiments of the present disclosure disclosed above are only used to help explain the present disclosure. The alternative embodiments do not describe all the details and do not limit the invention to the specific embodiments described. Obviously, according to the content of the embodiments of the present disclosure, many modifications and changes can be made. The present disclosure selects and describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present disclosure, so that those skilled in the art can well understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.

Claims

1. A video generation method, comprising: obtaining video generation data, wherein the video generation data comprises virtual camera parameters, a reference action sequence of a target object, and an object image of the target object; inputting the video generation data into a video generation model to obtain a target video of the target object, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit, and a time sequence attention unit in a pre-trained generation model based on a sample video and a sample object image, a sample action sequence, and sample camera parameters corresponding to the sample video, and the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit, and a generation unit in an initial generation model based on the sample video.

2. The method of claim 1, wherein the video generation model comprises an image encoding unit, an object encoding unit, an action encoding unit, a parameter encoding unit, a generation unit, and a decoding unit; the inputting the video generation data into a video generation model to obtain a target video of the target object comprises: inputting the object image into the image encoding unit to obtain image encoding features; inputting the image encoding features into the object encoding unit to obtain object encoding features; inputting the reference action sequence into the action encoding unit to obtain an action encoding sequence; inputting the virtual camera parameters into the parameter encoding unit to obtain parameter encoding features; inputting the object encoding features, the action encoding sequence, and the parameter encoding features into the generation unit to obtain video encoding features; inputting the video encoding features into the decoding unit to obtain the target video of the target object.

3. The method of claim 2, wherein the object encoding unit comprises a plurality of encoding blocks; the inputting the image encoding features into the object encoding unit to obtain object encoding features comprises: inputting the image encoding features into a first encoding block to obtain first encoding features output by the first encoding block, wherein the first encoding block is a first encoding block in the plurality of encoding blocks; inputting encoding features output by a previous encoding block of a second encoding block into the second encoding block to obtain second encoding features output by the second encoding block, wherein the second encoding block is any encoding block in the plurality of encoding blocks except the first encoding block; determining the object encoding features according to encoding features output by the plurality of encoding blocks respectively.

4. The method of claim 2 or 3, wherein the generation unit comprises a plurality of generation blocks, the object encoding unit comprises a plurality of encoding blocks, and the generation blocks and the encoding blocks correspond to each other one by one; the inputting the object encoding features, the action encoding sequence, and the parameter encoding features into the generation unit to obtain video encoding features comprises: determining encoding features corresponding to the generation blocks according to a correspondence between the generation blocks and the encoding blocks; input the action encoding sequence, the parameter encoding feature, and an encoding feature corresponding to a first generation block of the plurality of generation blocks into the first generation block to obtain a first generation feature output by the first generation block, wherein the first generation block is a first generation block of the plurality of generation blocks; input a generation feature output by a previous generation block of a second generation block, the parameter encoding feature, and an encoding feature corresponding to the second generation block into the second generation block to obtain a second generation feature output by the second generation block, wherein the second generation block is any generation block of the plurality of generation blocks other than the first generation block; in a case where the second generation block is a last generation block of the plurality of generation blocks, determine the second generation feature output by the second generation block as the video encoding feature.

5. The method of claim 4, wherein the second generation block comprises a self-attention unit, a cross-attention unit, a temporal attention unit, and a parameter cross-attention unit. The inputting a generation feature output by a previous generation block of a second generation block, the parameter encoding feature, and an encoding feature corresponding to the second generation block into the second generation block to obtain a second generation feature output by the second generation block comprises: inputting the generation feature output by the previous generation block of the second generation block into the self-attention unit to obtain a self-attention feature; inputting the self-attention feature and the encoding feature corresponding to the second generation block into the cross-attention unit to obtain a cross-attention feature; inputting the cross-attention feature into the temporal attention unit to obtain a temporal attention feature; inputting the temporal attention feature and the parameter encoding feature into the parameter cross-attention unit to obtain the second generation feature.

6. The method of any one of claims 1-5, wherein before the inputting the video generation data into the video generation model to obtain a target video of a target object, the method further comprises: obtaining a sample video, a sample object image corresponding to the sample video, a sample action sequence, and a sample camera parameter; inputting the sample object image, the sample action sequence, and the sample camera parameter into the pre-trained generation model to obtain a first prediction feature output by a generation unit in the pre-trained generation model; inputting the sample video into an image encoding unit in the pre-trained generation model to obtain a first sample feature; adjusting unit parameters of a parameter encoding unit, a temporal attention unit, and a parameter cross-attention unit in the pre-trained generation model according to the first prediction feature and the first sample feature to obtain the video generation model after training.

7. The method of claim 6, wherein the sample video comprises a plurality of sample video frames. Before the inputting the sample object image, the sample action sequence, and the sample camera parameter into the pre-trained generation model to obtain a first prediction feature output by a generation unit in the pre-trained generation model, the method further comprises: for a first sample video frame, determining a first sample object image and a first sample action image corresponding to the first sample video frame, wherein the first sample video frame is sampled from the plurality of sample video frames; inputting the first sample object image and the first sample action image into the initial generation model to obtain a second predicted feature output by a generation unit in the initial generation model; inputting the first sample video frame into an image encoding unit in the initial generation model to obtain a second sample feature; adjusting unit parameters of an object encoding unit, an action encoding unit and a generation unit in the initial generation model according to the second predicted feature and the second sample feature to obtain the pre-training generation model after training.

8. The method of any one of claims 1-7, wherein the obtaining video generation data comprises: receiving a reference video and an object image of the target object sent by a client; inputting the reference video into an action extraction model to obtain the reference action sequence; inputting the reference video into a parameter extraction model to obtain the virtual camera parameter.

9. The method of any one of claims 1-8, after the inputting the video generation data into a video generation model to obtain a target video of the target object, further comprising: sending the target video to the client; receiving result feedback information sent by the client, wherein the result feedback information is information fed back by the client on the target video; constructing model optimization data according to the result feedback information; and adjusting parameters of the video generation model using the model optimization data.

10. A method for generating a motion video of a virtual object, comprising: obtaining motion video generation data of a virtual object, wherein the motion video generation data comprises a virtual camera parameter, a reference action sequence of the virtual object and a virtual object image; inputting the motion video generation data into a video generation model to obtain a target motion video of the virtual object, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross-attention unit and a timing attention unit in a pre-training generation model based on a sample video and sample object images, sample action sequences and sample camera parameters corresponding to the sample video, and the pre-training generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit and a generation unit in an initial generation model based on the sample video.

11. A video editing method, comprising: obtaining an original video and target editing data, wherein the target editing data comprises at least one of a target object image, a target camera parameter and a target action sequence; analyzing the original video based on the target editing data to determine original video data in the original video; inputting the target editing data and the original video data into a video generation model to obtain a target editing video, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross-attention unit and a timing attention unit in a pre-training generation model based on a sample video and sample object images, sample action sequences and sample camera parameters corresponding to the sample video, and the pre-training generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit and a generation unit in an initial generation model based on the sample video.

12. A video generation model training method, comprising: obtaining a sample video and sample object images, sample action sequences and sample camera parameters corresponding to the sample video; inputting the sample object images, the sample action sequences and the sample camera parameters into a pre-trained generation model to obtain first predicted features output by a generation unit in the pre-trained generation model, wherein the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit and a generation unit in an initial generation model based on the sample video; inputting the sample video into an image encoding unit in the pre-trained generation model to obtain first sample features; adjusting unit parameters of a parameter encoding unit, a temporal attention unit and a parameter cross attention unit in the pre-trained generation model based on the first predicted features and the first sample features to obtain a trained video generation model.

13. An information processing method based on a video generation model, comprising: receiving a task generation request, wherein the task generation request comprises request information; obtaining a video generation model based on the request information, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit and a temporal attention unit in a pre-trained generation model based on sample videos and sample object images, sample action sequences and sample camera parameters corresponding to the sample videos, and the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit and a generation unit in an initial generation model based on the sample videos; generating task information based on the video generation model, wherein the task information is used to perform a target video task.

14. The method of claim 13, wherein the obtaining a video generation model based on the request information comprises: determining a target scene template from a plurality of preset scene templates based on a task scene identifier, and searching for a video generation model from a model library based on the target scene template, wherein the model library stores a plurality of video generation models, and the request information comprises the task scene identifier of the target video task; or searching for a video generation model from the model library based on a task model identifier, wherein the request information comprises the task model identifier of the target video task.

15. The method of claim 13, wherein the request information comprises sample videos and sample object images, sample action sequences and sample camera parameters corresponding to the sample videos of a target video task; the obtaining a video generation model based on the request information comprises: training a pre-trained generation model corresponding to the target video task based on the sample videos and the sample object images, sample action sequences and sample camera parameters corresponding to the sample videos to obtain a trained video generation model.

16. A task platform, comprising a request interface and a response unit; the request interface is configured to receive a task generation request, wherein the task generation request comprises request information; and the response unit is configured to generate task information based on a video generation model obtained based on the request information, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit and a temporal attention unit in a pre-trained generation model based on sample videos and sample object images, sample action sequences and sample camera parameters corresponding to the sample videos, and the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit and a generation unit in an initial generation model based on the sample videos. The response unit is configured to obtain a video generation model based on the request information, wherein the video generation model is obtained by adjusting parameters of a parameter encoding unit, a parameter cross attention unit, and a timing attention unit in a pre-trained generation model based on a sample video, a sample object image corresponding to the sample video, a sample action sequence, and sample camera parameters, and the pre-trained generation model is obtained by adjusting parameters of an object encoding unit, an action encoding unit, and a generation unit in an initial generation model based on the sample video; and generate task information based on the video generation model, wherein the task information is used to perform the target video task.

17. The task platform of claim 16, further comprising a model library, wherein, The model library stores a plurality of video generation models. The response unit is specifically configured to determine a target scene template from a plurality of preset scene templates based on a task scene identifier, and find a video generation model from the model library based on the target scene template, wherein the request information comprises the task scene identifier of the target video task; or find a video generation model from the model library based on a task model identifier, wherein the request information comprises the task model identifier of the target video task.

18. A computing device comprising: a memory and a processor; the memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method of any one of claims 1 to 9 or claim 10 or claim 11 or claim 12 or any one of claims 13 to 15.

19. A computer-readable storage medium storing computer programs / instructions, which, when executed by a processor, implement the steps of the method of any one of claims 1 to 9 or claim 10 or claim 11 or claim 12 or any one of claims 13 to 15.

20. A computer program product comprising computer programs / instructions, which, when executed by a processor, implement the steps of the method of any one of claims 1 to 9 or claim 10 or claim 11 or claim 12 or any one of claims 13 to 15.

Citation Information

Patent Citations

  • Image generation method, image generation model training method, device and equipment

    CN116580212A

  • Unsupervised learning of object representations from video sequences using spatial and temporal attention

    CN117255998A

  • Video generation method, deep learning model training method and device, equipment and storage medium

    CN118229815A

  • Video generation method, motion video generation method of virtual object, video editing method, video generation model training method and information processing method based on video generation model

    CN119031207A

  • Generating videos using sequences of generative neural networks

    US11908180B1

Cited By

  • Video generation neural networks with camera and subject motion inputs

    US20260170737A1