Method, apparatus, device and storage medium for generating a video
Patent Information
- Application Number
- CN202510368332.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2026-09-29
AI Technical Summary
然而,一些生成式模型所生成的视频内容的真实感较差,例如,不符合物理规律等
[0007]应当理解,本内容部分中所描述的内容并非旨在限定本公开的实施例的关键特征或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的描述而变得容易理解。
Smart Images

Figure CN122845877A_ABST
Abstract
Description
Technical Field
[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatus, devices, and computer-readable storage media for generating video. Background Technology
[0002] With the development of computer technology, generative models have been gradually applied to various task processing processes. For example, some generative models can generate video content based on input prompts. However, some generative models produce video content with poor realism, for example, failing to conform to physical laws. Summary of the Invention
[0003] In a first aspect of this disclosure, a method for generating video is provided. The method includes: receiving a video generation request; and processing the video generation request using a video generation model to generate video content, wherein the video generation model is trained based on a synthetic dataset, the synthetic dataset being generated by: using a scene generation unit to generate a synthetic video associated with a virtual 3D object; using a video description unit to generate descriptive text corresponding to the synthetic video; and constructing synthetic data samples in the synthetic dataset based on the synthetic video and the descriptive text.
[0004] In a second aspect of this disclosure, an apparatus for generating video is provided. The apparatus includes: a receiving module configured to receive a video generation request; and a generation module configured to process the video generation request using a video generation model to generate video content, wherein the video generation model is trained based on a synthetic dataset, the synthetic dataset being generated based on the following process: generating a synthetic video associated with a virtual 3D object using a scene generation unit; generating descriptive text corresponding to the synthetic video using a video description unit; and constructing synthetic data samples in the synthetic dataset based on the synthetic video and the descriptive text.
[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.
[0007] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0008] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0009] Figure 1 A schematic diagram is shown of an example environment in which embodiments of the present disclosure may be implemented;
[0010] Figure 2 An example process for training a video generation model according to some embodiments of this disclosure is shown;
[0011] Figure 3 A flowchart illustrating an example process for generating video according to some embodiments of this disclosure is shown;
[0012] Figure 4 A schematic structural block diagram of an example apparatus for generating video according to some embodiments of the present disclosure is shown; and
[0013] Figure 5 A block diagram of an electronic device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation
[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0015] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0016] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0017] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
[0018] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
[0019] As mentioned above, traditional generative models may have shortcomings in terms of physical realism in the video content they generate. For example, when the camera moves or objects deform, the generated video often fails to maintain the three-dimensional consistency of the objects. Furthermore, the movement of people or objects in the generated video may lack realism; for instance, when performing complex actions, a person's body movements may appear stiff or unnatural. When generating videos involving large-scale camera rotation or movement, existing models may fail to accurately simulate the motion effects of a real camera, resulting in videos that look unnatural.
[0020] Embodiments of this disclosure propose a scheme for generating videos. The scheme includes: receiving a video generation request; and processing the video generation request using a video generation model to generate video content, wherein the video generation model is trained based on a synthetic dataset, and the synthetic dataset is generated based on the following process: using a scene generation unit to generate a synthetic video associated with a virtual 3D object; using a video description unit to generate descriptive text corresponding to the synthetic video; and constructing synthetic data samples in the synthetic dataset based on the synthetic video and the descriptive text.
[0021] On the one hand, training video generation models using synthetic datasets can significantly improve the physical fidelity and realism of generated videos. Synthetic videos strictly adhere to physical laws during generation, making them physically closer to real videos. Furthermore, the diversity and scalability of synthetic datasets allow models to access a wider range of scenes and object types, thereby enhancing their generalization ability to different scenes and objects and demonstrating greater robustness in handling various complex scenarios.
[0022] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.
[0023] Example Environment
[0024] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, example environment 100 may include electronic device 110.
[0025] like Figure 1 As shown, the electronic device 110 can receive a user's video generation request 120. In some embodiments, the video generation request 120 may include a prompt message entered by the user. As an example, the prompt message may describe the content of the video to be generated, such as "a cat playing guitar".
[0026] In addition, the video generation request 120 may also specify other appropriate generation parameters. As an example, such generation parameters may include, but are not limited to: the length of the video content, the resolution of the video content, and the first or last frame of the video content.
[0027] Furthermore, the electronic device 110 can utilize the video generation model 130 to process the video generation request 120 to generate video content 140. As an example, the video generation model 130 can be any suitable type of generative model, such as a diffusion model.
[0028] The specific training process for video generation model 130 will be discussed in the following text. Figure 2 Detailed description.
[0029] In some embodiments, the electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).
[0030] In some embodiments, electronic device 110 may also be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Electronic device 110 may include, for example, computing systems / servers, such as mainframes, edge computing nodes, computing devices in a cloud environment, etc.
[0031] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0032] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.
[0033] Model training process
[0034] The process of generating video according to embodiments of the present disclosure will now be described with reference to the accompanying drawings. Figure 2 A schematic diagram 200 of a trained video generation model according to some embodiments of the present disclosure is shown.
[0035] like Figure 2 As shown, a synthetic dataset 240 can first be constructed for training the video generation model 130. Specifically, the scene generation unit 210 can be used to generate synthetic videos 220 associated with virtual 3D objects.
[0036] In some embodiments, the scene generation unit 210 may include a 3D scene generator, which can be used to generate a video scene containing a single 3D object. Specifically, the 3D scene generator may include multiple preset 3D assets, which may include various 3D models.
[0037] Furthermore, camera parameters associated with the virtual camera can be set. These parameters control the virtual camera's behavior, including its motion type, initial position, focus, and focal length. Specifically, the virtual camera's motion type determines its trajectory around the object, such as basic motion types like panning, tilting, and rotation. The initial position and focus parameters specify the camera's starting position and how it focuses on the main object. Additionally, the focal length parameter adjusts the camera's field of view, controlling the object's size on the screen to simulate different shooting angles and effects.
[0038] Furthermore, to further enhance the realism of the scene, the scene generation unit 210 can also perform joint modeling of lighting and environment. As an example, the scene generation unit 210 can support three main configurations: environment mapping, solid color indoor room, and empty scene.
[0039] In some examples, the environment map not only provides the background but also serves as the main light source. A solid-color interior room configuration uses two light sources for illumination, one above the object and the other placed elsewhere in the scene to simulate the lighting effects of an interior environment. An empty scene configuration uses either an environment map or two light sources for illumination, with the surrounding environment left blank; this configuration is suitable for scenes requiring a simpler background.
[0040] In some embodiments, to generate a large number of diverse synthetic videos, the above parameters can be defined through configuration information (e.g., configuration files). The scene generation unit 210 can parse the configuration information to set up the corresponding virtual scene and use the rendering engine to render the scene into a video.
[0041] In some embodiments, in addition to the object parameters of the virtual 3D object (e.g., type, size, etc.) and the camera parameters of the virtual camera (e.g., camera position, camera motion trajectory, etc.) mentioned above, the configuration information may also define scene parameters of the virtual scene associated with the virtual 3D object. As an example, scene parameters may include indicator lights and environmental parameters for constructing the virtual scene where the virtual 3D object is placed. For instance, scene parameters may indicate whether the virtual 3D object is placed in an outdoor environment or an indoor environment.
[0042] In some embodiments, the specific value of a parameter can also be determined by sampling from a preset range of values. For example, random sampling can be performed based on the probability distribution of the parameter to determine its specific value. Thus, during large-scale generation, each sampling generates a unique configuration file, which is then rendered into an independent synthetic video. In this way, embodiments of this disclosure can generate a large number of diverse synthetic videos with minimal human intervention.
[0043] Furthermore, the video description unit 230 can also be used to generate a caption for the synthesized video 220. In some embodiments, unlike generating a text description for a video directly using a visual model, the video description unit 230 can construct the final caption by combining multi-element descriptions of multiple preset elements in the synthesized video 220.
[0044] Specifically, the process of generating a synthetic video is achieved by combining one or more 3D objects, scene settings, and camera movement. Therefore, when generating descriptive text, the video description unit 230 first generates detailed descriptive text for each individual element (such as 3D objects, virtual scenes, virtual camera movement, etc.). For example, for a video containing dancers, the video description unit 230 can separately describe the dancers' movements (e.g., "performing street dance moves"), the scene background (e.g., "in an indoor environment"), and the camera's movement (e.g., "the camera rotates around the dancer").
[0045] In some embodiments, the video description unit 230 may utilize a visual model to generate element descriptions of individual elements based on the synthesized video 220. In some embodiments, the visual model may also generate corresponding element descriptions based on synthesis parameters used to generate the synthesized video 220.
[0046] As an example, when describing camera motion, the visual model can generate corresponding descriptive text based on the video content 220 and the camera motion parameters used by the scene generation unit 210.
[0047] Furthermore, this method of generating descriptive text significantly reduces the required time cost. Assuming a composite video 220 has N different 3D objects, M scene settings, and C camera motion methods, traditional video description methods would require generating descriptive text for each of the N×M×C different video combinations. However, by generating descriptive text for each element separately, only N+M+C element descriptive texts need to be generated, and then these texts can be combined to generate the descriptive text for any video. This not only significantly reduces the complexity of generating descriptive text but also improves the speed and efficiency of descriptive text generation.
[0048] Furthermore, synthetic data samples in the synthetic video dataset 240 can be constructed based on the generated synthetic video 220 and its corresponding descriptive text. The following section will further describe how to use the synthetic video dataset 240 to train the video generation model 130.
[0049] like Figure 2 As shown, the video generation model 130 can be trained using a combination of real video dataset 250 and synthetic video dataset 240. The real video dataset 250 may include real video content that has been captured.
[0050] Furthermore, to balance the distribution differences between real and synthetic videos, a reference generation model 260 can be used to assist in training the video generation model 130. Specifically, the reference generation model 260 can, for example, have the same model structure as the real generation model 130, such as a diffusion model.
[0051] Furthermore, the reference generative model 260 can be trained, for example, solely based on the synthetic video dataset 240. Specifically, training samples can be constructed using the synthetic video dataset 240, and such training samples may include the synthetic video 220 and its corresponding descriptive text. Unlike the descriptive text used to train the video generation model 130, the descriptive text used to train the reference generative video 260 may omit the descriptive content of the synthetic video 220 regarding the target element, and such descriptive content can be retained in the descriptive text used to train the video generation model 130.
[0052] For example, since the synthetic video 220 has relatively weak realism in its depiction of the virtual 3D objects' movements, the descriptive text used to train the reference generative model 260 may not include descriptions of the objects' movements. This allows the reference generative model to focus more on the visual style of the synthetic video.
[0053] Furthermore, during the training of video generation model 130, the same input prompts can be provided to both video generation model 130 and reference generation model 260. Taking video generation model 130 and reference generation model 260 as diffusion models as an example, feature fusion can be performed on the features of the two models during the noise reduction process in box 270.
[0054] Specifically, the feature fusion process can be represented as:
[0055]
[0056] Among them, V θ This indicates that the video generation model is 130, V σ Indicates reference generative model 260, l k-1Let represent the first feature at the (k-1)th denoising time step (i.e., the first time step), and t represent the positive cue information of the video generation model 130. This represents the positive feedback information from the reference generation model 260, and n represents the negative feedback information from the video generation model 130. This indicates a negative hint from the reference generative model 260, where α and β are weighting coefficients. This represents the first guidance information corresponding to the second denoising time step, determined using the reference generation model based on the input prompt information and the first feature corresponding to the first denoising time step; f(V θ , l k-1 (t, n) represents the second guidance information corresponding to the second denoising time step, determined by the video generation model based on the input prompt information and the first feature corresponding to the first denoising time step. As an example, the second guidance information can be represented as f θ (V θ (l, t, n) = V θ (l,t)-V θ (l, n), that is, the second guidance information can characterize the difference between the noise reduction results corresponding to positive and negative prompts.
[0057] Therefore, based on the above formula, the second feature l corresponding to the second noise reduction time step can be determined based on the noise reduction result of the video generation model at the second noise reduction time step, the first guiding information, and the second guiding information. k .
[0058] Furthermore, the generation result of the video generation model 130 can be determined based on the second feature, and the video generation model 130 can be trained based on the generation result and the synthesized video. For example, after completing multiple rounds of noise reduction, video features can be determined, and the generation result can be obtained based on the video features. The difference between the generated result and the synthesized video can be used to train the video generation model 130.
[0059] Furthermore, to evaluate the model's performance during training, both quantitative and qualitative evaluation metrics can be used. Quantitative evaluation can measure the physical fidelity of the generated video using metrics such as 3D reconstruction error and human pose estimation confidence. These metrics effectively assess the performance of the generated video in terms of 3D consistency and human pose integrity. Qualitative evaluation allows professionals to subjectively assess the physical fidelity of the generated video to ensure that the model's output conforms to human visual perception.
[0060] Through the training process described above, on the one hand, training the video generation model using synthetic datasets can significantly improve the physical fidelity and realism of the generated videos. Synthetic videos strictly adhere to physical laws during generation, making them physically closer to real videos. Furthermore, the diversity and scalability of synthetic datasets allow the model to access a wider range of scenes and object types, thereby enhancing its generalization ability to different scenes and objects and demonstrating greater robustness when handling various complex scenarios.
[0061] Example process
[0062] Figure 3 A flowchart of an example process 300 for generating video according to some embodiments of the present disclosure is shown. Process 300 can be implemented at electronic device 110.
[0063] As shown in the figure, in box 310, electronic device 110 receives a video generation request.
[0064] In box 320, electronic device 110 uses a video generation model to process a video generation request in order to generate video content. The video generation model is trained on a synthetic dataset, which is generated based on the following process: using a scene generation unit to generate a synthetic video associated with a virtual 3D object; using a video description unit to generate descriptive text corresponding to the synthetic video; and constructing synthetic data samples in the synthetic dataset based on the synthetic video and the descriptive text.
[0065] In some embodiments, generating a composite video associated with a virtual 3D object using a scene generation unit includes: obtaining configuration information describing a set of parameters for generating the composite video; and processing the configuration information using the scene generation unit to generate the composite video.
[0066] In this way, embodiments of the present disclosure can use the configuration information to precisely control the generation process of the synthesized video, thereby improving the controllability of the synthesized video.
[0067] In some embodiments, the value of at least one parameter in a set of parameters is determined by sampling from a preset value range.
[0068] In this way, embodiments of the present disclosure can improve the randomness and diversity of synthesized videos.
[0069] In some embodiments, a set of parameters includes at least one of the following: object parameters of the virtual 3D object; camera parameters of the virtual camera associated with the synthesized video; and scene parameters of the virtual scene associated with the virtual 3D object.
[0070] In this way, the embodiments of this disclosure can ensure that the synthesized video conforms to physical laws and has a sense of realism.
[0071] In some embodiments, generating descriptive text corresponding to the synthesized video using the video description unit includes: generating multi-element description content of the synthesized video about multiple preset elements using the video description unit; and constructing descriptive text corresponding to the synthesized video based on the multi-element description content.
[0072] In this way, the embodiments of this disclosure can improve the accuracy of descriptive text and enhance the model's ability to understand video content.
[0073] In some embodiments, at least one of the multiple element description contents is determined based on the synthesis parameters of the synthesized video.
[0074] In this way, the embodiments of this disclosure can more accurately describe the synthesized video, thereby improving the training effect of the model.
[0075] In some embodiments, the plurality of preset elements include at least one of the following: a virtual 3D object; a virtual scene associated with the virtual 3D object; and a virtual camera associated with the synthesized video.
[0076] In this way, embodiments of this disclosure can cover the key elements of video generation, thereby providing more comprehensive semantic information.
[0077] In some embodiments, the video generation model is trained on a training dataset, which includes a synthetic dataset and a real dataset, with the real dataset including captured video content.
[0078] In this way, embodiments of the present disclosure can combine synthetic data and real data for training, enabling the model to learn the physical fidelity of synthetic data and the visual style of real data simultaneously, thereby generating videos that are both physically accurate and realistic.
[0079] In some embodiments, the video generation model is a diffusion model, and the video generation model is trained based on the following process: using a reference generation model to determine first guidance information corresponding to a second denoising time step based on input prompt information and a first feature corresponding to a first denoising time step, wherein the input prompt information is determined based on descriptive text; using the video generation model to determine second guidance information corresponding to the second denoising time step based on the input prompt information and the first feature; determining a second feature corresponding to the second denoising time step based on the denoising result of the video generation model at the second denoising time step, the first guidance information, and the second guidance information; determining the generation result of the video generation model based on the second feature; and training the video generation model based on the generation result and the synthesized video.
[0080] In this way, embodiments of the present disclosure can use a reference generative model to capture visual patterns in synthetic data and suppress these patterns during training, which can effectively reduce visual artifacts introduced by synthetic data and improve the realism of generated videos.
[0081] In some embodiments, input prompts include positive and negative prompts determined based on descriptive text.
[0082] In this way, the embodiments of this disclosure can more accurately determine whether the model generates video content that meets the requirements.
[0083] In some embodiments, the descriptive text is a first descriptive text, and the reference generative model is trained based on the following process: constructing training samples using a synthetic dataset, the training samples including a synthetic video and a second descriptive text, the second descriptive text excluding descriptions of target elements in the synthetic video and the first descriptive text including descriptions of target elements in the synthetic video; and training the reference generative model using the training samples.
[0084] In this way, embodiments of the present disclosure enable the reference video model to focus on capturing visual patterns in the synthetic data, thereby avoiding the generation of visual artifacts.
[0085] In some embodiments, the target element includes the action of a virtual 3D object.
[0086] In this way, embodiments of this disclosure can enhance the realism of the generated video content.
[0087] Example devices and equipment
[0088] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 4 A schematic structural block diagram of an example device 400 for generating video according to certain embodiments of the present disclosure is shown. Device 400 may be implemented as or included in electronic device 110. Various modules / components in device 400 may be implemented by hardware, software, firmware, or any combination thereof.
[0089] like Figure 4 As shown, the apparatus 400 includes a receiving module 410 configured to receive a video generation request; and a generation module 420 configured to process the video generation request using a video generation model to generate video content. The video generation model is trained based on a synthetic dataset, which is generated through the following process: using a scene generation unit to generate a synthetic video associated with a virtual 3D object; using a video description unit to generate descriptive text corresponding to the synthetic video; and constructing synthetic data samples in the synthetic dataset based on the synthetic video and the descriptive text.
[0090] In some embodiments, generating a composite video associated with a virtual 3D object using a scene generation unit includes: obtaining configuration information describing a set of parameters for generating the composite video; and processing the configuration information using the scene generation unit to generate the composite video.
[0091] In some embodiments, the value of at least one parameter in a set of parameters is determined by sampling from a preset value range.
[0092] In some embodiments, a set of parameters includes at least one of the following: object parameters of the virtual 3D object; camera parameters of the virtual camera associated with the synthesized video; and scene parameters of the virtual scene associated with the virtual 3D object.
[0093] In some embodiments, generating descriptive text corresponding to the synthesized video using the video description unit includes: generating multi-element description content of the synthesized video about multiple preset elements using the video description unit; and constructing descriptive text corresponding to the synthesized video based on the multi-element description content.
[0094] In some embodiments, at least one of the multiple element description contents is determined based on the synthesis parameters of the synthesized video.
[0095] In some embodiments, the plurality of preset elements include at least one of the following: a virtual 3D object; a virtual scene associated with the virtual 3D object; and a virtual camera associated with the synthesized video.
[0096] In some embodiments, the video generation model is trained on a training dataset, which includes a synthetic dataset and a real dataset, with the real dataset including captured video content.
[0097] In some embodiments, the video generation model is a diffusion model, and the video generation model is trained based on the following process: using a reference generation model to determine first guidance information corresponding to a second denoising time step based on input prompt information and a first feature corresponding to a first denoising time step, wherein the input prompt information is determined based on descriptive text; using the video generation model to determine second guidance information corresponding to the second denoising time step based on the input prompt information and the first feature; determining a second feature corresponding to the second denoising time step based on the denoising result of the video generation model at the second denoising time step, the first guidance information, and the second guidance information; determining the generation result of the video generation model based on the second feature; and training the video generation model based on the generation result and the synthesized video.
[0098] In some embodiments, input prompts include positive and negative prompts determined based on descriptive text.
[0099] In some embodiments, the descriptive text is a first descriptive text, and the reference generative model is trained based on the following process: constructing training samples using a synthetic dataset, the training samples including a synthetic video and a second descriptive text, the second descriptive text excluding descriptions of target elements in the synthetic video and the first descriptive text including descriptions of target elements in the synthetic video; and training the reference generative model using the training samples.
[0100] In some embodiments, the target element includes the action of a virtual 3D object.
[0101] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 5 The electronic device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 5 The electronic device 500 shown can be used to achieve Figure 1 Electronic devices 110.
[0102] like Figure 5 As shown, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500.
[0103] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.
[0104] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 5As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
[0105] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0106] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0107] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0108] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0109] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0110] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0111] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0112] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for generating video, comprising: Receive video generation request; as well as The video generation request is processed using a video generation model to generate video content, wherein the video generation model is trained on a synthetic dataset, which is generated based on the following process: Using scene generation units, synthesized videos associated with virtual 3D objects are generated; Using a video description unit, generate description text corresponding to the synthesized video; as well as Based on the synthesized video and the descriptive text, synthesized data samples are constructed in the synthesized dataset.
2. The method according to claim 1, wherein generating a synthetic video associated with a virtual 3D object using a scene generation unit comprises: Obtain configuration information, which describes a set of parameters used to generate the synthesized video; as well as The scene generation unit processes the configuration information to generate the synthesized video.
3. The method according to claim 2, wherein the value of at least one parameter in the set of parameters is determined by sampling from a preset value range.
4. The method of claim 2, wherein the set of parameters includes at least one of the following: The object parameters of the virtual 3D object; Camera parameters of the virtual camera associated with the synthesized video; Scene parameters of the virtual scene associated with the virtual 3D object.
5. The method according to claim 1, wherein generating descriptive text corresponding to the synthesized video using a video description unit comprises: Using the video description unit, generate multi-element description content of the synthesized video about multiple preset elements; as well as Based on the descriptions of the multiple elements, a descriptive text corresponding to the synthesized video is constructed.
6. The method according to claim 5, wherein at least one element description content of the plurality of element description contents is determined based on the synthesis parameters of the synthesized video.
7. The method of claim 5, wherein the plurality of preset elements includes at least one of the following: The virtual three-dimensional object; The virtual scene associated with the virtual 3D object; A virtual camera associated with the synthesized video.
8. The method according to claim 1, wherein the video generation model is trained based on a training dataset, the training dataset including the synthetic dataset and a real dataset, the real dataset including the captured video content.
9. The method of claim 1, wherein the video generation model is a diffusion model, and the video generation model is trained based on the following process: Using a reference generation model, based on input prompt information and a first feature corresponding to the first noise reduction time step, a first guidance information corresponding to the second noise reduction time step is determined, wherein the input prompt information is determined based on the description text; The video generation model uses the input prompt information and the first feature to determine the second guidance information corresponding to the second noise reduction time step; Based on the denoising result of the video generation model at the second denoising time step, the first guidance information, and the second guidance information, a second feature corresponding to the second denoising time step is determined; Based on the second feature, the generation result of the video generation model is determined; as well as The video generation model is trained based on the generated results and the synthesized video.
10. The method of claim 9, wherein the input prompt information includes positive prompt information and negative prompt information determined based on the descriptive text.
11. The method of claim 9, wherein the descriptive text is a first descriptive text, and the reference generation model is trained based on the following process: Using the synthetic dataset, training samples are constructed, the training samples including the synthetic video and a second descriptive text, wherein the second descriptive text does not include descriptions of target elements in the synthetic video, and the first descriptive text includes descriptions of the target elements in the synthetic video; and The reference generative model is trained using the training samples.
12. The method of claim 11, wherein the target element includes the action of the virtual 3D object.
13. An apparatus for generating video, comprising: The receiving module is configured to receive video generation requests; as well as A generation module is configured to process the video generation request using a video generation model to generate video content, wherein the video generation model is trained based on a synthetic dataset, which is generated based on the following process: Using scene generation units, synthesized videos associated with virtual 3D objects are generated; Using a video description unit, generate description text corresponding to the synthesized video; as well as Based on the synthesized video and the descriptive text, synthesized data samples are constructed in the synthesized dataset.
14. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 12 when executed by the at least one processing unit.
15. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 12.