Method and apparatus for generating video, and device and storage medium

By using a method to generate videos, a diffusion model and dynamic conditional guidance parameters are employed to generate high-quality, highly stable videos based on reference images and audio control signals. This solves the problems of inconsistent video generation and color saturation in existing technologies, and enables high-quality video generation in complex scenes.

WO2026157632A1PCT designated stage Publication Date: 2026-07-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2025-12-12
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing technologies are insufficient to meet the demands for high-quality and stable video generation in complex scenarios in the field of digital human video-driven systems, especially in terms of unstable response to audio control signals.

Method used

By acquiring reference images and audio control signals, a first model is used to generate descriptive text, and a second model is used to generate target video based on the reference images, descriptive text, and control signals. By combining a diffusion model and dynamic conditional guidance parameters, the consistency and stability of the generated video and audio control signals are ensured.

Benefits of technology

It improves the quality and stability of video driving, ensures a close correlation between the generated video and audio control signals, and solves the problems of inconsistency in video generation and exaggerated color saturation in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025142100_30072026_PF_FP_ABST
    Figure CN2025142100_30072026_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure relate to a method and apparatus for generating a video, and a device and a computer-readable storage medium. The method provided herein comprises: acquiring a reference image and an audio control signal (310); using a first model to generate description text of the reference image (320); and using a second model to generate a target video on the basis of the reference image, the description text and the control signal, wherein the target video comprises motion content associated with a preset object in the reference image, and the motion content corresponds to the audio control signal (330).
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, devices and storage media for generating video

[0001] This application claims priority to Chinese Patent Application No. 202510097055.7, filed on January 21, 2025, entitled "Method, Apparatus, Device and Storage Medium for Generating Video", the entire contents of which are incorporated herein by reference. Technical Field

[0002] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatus, devices, and computer-readable storage media for generating video. Background Technology

[0003] With the rapid development of computer technology, the field of machine learning has made significant progress. Some solutions enable users to drive image generation through control signals, thereby creating dynamic video content. The core of this technology lies in transforming input control signals (such as audio) into visual dynamic representations, providing users with a completely new interactive experience.

[0004] For example, in the field of digital humans, users can input audio signals to drive digital humans to perform actions that match the audio content, enabling digital humans to react in real time based on changes in tone and rhythm of speech. This capability not only enriches the ways digital content is created. Summary of the Invention

[0005] In a first aspect of this disclosure, a method for generating a video is provided. The method includes: acquiring a reference image and an audio control signal; generating descriptive text for the reference image using a first model; and generating a target video based on the reference image, the descriptive text, and the control signal using a second model. The target video includes motion content associated with a preset object in the reference image, the motion content corresponding to the audio control signal.

[0006] In a second aspect of this disclosure, an apparatus for generating video is provided. The apparatus includes: an acquisition module configured to acquire a reference image and an audio control signal; a text generation module configured to generate descriptive text of the reference image using a first model; and a video generation module configured to generate a target video based on the reference image, the descriptive text, and the control signal using a second model, the target video including motion content associated with a preset object in the reference image, the motion content corresponding to the audio control signal.

[0007] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.

[0008] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the first aspect.

[0009] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.

[0010] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0012] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure may be implemented;

[0013] Figure 2 illustrates an example model for generating video according to some embodiments of the present disclosure;

[0014] Figure 3 illustrates a flowchart of an example process for generating video according to some embodiments of the present disclosure;

[0015] Figure 4 shows a schematic structural block diagram of an example apparatus for generating video according to some embodiments of the present disclosure; and

[0016] Figure 5 shows a block diagram of an electronic device capable of implementing several embodiments of the present disclosure. Detailed Implementation

[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0018] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0019] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0020] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0021] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.

[0022] In recent years, with the rapid development of computer technology, diffusion models have become a powerful tool for generating high-quality images and videos. However, despite the significant progress made in diffusion model technology, existing research still faces many challenges in its application to digital human video-driven fields. Currently, most related research is either based on the U-Net framework or limited to portrait scenes, making it difficult to meet the needs of high-quality, high-stability video generation in complex scenarios.

[0023] Embodiments of this disclosure propose a scheme for generating video. The scheme includes: acquiring a reference image and an audio control signal; generating descriptive text for the reference image using a first model; and generating a target video based on the reference image, the descriptive text, and the control signal using a second model. The target video includes motion content associated with a preset object in the reference image, and the motion content corresponds to the audio control signal.

[0024] In this way, embodiments of the present disclosure can control the video generation process based on the descriptive text of the image to be driven, thereby improving the quality of video driving.

[0025] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.

[0026] Example Environment

[0027] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in Figure 1, the example environment 100 may include an electronic device 110.

[0028] As shown in Figure 1, the electronic device 110 can receive a control signal 130 and a reference image 120. The control signal 130 can cover multiple modalities, including but not limited to text, audio, or video content. As will be explained in detail below, the electronic device 110 can generate a target video 150 by inputting the reference image 120 and the control signal 130 into the model 140.

[0029] Taking a specific scenario as an example, the target video 150 can display dynamic content related to a preset object in the reference image 120, and this dynamic content is closely related to the received control signal 130.

[0030] In one example scenario, control signal 130 may contain both reference video content and audio content. In this case, the dynamic content in target video 150 may include: body movements generated based on skeletal information from the reference video content, and mouth movements generated based on the audio content.

[0031] In some embodiments, electronic device 110 may be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 may also support any type of user-facing interface (such as "wearable" circuitry).

[0032] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0033] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0034] Example Model

[0035] Figure 2 illustrates an example model for generating video according to some embodiments of the present disclosure. The specific process of the model generating the target video will be described below with reference to Figure 2.

[0036] As shown in Figure 2, model 245 can acquire reference token 225 corresponding to reference image 220, and can acquire one or more control signals for driving reference image 220.

[0037] As an example, a reference token 225 can be generated by encoding a reference image 220. Furthermore, model 245 can also acquire an audio token generated based on audio control signal 240. As an example, audio token 240 can be generated by encoding audio control signal 240 (e.g., an audio segment).

[0038] As shown in Figure 2, model 245 can generate target video 250 based on the received reference token 225 and audio token 240. In some embodiments, model 245 can be implemented based on a diffusion model, which can determine video features by denoising the noise 205 to generate the corresponding target video 250.

[0039] In some embodiments, to improve the quality of video driving, the electronic device 110 may also utilize a first model to generate descriptive text 230 for the reference image 220. As an example, the first model may further include an appropriate visual model, and the descriptive text 230 may be used to describe the content of the reference image 220. As an example, the descriptive text 230 may also be referred to as the caption of the reference image 220.

[0040] As shown in the figure, the electronic device 110 can also generate a text token 215 based on the descriptive text 230. For example, by processing the descriptive text 230 using an encoding model, a corresponding text token 215 can be generated. Subsequently, the model 245 can combine the text token, reference token 225, and audio token 240 to perform noise reduction processing on the noise 205, thereby generating video features.

[0041] In some scenarios, the electronic device 110 can also acquire reference video content as a control signal. Furthermore, skeleton information 235 can be determined based on the reference video content. For example, skeleton information 235 can characterize a set of key points of a preset object in the reference video content.

[0042] Furthermore, the skeleton information 235 can be processed using the coding unit to extract skeleton features. For example, the feature size of the skeleton features can be consistent with that of the video token 210. Then, the video token 210 and the skeleton features can be concatenated along the channel dimension to form the input video features of model 245.

[0043] In some embodiments, when generating the first video segment of the target video 250, the video token 210 may be set to empty, for example. Furthermore, the first video segment can be used to generate subsequent video segments. As an example, the second video segment may also have a preset duration. For example, the video token 210 used to generate candidate video segments may be determined based on one or more video frames in the first video segment.

[0044] In some scenarios, the text token 215 may also include information corresponding to a prompt word entered by the user. For example, the target text can be obtained by combining the descriptive text 230 and the prompt word entered by the user. This text can, for example, be provided as a positive prompt word for model 245.

[0045] Furthermore, model 245 can generate target video 250 based on positive prompts, negative prompts, reference images, and audio control signals. Specifically, negative prompts may include, for example, preset text content that indicates one or more descriptions about unwanted video content. Conversely, positive prompts may include text cues that explicitly describe the generation target.

[0046] As an example, positive cue words can include descriptive text about a reference image, such as "a person is sitting in a chair." Negative cue words can include pre-defined text content, such as "avoid blurry, low resolution." Positive cue words can be encoded as conditional vectors and fed into a diffusion model to guide the model in generating content that conforms to the description. Similarly, negative cue words can be encoded as conditional vectors to suppress unwanted features.

[0047] In some scenarios, if the user does not input a prompt, the positive prompt may include, for example, only the descriptive text 230 of the reference image 220.

[0048] In some embodiments, model 245 can also control the generation process of target video 250 based on dynamic conditional guidance coefficients. Specifically, conditional guidance parameters can be determined based on the denoising time step of the diffusion model, and the conditional guidance parameters are positively correlated with the denoising time step. As an example, conditional guidance parameters may include CFG (Classifier-Free Guidance Scale) coefficients. CFG coefficients can be used to improve the quality and diversity of generated images while avoiding generation bias caused by over-reliance on conditional information.

[0049] In some embodiments, model 245 can generate video features of the target video based on conditional guidance parameters and feature difference information, wherein the feature difference information indicates the difference between a first video feature corresponding to a positive cue word and a second video feature corresponding to a negative cue word. For example, this process can be represented as: M(negative cue word, empty audio) + scale_factor * (M(positive cue word, audio) - M(negative cue word, audio)), where M represents the denoising process at one time step of the model, and scale_factor represents the conditional guidance parameters. M(positive cue word, audio) represents the first video feature generated based on the positive cue word and audio control signal, and M(negative cue word, audio) represents the second video feature generated based on rich text and audio control signal.

[0050] In some embodiments, the conditional guidance parameter (e.g., the CFG coefficient) can be positively correlated with the denoising time step. It should be understood that the denoising time step gradually decreases during the denoising process of the diffusion model. That is, the CFG coefficient can decrease as the denoising process progresses.

[0051] Therefore, the embodiments of this disclosure can ensure the consistency between the generated video and the input text and audio, and utilize dynamic CFG coefficients to solve problems such as exaggerated motion and color saturation.

[0052] In some embodiments, during the denoising process of model 245, model 245 may also determine a denoising gradient at each denoising time step, wherein the denoising gradient includes a noise perturbation term. As an example, model 245 may determine the denoising gradient of the model at each denoising time step and may add a noise perturbation term at each step of gradient descent to improve the stability of the generated result. In some embodiments, the noise perturbation term added during the denoising process is determined based on Langevin dynamics.

[0053] The training process of Model 245 will be further described below. Traditional diffusion models are usually trained using pre-trained text-to-image or text-to-video data, which makes the model lack stability in audio-driven scenarios.

[0054] In some embodiments, the training dataset used to train model 245 may include a first data subset and a second data subset. The first data subset includes multiple text-video pairs, and the second data subset includes multiple audio-video pairs.

[0055] In some embodiments, the first data subset corresponds to a first proportion of the training dataset, the second data subset corresponds to a second proportion of the training dataset, and the difference between the first proportion and the second proportion is less than a threshold. For example, the first data subset and the second data subset can each participate in the training at a 50% ratio. Based on the latency results, such a training data ratio can both ensure the accuracy of audio-driven lip-syncing and improve the generalization of the model.

[0056] In some embodiments, text conditional injection can be retained during model training to improve the model's ability to fit large-scale, diverse training data. Furthermore, by providing text guidance during testing, embodiments of this disclosure can ensure the stability of the video content generated by the model.

[0057] Furthermore, embodiments of this disclosure can utilize larger datasets for training, rather than relying on higher-quality training data. By relaxing the selection criteria for audio data, embodiments of this disclosure can increase the scale of training data and improve the quality of the model's generation of video content based on audio control signals.

[0058] Example process

[0059] Figure 3 illustrates a flowchart of an example process 300 for generating video according to some embodiments of the present disclosure. Process 300 can be implemented at electronic device 110.

[0060] As shown in the figure, in box 310, electronic device 110 acquires reference image and audio control signal.

[0061] In box 320, electronic device 110 uses the first model to generate descriptive text for the reference image.

[0062] In box 330, electronic device 110 uses a second model to generate a target video based on a reference image, descriptive text, and control signals. The target video includes motion content associated with a preset object in the reference image, and the motion content corresponds to the audio control signals.

[0063] In some embodiments, generating a target video using a second model based on a reference image, descriptive text, and control signals includes: constructing positive cue words associated with the second model based on the descriptive text; and generating the target video using the second model based on the positive cue words, negative cue words, a reference image, and audio control signals, wherein the negative cue words include preset text content.

[0064] In some embodiments, the second model includes a diffusion model, and generating a target video using the second model based on positive cue words, negative cue words, a reference image, and an audio control signal includes: determining conditional guidance parameters based on a denoising time step of the diffusion model, wherein the conditional guidance parameters are positively correlated with the denoising time step; and generating video features of the target video based on the conditional guidance parameters and feature difference information, wherein the feature difference information indicates the difference between a first video feature corresponding to a positive cue word and a second video feature corresponding to a negative cue word.

[0065] In some embodiments, the diffusion model is configured to determine a denoising gradient at each denoising time step, the denoising gradient including a noise perturbation term.

[0066] In some embodiments, the noise disturbance term is determined based on Langevin dynamics.

[0067] In some embodiments, the positive prompt also includes the input prompt.

[0068] In some embodiments, the second model is trained based on a training dataset, which includes a first data subset and a second data subset, wherein the first data subset includes multiple text-video pairs and the second data subset includes multiple audio-video pairs.

[0069] In some embodiments, the first data subset corresponds to a first proportion of the training dataset, the second data subset corresponds to a second proportion of the training dataset, and the difference between the first proportion and the second proportion is less than a threshold.

[0070] Example devices and equipment

[0071] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. Figure 4 shows a schematic structural block diagram of an example apparatus 400 for generating video according to certain embodiments of this disclosure. Apparatus 400 may be implemented as or included in electronic device 110. The various modules / components in apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0072] As shown in Figure 4, the device 400 includes an acquisition module 410 configured to acquire a reference image and an audio control signal; a text generation module 420 configured to generate descriptive text of the reference image using a first model; and a video generation module 430 configured to generate a target video based on the reference image, the descriptive text, and the control signal using a second model. The target video includes motion content associated with a preset object in the reference image, and the motion content corresponds to the audio control signal.

[0073] In some embodiments, the video generation module 430 is further configured to: construct positive cue words associated with the second model based on descriptive text; and generate a target video using the second model based on the positive cue words, negative cue words, reference images, and audio control signals, wherein the negative cue words include preset text content.

[0074] In some embodiments, the second model includes a diffusion model, and the video generation module 430 is further configured to: determine conditional guidance parameters based on the denoising time step of the diffusion model, wherein the conditional guidance parameters are positively correlated with the denoising time step; and generate video features of the target video based on the conditional guidance parameters and feature difference information, wherein the feature difference information indicates the difference between the first video feature corresponding to the positive prompt word and the second video feature corresponding to the negative prompt word.

[0075] In some embodiments, the diffusion model is configured to determine a denoising gradient at each denoising time step, the denoising gradient including a noise perturbation term.

[0076] In some embodiments, the noise disturbance term is determined based on Langevin dynamics.

[0077] In some embodiments, the positive prompt also includes the input prompt.

[0078] In some embodiments, the second model is trained based on a training dataset, which includes a first data subset and a second data subset, wherein the first data subset includes multiple text-video pairs and the second data subset includes multiple audio-video pairs.

[0079] In some embodiments, the first data subset corresponds to a first proportion of the training dataset, the second data subset corresponds to a second proportion of the training dataset, and the difference between the first proportion and the second proportion is less than a threshold.

[0080] Figure 5 shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 500 shown in Figure 5 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 500 shown in Figure 5 can be used to implement the electronic device 110 of Figure 1.

[0081] As shown in Figure 5, the electronic device 500 is in the form of a general-purpose electronic device. Components of the electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in the memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 500.

[0082] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.

[0083] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0084] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0085] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0086] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0087] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0088] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0089] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0090] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0091] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for generating video, comprising: Acquire reference images and audio control signals; The first model is used to generate descriptive text for the reference image; as well as Using a second model, a target video is generated based on the reference image, the descriptive text, and the control signal. The target video includes motion content associated with a preset object in the reference image, and the motion content corresponds to the audio control signal.

2. The method of claim 1, wherein generating the target video based on the reference image, the descriptive text, and the control signal using the second model comprises: Based on the descriptive text, construct positive prompt words associated with the second model; as well as Using the second model, the target video is generated based on the positive prompt words, negative prompt words, the reference image, and the audio control signal, wherein the negative prompt words include preset text content.

3. The method according to claim 2, wherein the second model includes a diffusion model, and generating the target video using the second model based on the positive cue word, the negative cue word, the reference image, and the audio control signal includes: Based on the denoising time step of the diffusion model, a conditional guidance parameter is determined, wherein the conditional guidance parameter is positively correlated with the denoising time step; Based on the conditional guidance parameters and feature difference information, video features of the target video are generated, wherein the feature difference information indicates the difference between the first video feature corresponding to the positive prompt word and the second video feature corresponding to the negative prompt word.

4. The method of claim 3, wherein the diffusion model is configured as follows: At each denoising time step, a denoising gradient is determined, which includes a noise perturbation term.

5. The method of claim 4, wherein the noise disturbance term is determined based on Langevin dynamics.

6. The method of claim 2, wherein the positive prompt word further includes an input prompt word.

7. The method of claim 1, wherein the second model is trained based on a training dataset, the training dataset comprising a first data subset and a second data subset, wherein the first data subset comprises multiple text-video pairs, and the second data subset comprises multiple audio-video pairs.

8. The method of claim 7, wherein the first data subset corresponds to a first proportion of the training dataset, the second data subset corresponds to a second proportion of the training dataset, and the difference between the first proportion and the second proportion is less than a threshold.

9. An apparatus for generating video, comprising: The acquisition module is configured to acquire reference images and audio control signals; The text generation module is configured to generate descriptive text for the reference image using the first model; as well as The video generation module is configured to generate a target video using a second model based on the reference image, the descriptive text, and the control signal. The target video includes motion content associated with a preset object in the reference image, and the motion content corresponds to the audio control signal.

10. An electronic device, comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 8 when executed by the at least one processor.

11. A computer-readable storage medium having stored thereon computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 8.

12. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 8.