Video generation method and apparatus, device, and medium

By generating videos using a target network model and multiple parallel generative networks, the high cost and low efficiency of animation video generation in existing technologies are solved, enabling personalized and highly controllable video generation and improving user experience.

WO2026001087A1PCT designated stage Publication Date: 2026-01-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/082423
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-26
Filing Date
2025-03-13
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Current technologies for generating animated videos require professional production, which is costly and inefficient. Furthermore, AI-generated methods struggle to meet personalized customization needs and offer poor controllability.

Method used

A target network model is adopted. By acquiring the target cue image and at least two cue texts, multiple parallel generative networks are used to generate a video containing a second target object. Combining conditional injection network and temporal attention mechanism, the accuracy and controllability of the object's actions in the video are ensured.

Benefits of technology

It enables low-cost and efficient generation of personalized videos. The objects in the videos have the external characteristics of the first target object and can perform different actions, which improves the controllability and fun of the generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025082423_02012026_PF_FP_ABST
    Figure CN2025082423_02012026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a video generation method and apparatus, a device, and a medium. The method comprises: acquiring a target prompt image and at least two pieces of prompt text, the target prompt image comprising a first target object, and the at least two pieces of prompt text being used for describing different object actions; and on the basis of the target prompt image and the at least two pieces of prompt text, using a target network model to generate a target video comprising a second target object, the second target object having at least some external features of the first target object, and the target video presenting dynamic pictures of the second target object executing different object actions.
Need to check novelty before this filing date? Find Prior Art

Description

Video generation method, apparatus, device and medium

[0001] Cross-reference to Related Applications

[0002] The present application claims priority to the Chinese patent application No. 202410841770.2, filed on June 26, 2024, and entitled "Video generation method, apparatus, device and medium", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0003] The present disclosure relates to the technical field of multimedia processing, and particularly relates to a video generation method, apparatus, device and medium. BACKGROUND

[0004] With the development of multimedia technology, more and more scenarios need to use videos with strong appeal or interest to improve the visual experience of users. For example, in a live broadcast scenario, a gift animation video can be presented on the terminal interface of all viewers based on a virtual gift selected by a user and presented to a host. In a social scenario, users can also present interesting animation videos to each other. SUMMARY

[0005] The present disclosure provides a video generation method, apparatus, device and medium.

[0006] The present disclosure provides a video generation method, which comprises: obtaining a target prompt image and at least two prompt texts; wherein the target prompt image comprises a first target object, and the at least two prompt texts are used to describe different object actions; based on the target prompt image and the at least two prompt texts, a target video containing a second target object is generated by using a target network model; wherein the second target object has at least part of the external characteristics of the first target object, and the target video presents a dynamic picture of the second target object performing the different object actions.

[0007] Optionally, the target network model comprises a plurality of generation networks, and the plurality of generation networks are arranged in parallel; the target prompt image is used as the input of each generation network, and the at least two prompt texts are used as the input of at least two generation networks, and different prompt texts correspond to different generation networks.

[0008] Optionally, the generating, based on the target prompt image and the at least two prompt texts, a target video containing a second target object by using a target network model comprises: obtaining object text features and prompt text features corresponding to the at least two prompt texts based on the target prompt image and the at least two prompt texts; inputting the object text features into each of the generation networks and inputting the prompt text features corresponding to the at least two prompt texts into different generation networks; generating, by the multiple generation networks, respective output images based on respective input information; wherein the output images contain the second target object, and the second target object in different output images has different action forms; and obtaining, based on a parallel arrangement order of the multiple generation networks and the output images of the multiple generation networks, a target video containing the second target object.

[0009] Optionally, the target network model further comprises a conditional injection network and at least two text encoders; and the obtaining, based on the target prompt image and the at least two prompt texts, object text features and prompt text features corresponding to the at least two prompt texts comprises: obtaining first target object features in the target prompt image by using an object recognition model, and mapping the first target object features into object text features by using the conditional injection network; and respectively encoding the at least two prompt texts by using the at least two text encoders to obtain the prompt text features corresponding to the at least two prompt texts.

[0010] Optionally, the generating, by the multiple generation networks, respective output images based on respective input information comprises: generating, by the multiple generation networks, respective output images based on respective input information, target association information and a parallel arrangement order by using a time sequence attention mechanism; wherein the target association information comprises target features output by the same designated network layer of the multiple generation networks.

[0011] Optionally, the multiple generation networks are divided into a first type of network and a second type of network, the input of the first type of network comprises the prompt texts, and the second type of network is a network other than the first type of network in the multiple generation networks; the second target object in the output image corresponding to the first type of network has a first action form, the first action form corresponds to an object action described by the prompt text of the first type of network; the second target object in the output image corresponding to the second type of network has a second action form; and the second action form corresponding to the second type of network located between two adjacent first type of networks is a gradual change form used to connect the first action forms corresponding to the two adjacent first type of networks.

[0012] Optionally, the plurality of generation networks have the same structure and share parameters.

[0013] Optionally, the method further comprises: performing special effect processing on the target video to obtain a special effect video.

[0014] The embodiments of the present disclosure further provide a video generation device, including: an acquisition module, configured to acquire a target prompt image and at least two prompt texts; wherein the target prompt image includes a first target object, and the at least two prompt texts are used to describe different object actions; and a video generation module, configured to generate a target video containing a second target object based on the target prompt image and the at least two prompt texts by using a target network model; wherein the second target object has at least part of the external characteristics of the first target object, and the target video presents a dynamic picture of the second target object performing the different object actions.

[0015] The embodiments of the present disclosure further provide an electronic device, including: a processor; a memory for storing executable instructions of the processor; and the processor, configured to read the executable instructions from the memory and execute the instructions to implement the video generation method provided by the embodiments of the present disclosure.

[0016] The embodiments of the present disclosure further provide a computer readable storage medium, which stores a computer program, and the computer program is used to execute the video generation method provided by the embodiments of the present disclosure.

[0017] It should be understood that the contents described in this part are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0018] The accompanying drawings incorporated in the specification hereof and forming a part thereof illustrate embodiments consistent with the present disclosure and together with the description serve to explain the principles of the present disclosure.

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced here. Obviously, for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0020] FIG. 1 is a flow diagram of a video generation method according to an embodiment of the present disclosure;

[0021] FIG. 2 is a structural schematic diagram of a target network model according to an embodiment of the present disclosure;

[0022] FIG. 3 is a structural schematic diagram of a target network model according to an embodiment of the present disclosure;

[0023] FIG. 4 is a structural schematic diagram of a generation network according to an embodiment of the present disclosure;

[0024] FIG. 5 is a schematic diagram of a timing attention mechanism according to an embodiment of the present disclosure;

[0025] FIG. 6 is a structural schematic diagram of a video generation apparatus according to an embodiment of the present disclosure;

[0026] FIG. 7 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] As described above, more and more scenarios need to use videos with high infectivity or interest to improve the visual experience of users. However, most of the existing technologies need professional people to make animation videos, and the cost required for generating animation videos is high and the efficiency is low. The way of generating animation videos with the help of artificial intelligence has poor controllability and is difficult to meet the personalized customization demand.

[0028] In order to enable a person skilled in the art to more clearly understand the above-mentioned purposes, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0029] In the following description, many specific details are set forth in order to provide a thorough understanding of the present disclosure, but the present disclosure can also be implemented in other ways different from those described herein; obviously, the embodiments in the description are only some of the embodiments of the present disclosure, not all the embodiments.

[0030] FIG. 1 is a flow schematic diagram of a video generation method according to an embodiment of the present disclosure, which can be executed by a video generation apparatus. The apparatus can be implemented by software and / or hardware, and can be integrated in an electronic device. As shown in FIG. 1, the method mainly includes the following steps S102-S104:

[0031] In step S102, a target prompt image and at least two prompt texts are obtained; wherein the target prompt image includes a first target object, and the at least two prompt texts are used to describe different object actions.

[0032] The first target object is not limited by the embodiments of the present disclosure, such as the first target object can be a person, an animal, a vehicle, a robot, etc. The object action described in the prompt text is not limited by the embodiments of the present disclosure, such as the object action can be opening eyes, closing eyes, lifting head, lowering head, walking, running, etc.

[0033] In step S104, a target video containing a second target object is generated based on the target prompt image and the at least two prompt texts by using a target network model; wherein the second target object has at least part of the external characteristics of the first target object, and the target video presents a dynamic picture of the second target object performing different object actions. For example, assuming that the two input prompt texts are used to describe two actions of opening eyes and closing eyes, the target video can present a dynamic picture of the second target object from opening eyes to closing eyes.

[0034] The structure of the target network model is not limited by the embodiments of the present disclosure, and the target network model is exemplarily a generative model. In actual application, the target network model in the embodiments of the present disclosure can be trained by using a basic diffusion model and a LoRA (Low-Rank Adaptation) model, and the target network model can generate an object image with a specific style. Taking a person as the first target object, the target network model can generate an image containing a virtual person image with a specific style, such as a cartoon virtual person, a virtual person in ancient costume, etc. with the external characteristics of the first target object. Specifically, the parameters of the basic diffusion model can be adjusted based on the LoRA technology. The LoRA technology is a light-weight model fine-tuning technology, which can assist model training by inserting some network layers (which can be referred to as LoRA layers) in the basic diffusion model. The LoRA technology can customize and fine-tune a large basic model with a small amount of sample data with specific objects, so as to generate a generative model with specific object generation capability. In addition, the embodiments of the present disclosure can add specific functional modules such as a time sequence attention module to the original network structure of the generative model, and generate a plurality of images with a gradual change effect of the action of the second target object based on the time sequence attention mechanism by using the functional module, so as to form a target video.

[0035] It should be noted that in related generative artificial intelligence technology, only one prompt text is provided for the model, while at least two prompt texts are obtained in the embodiments of the present disclosure, which helps the model to controllably and reliably generate an object animation video presenting different object actions based on the at least two prompt texts.

[0036] The technical solution provided by the embodiments of the present disclosure can efficiently and conveniently generate a target video containing a second target object by using a target network model based on a target prompt image (including a first target object) and at least two prompt texts (used to describe different object actions), without the need for professionals to make an animation video, thereby reducing the generation cost of the animation video and improving the generation efficiency of the animation video. Moreover, under the constraint of the target prompt image, the second target object in the target video has at least part of the external characteristics of the first target object, and under the constraint of the at least two prompt texts, the second target object can perform different object actions. The above-mentioned manner can realize the effect of generating a personalized customized video for the first target object, and through the constraint of the target prompt image and the at least two prompt texts, the video generation process also has strong controllability, and an animation video that can present a second target object similar in external characteristics to the first target object is generated, which has strong interest and appeal to users and can better meet the user demand.

[0037] In some embodiments, the target network model contains a plurality of generation networks, and the plurality of generation networks are arranged in parallel; the target prompt image is used as the input of each generation network, the at least two prompt texts are used as the input of at least two generation networks, and different prompt texts correspond to different generation networks. In some specific implementation examples, the structures of the plurality of generation networks are the same, and the plurality of generation networks share parameters, which is more convenient for training and information interaction and can better control the matching degree of the output image. The input of each generation network also includes a noise image, and the noise images of different generation networks are different. The noise image corresponding to each generation network can be randomly generated and can be implemented by using related technologies, which will not be described here.

[0038] For ease of understanding, refer to the structure diagram of a target network model shown in FIG. 2, which shows the parallel generation network 1, generation network 2, …, generation network 10 to generation network N. Each generation network has an output image, and all the output images arranged in order are the target video. As shown in FIG. 2, the input of each generation network contains a target prompt image and a corresponding noise image, but only the input of part of the generation networks includes a prompt text, such as the generation network 2, which can be regarded as that the input content of the prompt text is empty or default. Of course, FIG. 2 is only illustrative, and in actual application, a corresponding prompt text can also be set for each generation network. In addition, in the case where the number of prompt texts is less than the number of generation networks, the generation network whose input contains the prompt text can be flexibly specified according to the demand, which is not limited here.

[0039] On the basis of the foregoing, the step S104, i.e., the step of generating a target video containing a second target object based on the target prompt image and the at least two prompt texts by using the target network model, can be performed according to the following steps A to C.

[0040] Step A: based on the target prompt image and the at least two prompt texts, obtaining an object text feature and prompt text features corresponding to the at least two prompt texts respectively; in some specific implementation examples, the target network model further includes a conditional injection network and at least two text encoders, and the step A can be performed according to the following steps A1 to A2.

[0041] Step A1: obtaining a first target object feature in the target prompt image by using an object recognition model, and mapping the first target object feature to the object text feature by using the conditional injection network to inject the object text feature into the generation network. The structure of the conditional injection network is not limited in the embodiments of the present disclosure, such as that it can contain functional units such as feature mapping units, the conditional injection network can not only perform further feature extraction on the received data, but also perform processing such as feature mapping, map the object feature (i.e., the first target object feature) to the prompt text feature, and inject the prompt text feature as conditional information into the generation process of the target network, which can guide the image generation to a certain extent and guarantee that the second target object contained in the output image has similar external features to the first target object.

[0042] Step A2: encoding the at least two prompt texts by using the at least two text encoders respectively to obtain the prompt text features corresponding to the at least two prompt texts respectively.

[0043] Step B: inputting the object text feature into each generation network, and inputting the prompt text features corresponding to the at least two prompt texts into different generation networks respectively; generating output images by the multiple generation networks based on the input information respectively; wherein, the output images contain the second target object, and the second target object presented in different output images has different action forms.

[0044] In addition, it should be noted that the target network model provided by the embodiments of the present disclosure has two obvious differences compared with the existing conventional video generation model or text-to-image model: (1) It is not limited to generating images or videos from text, but introduces a target prompt image based on the prompt text, so as to inject the characteristics of the first target object into the target network model, which helps to achieve personalized customization effect for the first target object. (2) The existing technology basically provides a prompt text for the generation model. Even if the generation model has multiple generation networks, the multiple generation networks all share the same prompt text. However, the embodiments of the present disclosure can obtain at least two prompt texts for describing the actions of different objects, and can be input into different generation networks. Some generation networks can not have a prompt text, but can learn the content to be generated by themselves based on the feature information of other networks and according to the time sequence attention mechanism and the like. The above-mentioned manner can effectively guarantee the accuracy and controllability of the object actions presented in the generated animation video.

[0045] For ease of understanding, with reference to FIG. 3, a structural schematic diagram of a target network model is shown, which is based on FIG. 2. In FIG. 3, the object recognition model, the conditional injection network and the text encoder are further illustrated. Exemplarily, the conditional injection network can include an Adapter unit. The conditional injection network can map the received first target object features to the corresponding feature space. The mapped features (i.e., object text features) can be injected into the image generation process of the generation network through part of the network in the generation network, such as the cross-attention unit of the denoising network (e.g., Unet) in the generation network, or the LoRA layer added in the generation network. In this way, the conditional injection network can inject the first target object features and other conditional information into the image generation process, so that the second target object contained in the output image of the generation network can reflect at least part of the external characteristics of the first target object to a certain extent. Similarly, by inputting the prompt text features into the generation network, the second target object contained in the output image of the generation network can be effectively constrained to perform the object action described in the prompt text.

[0046] In some specific implementation examples, the multiple generation networks generate respective output images based on respective corresponding input information, target association information and parallel arrangement order by using the time sequence attention mechanism; wherein the target association information includes target features output by the same specified network layer corresponding to each of the multiple generation networks. In actual application, the same specified network layer corresponding to each of the multiple generation networks can be, for example, a certain network layer of the same Unet unit of the multiple generation networks. The features at each position in the feature information output by the specified network layer or the features at the specified position can all be used as the target features.

[0047] For ease of understanding, with reference to a structural schematic diagram of a generation network shown in FIG. 4, it is illustrated that the generation network includes an encoder, a denoising network and a decoder, wherein the input of the encoder is a noisy image, the output of the decoder is an image, and it is also illustrated that the features of the conditions injected into the network output can be injected into each Unet unit. In addition, in actual application, if the input of the generation network contains prompt text, the features of the prompt text will also be input into each Unet unit, which is not illustrated in FIG. 4. In FIG. 4, the denoising network is taken as an example of Unet network, which includes multiple Unet units (Unet Block), and the difference from the conventional Unet network is that a time sequence attention module is inserted, which mainly performs information processing based on the time sequence attention mechanism. Specifically, the time sequence attention mechanism is used to realize information interaction between N generation networks, so that each generation network can determine the features it needs to generate based on the features of other generation networks it has obtained, thereby ensuring that the output images corresponding to the N generation networks can realize the animation effect of gradual change of the second target object action in sequence, such as realizing the dynamic effect from opening eyes, gradually closing upper and lower eyelids to completely closing eyes. As can be seen from FIG. 4, in the N parallel generation networks, the time sequence attention modules with the same position can perform time sequence attention processing based on their respective corresponding target association information, arrangement order of the target network and other information. Taking the target association information 1 as an example, it contains the feature information output by the same specified network layer of the first module (the first Unet unit and / or the first time sequence attention module) of the N generation networks. The positions of different time sequence attention modules are different, and the corresponding target association information is different. In actual application, the time sequence attention module can be implemented by related technologies, for example, it can be implemented by an Animate diff module (animation difference module), which can be integrated into models such as generation networks, and learn motion priors from video data sets, so that the image generation model can be extended to a video generation model.

[0048] Specifically, the time sequence attention mechanism allows the generation network in the target network model to adaptively focus on the time sequence information when processing the N sequence data, capture the dependency and context information in the sequence data, and assign different weights to the features with the same position in the N sequence data. Referring to a principle diagram of a time sequence attention mechanism shown in FIG. 5, for each generation network, the time sequence attention mechanism can be used to obtain the target features corresponding to the same specified network layer output of each generation network. For example, the features with the same position in the output features of the specified network layer of the i-th Unet unit of each generation network are assigned with corresponding weights and are subjected to weighted fusion processing. In FIG. 5, the features with the same position in the output features of the specified network layer of the i-th Unet unit of different generation networks are shown in gray lattice, i.e., the aforementioned target features. In actual application, each feature in the 3*3 features shown in FIG. 5 can be sequentially used as the target feature and is subjected to weighted fusion processing in combination with the corresponding features of other generation networks, so as to obtain the features corresponding to the corresponding position of the generation network based on the weighted fusion features, thereby ensuring the rationality of the output image of the generation network. For example, the output image of the generation network 1 is a person opening eyes image, the output image of the generation network 10 is a person closing eyes image, and the distances between the upper and lower eyelids of the output images of the generation networks 2 to 9 are different, such as gradually decreasing, thereby presenting the action gradient effect of the person gradually transitioning from opening eyes to closing eyes. It can be understood that although only part of the network inputs contain prompt texts for describing actions, such as the prompt text of the generation network 1 describing the opening eyes action and the prompt text of the generation network 10 describing the closing eyes action, for the generation networks 2 to 9 whose inputs do not contain prompt texts (which can be understood as empty prompt texts or default prompt texts), based on the obtained target association information of other generation networks and the arrangement order of the multiple networks, the time sequence attention mechanism can be used to reasonably and reliably determine the content required to be output by itself, thereby generating an animation video with action continuity.

[0049] Step C: obtaining a target video containing a second target object based on the parallel arrangement order of the multiple generation networks and the output images of the multiple generation networks.

[0050] Each generation network has a corresponding arrangement serial number, and the parallel arrangement order of the multiple generation networks can be known based on the arrangement serial numbers of the generation networks. Then, the target video can be obtained by sequentially arranging the output images of the multiple generation networks according to the serial numbers.

[0051] In some embodiments, the plurality of generation networks are divided into a first type of network and a second type of network, an input of the first type of network includes the prompt text, and the second type of network is a network other than the first type of network in the plurality of generation networks; in other words, the input of the second type of network does not include the prompt text, and the input content corresponding to the prompt text is empty or the prompt text is default. For example, the generation network 1, the generation network 10, and the generation network N in FIG. 2 are the first type of network, and the generation network 2 to the generation network 9 are the second type of network.

[0052] The second target object in the output image corresponding to the first type of network has a first action form, and the first action form corresponds to the object action described in the prompt text of the first type of network; the second target object in the output image corresponding to the second type of network has a second action form; and the second action form corresponding to the second type of network located between two adjacent first type of networks is a gradual form used to connect the first action forms corresponding to the two adjacent first type of networks. For example, the prompt text of the generation network 1 describes an open-eye action, and the first action form is an open-eye form; the prompt text of the generation network 10 describes a close-eye action, and the corresponding first action form is a close-eye form; and the second action forms generated by the generation network 2 to the generation network 9 are intermediate forms gradually connecting from open-eye to close-eye. The specific implementation principle can be referred to the foregoing related content, and will not be described here.

[0053] In some embodiments, the video generation method provided by the embodiments of the present disclosure further includes: performing special effect processing on the target video to obtain a special effect video. The special effect processing includes but is not limited to speed processing, mask processing, etc., and can be determined according to the application scene and user demand of the target video, and the special effect processing mode is not limited here.

[0054] The application scenarios of the target video of the embodiments of the present disclosure are not limited, such as live broadcast scenarios, social scenarios, and the like. Taking a live broadcast scenario as an example, the first target object can be an anchor in a target live broadcast room, and the target prompt image is an anchor image. Then, at least two prompt texts for describing actions can be pre-set. After that, a target network model can be used to generate an animation video of a virtual object having external characteristics of the anchor, and the virtual object can perform the actions described in the prompt texts. The animation video can be used as a customized gift video for the anchor. When a specific gift sending request is initiated by a viewer, the animation video can be played on the interface of all user terminals in the target live broadcast room, presenting the viewer's visual effect of sending a customized image gift to the anchor, and enhancing the interest and appeal of the live broadcast. It can be understood that, if the prior art is used, only professionals can make personalized animation videos for anchors, which requires a high cost and is low in efficiency. In the current situation of a large number of anchors, it is impossible to customize animation videos for all anchors. However, by using the above method provided by the embodiments of the present disclosure, a corresponding personalized customized animation can be quickly and conveniently generated for the anchor, the object actions in the animation can be adjusted according to requirements, and the user requirements can be better met. In other words, one or more embodiments of the present disclosure can reduce the cost of generating an animation video and improve the video production efficiency, and also make the video generation process controllable, generate an animation video of a second target object that can present similar external characteristics to the first target object, and have strong interest and appeal to the user, which can better meet the user requirements.

[0055] Corresponding to the foregoing video generation method, the embodiments of the present disclosure further provide a video generation apparatus. FIG. 6 is a structural schematic diagram of a video generation apparatus provided by an embodiment of the present disclosure. The apparatus can be implemented by software and / or hardware, and can be integrated in an electronic device. As shown in FIG. 6, the video generation apparatus includes:

[0056] The acquisition module 602 is configured to acquire a target prompt image and at least two prompt texts. The target prompt image includes a first target object, and the at least two prompt texts are used to describe different object actions.

[0057] The video generation module 604 is configured to generate a target video containing a second target object by using a target network model based on the target prompt image and the at least two prompt texts. The second target object has at least part of the external characteristics of the first target object, and the target video presents dynamic pictures of the second target object performing different object actions.

[0058] The device provided by the embodiments of the present disclosure can efficiently and conveniently generate a target video containing a second target object by using a target network model based on a target prompt image (containing a first target object) and at least two prompt texts (used to describe different object actions), without the need for professionals to make the video, thereby reducing the video generation cost and improving the video generation efficiency. Moreover, under the constraint of the target prompt image, the second target object in the target video has at least part of the external characteristics of the first target object, and under the constraint of the at least two prompt texts, the second target object can perform different object actions. The above-mentioned manner can realize the effect of generating a personalized customized video for the first target object, and through the constraint of the target prompt image and the at least two prompt texts, the video generation process also has strong controllability, and an animation video containing a second target object similar to the first target object in external characteristics is generated, which has strong interest and appeal to users and can better meet the user demand.

[0059] In some embodiments, the target network model contains a plurality of generation networks, and the plurality of generation networks are arranged in parallel; the target prompt image is used as the input of each generation network, and the at least two prompt texts are used as the input of at least two generation networks, and different prompt texts correspond to different generation networks.

[0060] In some embodiments, the video generation module 604 is specifically configured to: based on the target prompt image and the at least two prompt texts, obtain object text features and prompt text features corresponding to the at least two prompt texts respectively; input the object text features into each generation network, and input the prompt text features corresponding to the at least two prompt texts respectively into different generation networks; generate respective output images by the plurality of generation networks based on respective input information; wherein the output images contain the second target object, and the second target object in different output images has different action forms; based on the parallel arrangement order of the plurality of generation networks and the respective output images of the plurality of generation networks, obtain a target video containing a second target object.

[0061] In some embodiments, the video generation module 604 is specifically configured to: obtain first target object features in the target prompt image by using an object recognition model, and map the first target object features to object text features by using the conditional injection network; encode the at least two prompt texts respectively by using the at least two text encoders to obtain prompt text features corresponding to the at least two prompt texts respectively.

[0062] In some embodiments, the video generation module 604 is specifically configured to: generate, by the plurality of generation networks, the respective output images based on the respective corresponding input information, target association information, and parallel arrangement order by using a time sequence attention mechanism; and wherein the target association information comprises target features output by the same designated network layer of the plurality of generation networks.

[0063] In some embodiments, the plurality of generation networks are divided into a first type of network and a second type of network, the input of the first type of network comprises the prompt text, and the second type of network is a network of the plurality of generation networks other than the first type of network; the second target object in the output image corresponding to the first type of network has a first action form, the first action form corresponds to the object action described in the prompt text of the first type of network; the second target object in the output image corresponding to the second type of network has a second action form; and the second action form corresponding to the second type of network located between two adjacent first type of networks is a gradual change form used to connect the first action forms corresponding to the two adjacent first type of networks.

[0064] In some embodiments, the plurality of generation networks have the same structure, and the plurality of generation networks share parameters; the input of each of the plurality of generation networks further comprises a noise map, and the noise maps of different generation networks are different.

[0065] In some embodiments, the apparatus further comprises a special effect processing module configured to perform special effect processing on the target video to obtain a special effect video.

[0066] The video generation apparatus provided by the embodiments of the present disclosure can perform the video generation method provided by any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of performing the method.

[0067] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the apparatus embodiments described above can refer to the corresponding process in the method embodiments, which will not be repeated here.

[0068] The embodiments of the present disclosure provide an electronic device, which comprises a storage device having a computer program stored thereon, and a processing device configured to execute the computer program in the storage device to implement the steps of any method in the present disclosure.

[0069] Reference is now made to FIG. 7, which shows a structural diagram of an electronic device 700 suitable for implementing embodiments of the present disclosure. The terminal device in embodiments of the present disclosure can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a car terminal (e.g., a car navigation terminal), and the like, as well as a stationary terminal such as a digital TV, a desktop computer, and the like. The electronic device shown in FIG. 7 is merely an example, and should not impose any limitation on the functions and use range of embodiments of the present disclosure.

[0070] As shown in FIG. 7, the electronic device 700 can include a processing device (e.g., a central processor, a graphic processor, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded into a random access memory (RAM) 703 from a storage device 708. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0071] In general, the following devices can be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage device 708 including, for example, a magnetic tape, a hard disk, and the like; and a communication device 709. The communication device 709 can allow the electronic device 700 to communicate with other devices wirelessly or via a wire to exchange data. Although FIG. 7 shows the electronic device 700 having various devices, it should be understood that all of the shown devices are not required to be implemented or possessed. More or fewer devices can be alternatively implemented or possessed.

[0072] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product including a computer program carried on a non-transitory computer readable medium, the computer program containing program code for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-described functions defined in the methods of embodiments of the present disclosure are performed.

[0073] In addition to the method and device described above, the embodiments of the present disclosure can also be a computer program product, which includes computer program instructions that make the processor execute the image processing method provided by the embodiments of the present disclosure when the processor runs. The computer program product can be written in any combination of one or more programming languages to execute the program code of the embodiments of the present disclosure, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" language or similar programming languages. The program code can be executed entirely on a user computing device, partially on a user device, as an independent software package, partially on a user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0074] In addition, the embodiments of the present disclosure can also be a computer readable storage medium, which stores computer program instructions, and the computer program instructions make the processor execute the video generation method provided by the embodiments of the present disclosure when the processor runs.

[0075] The computer readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples (non-exhaustive list) of readable storage medium include: electrical connection with one or more conductive wires, portable disk, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the above.

[0076] The embodiments of the present disclosure also provide a computer program product, which includes computer programs / instructions that are executed by a processor to implement the video generation method in the embodiments of the present disclosure.

[0077] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained according to relevant laws and regulations.

[0078] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed by the user will need to acquire and use personal information of the user. Thus, the user can autonomously select whether to provide the personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium performing the operation of the technical solution of the present disclosure according to the prompt information.

[0079] As an optional but non-limiting implementation, in response to receiving an active request of a user, the manner of sending a prompt information to the user may, for example, be a pop-up window manner, and the prompt information may be presented in the pop-up window in the form of text. In addition, the pop-up window may also carry a selection control for the user to select “agree” or “disagree” to provide personal information to the electronic device.

[0080] It can be understood that the above notification and acquisition of user authorization process is only illustrative and does not limit the implementation of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0081] It should be noted that in this document, relational terms such as “first” and “second”, and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms “comprises”, “comprising”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without more limitations, an element preceded by “comprises a...” does not, without more limitations, foreclose the existence of additional identical elements in the process, method, article, or apparatus that includes the recited element.

[0082] The above description is merely one specific implementation of the present disclosure, which enables those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A video generation method, comprising: Acquire a target prompt image and at least two prompt texts; wherein the target prompt image includes a first target object, and the at least two prompt texts are used to describe different object actions; Based on the target prompt image and the at least two prompt texts, a target video containing a second target object is generated using a target network model; wherein the second target object has at least some of the external features of the first target object, and the target video presents dynamic images of the second target object performing the different object actions.

2. The method according to claim 1, wherein the target network model comprises multiple generator networks, and the multiple generator networks are arranged in parallel; the target prompt image is used as input to each generator network, and the at least two prompt texts are used as input to at least two generator networks, and different prompt texts correspond to different generator networks.

3. The method according to claim 2, wherein generating a target video containing a second target object using a target network model based on the target cue image and the at least two cue texts includes: Based on the target prompt image and the at least two prompt texts, obtain the object text features and the prompt text features corresponding to each of the at least two prompt texts; The object text features are input into each of the generation networks, and the prompt text features corresponding to each of the at least two prompt texts are input into different generation networks; the multiple generation networks generate their respective output images based on their respective input information; wherein, the output image contains the second target object, and the action form of the second target object presented in different output images is different; Based on the parallel arrangement of the multiple generator networks and the output images of each generator network, a target video containing the second target object is obtained.

4. The method according to claim 3, wherein the target network model further comprises a conditional injection network and at least two text encoders; the step of obtaining object text features and corresponding prompt text features for each of the at least two prompt texts based on the target prompt image and the at least two prompt texts includes: The first target object features in the target prompt image are obtained using an object recognition model, and the first target object features are mapped to object text features through the conditional injection network; The at least two text encoders are used to encode the at least two prompt texts respectively to obtain the prompt text features corresponding to each of the at least two prompt texts.

5. The method according to claim 3, wherein generating respective output images through the plurality of generation networks based on their respective input information comprises: The multiple generator networks generate their respective output images using a temporal attention mechanism based on their respective input information, target association information, and parallel arrangement order; wherein, the target association information includes target features output by the same specified network layer corresponding to each of the multiple generator networks.

6. The method according to claim 2, wherein the plurality of generating networks are divided into a first type of network and a second type of network, the input of the first type of network includes the prompt text, and the second type of network is the network other than the first type of network among the plurality of generating networks; The second target object in the output image corresponding to the first type of network has a first action form, and the first action form corresponds to the object action described by the prompt text of the first type of network. The second target object in the output image corresponding to the second type of network has a second action form; and the second action form corresponding to the second type of network located between two adjacent first type networks is a gradient form used to connect the first action forms corresponding to the two adjacent first type networks.

7. The method according to claim 2, wherein the plurality of generating networks have the same structure and share parameters; the input of each of the plurality of generating networks further includes a noise map, and the noise maps of different generating networks are different.

8. The method according to claim 1, wherein the method further comprises: The target video is processed with special effects to obtain a special effects video.

9. A video generation apparatus, comprising: An acquisition module is used to acquire a target prompt image and at least two prompt texts; wherein, the target prompt image includes a first target object, and the at least two prompt texts are used to describe different object actions; The video generation module is used to generate a target video containing a second target object based on the target prompt image and the at least two prompt texts using a target network model; wherein the second target object has at least some of the external features of the first target object, and the target video presents a dynamic scene of the second target object performing the different object actions.

10. An electronic device, wherein the electronic device comprises: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the video generation method according to any one of claims 1-8.

11. A computer-readable storage medium, wherein the storage medium stores a computer program for performing the video generation method according to any one of claims 1-8.

Citation Information

Patent Citations

  • End-to-end animation generation method and device and electronic equipment

    CN110047121A

  • Animation generation method and device, computer readable storage medium and terminal

    CN116309965A

  • Video generation method, video model training method and electronic equipment

    CN117633296A

  • Video generation method and device, readable medium and electronic equipment

    CN117896592A

  • Video generation method and device, computer equipment and medium

    CN118138854A