Video generation method and device, equipment and medium
By adopting a multi-stage processing method in the video generation network architecture, the problems of low generation efficiency and unstable spatial and temporal consistency of existing video generation methods are solved, and efficient and stable video generation is achieved.
Patent Information
- Application Number
- CN202510534037.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The existing video generation methods have the problem of low generation efficiency, especially the model lacks the ability to understand the real physical world, which leads to unstable spatial and spatial consistency of the generated videos.
A video generation method is adopted to perform multi-stage processing through video generation network architecture, including a first encoder, a reference subnet, a decomposition subnet, a second encoder, a denoising subnet, a decoder and a motion prediction subnet. The specific steps include: encoding the reference image to obtain the potential vector, using text prompt words to query to obtain the text vector, obtain the first vector through attention calculation, decomposing the image and text to obtain the key stage text, encoding the key stage text to obtain the key vector, performing denoising processing to obtain the second vector, decoding to obtain the key frame, and converting the motion prediction subnet and stitching the frame to obtain the target video.
Through multi-stage processing, the generation efficiency of video generation is improved, the time and space consistency of the video subject is ensured, the understanding of the real physical world is enhanced, the difficulty of complete video generation is reduced, the video quality is maintained, and the artifacts are reduced.
Smart Images

Figure CN120186433A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular, to a video generation method, apparatus, device, and medium. Background Art
[0002] In recent years, in computer vision, conditional graph-to-video generation has gradually become a research hotspot, and graph-to-video methods have achieved great success. However, these methods cannot well solve the problem of unstable spatio-temporal consistency of the main body in the generated video, and the model lacks the ability to understand the real physical world, resulting in low generation efficiency.
[0003] Therefore, the existing video generation methods have the problem of low generation efficiency. Summary of the Invention
[0004] Embodiments of the present invention provide a video generation method, apparatus, device, and medium, aiming to solve the problem of low generation efficiency existing in the existing video generation methods.
[0005] In a first aspect, an embodiment of the present invention provides a video generation method, which is applied to a video generation network architecture. The video generation network architecture includes a first encoder, a reference sub-network, a decomposition sub-network, a second encoder, a denoising sub-network, a decoder, and a motion prediction sub-network. The video generation method includes: Inputting the obtained reference image into the first encoder for encoding to obtain a latent vector; Performing a query process on the obtained text prompt in a preset dictionary library to obtain a text vector; Calculating the attention of the latent vector and the text vector by using the reference sub-network to obtain a first vector; Inputting the reference image and the text prompt into the decomposition sub-network for decomposition to obtain a plurality of key stage texts; Inputting the plurality of key stage texts into the second encoder for encoding to obtain a plurality of key vectors; Performing denoising on the plurality of key vectors and the first vector by using the denoising sub-network to obtain a plurality of second vectors; Inputting the plurality of second vectors into the decoder for decoding to obtain a plurality of key frames; Performing conversion on the plurality of key frames by using the motion prediction sub-network to obtain a plurality of converted frames; Performing splicing on the plurality of converted frames to obtain a target video.
[0006] Second aspect, an embodiment of the present invention further provides a video generation device, which is applied to a video generation network architecture. The video generation network architecture includes a first encoder, a reference sub-network, a decomposition sub-network, a second encoder, a denoising sub-network, a decoder, and a motion prediction sub-network; the video generation device includes: A first encoding unit, configured to input the acquired reference image into the first encoder for encoding to obtain a latent vector; A query unit, configured to perform a query process on the acquired text prompt in a preset dictionary library to obtain a text vector; A calculation unit, configured to perform an attention calculation on the latent vector and the text vector by using the reference sub-network to obtain a first vector; A decomposition unit, configured to input the reference image and the text prompt into the decomposition sub-network for decomposition to obtain a plurality of key stage texts; A second encoding unit, configured to input the plurality of key stage texts into the second encoder for encoding to obtain a plurality of key vectors; A denoising unit, configured to perform denoising on the plurality of key vectors and the first vector by using the denoising sub-network to obtain a plurality of second vectors; A decoding unit, configured to input the plurality of second vectors into the decoder for decoding to obtain a plurality of key frames; A conversion unit, configured to perform a conversion process on the plurality of key frames by using the motion prediction sub-network to obtain a plurality of converted frames; A splicing unit, configured to splice the plurality of converted frames to obtain a target video.
[0007] Third aspect, an embodiment of the present invention further provides an electronic device, which includes a memory and a processor. A computer program is stored on the memory. When the processor executes the computer program, the method described in the first aspect above is implemented.
[0008] Fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium. The storage medium stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the method described in the first aspect above can be implemented.
[0009] The present invention provides a video generation method, apparatus, device, and medium. The video generation method is applied to a video generation network architecture, which includes a first encoder, a reference sub-network, a decomposition sub-network, a second encoder, a denoising sub-network, a decoder, and a motion prediction sub-network. The video generation method includes: inputting the acquired reference image into the first encoder for encoding to obtain a latent vector; querying the acquired text prompt in a preset dictionary library to obtain a text vector; calculating the attention of the latent vector and the text vector by using the reference sub-network to obtain a first vector; inputting the reference image and the text prompt into the decomposition sub-network for decomposition to obtain a plurality of key stage texts; inputting the plurality of key stage texts into the second encoder for encoding to obtain a plurality of key vectors; denoising the plurality of key vectors and the first vector by using the denoising sub-network to obtain a plurality of second vectors; inputting the plurality of second vectors into the decoder for decoding to obtain a plurality of key frames; converting the plurality of key frames by using the motion prediction sub-network to obtain a plurality of converted frames; and splicing the plurality of converted frames to obtain a target video. It can be seen that the embodiments of the present invention implement multi-stage processing of the reference image and the text prompt through the first encoder, the reference sub-network, the decomposition sub-network, the second encoder, the denoising sub-network, the decoder, and the motion prediction sub-network, so as to ensure the spatio-temporal consistency of the video subject, increase the understanding ability of the real physical world, reduce the difficulty of generating a complete video, maintain the stability of the video quality, reduce artifacts, and thus improve the generation efficiency. Description of the Drawings
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 It is a flowchart of the video generation method provided by the embodiments of the present invention; Figure 2 It is a schematic block diagram of the video generation apparatus provided by the embodiments of the present invention; Figure 3 It is a schematic block diagram of the electronic device provided by the embodiments of the present invention. Detailed Embodiments
[0012] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0013] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0014] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0015] It should also be further understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. An embodiment of the present invention provides a video generation method, apparatus, device and medium. The video generation method can be applied to a video generation network architecture, which includes a first encoder, a reference sub-network, a decomposition sub-network, a second encoder, a denoising sub-network, a decoder and a motion prediction sub-network. The video generation method and the video generation network architecture are specifically applicable to the application scenarios of digital marketing for video material generation in various fields. For example, it can be video material generation in fields such as the medical and health field, the food field or the automotive field, etc. The present invention encodes the obtained reference image by inputting it into the first encoder to obtain a latent vector; queries the obtained text prompt in a preset dictionary library to obtain a text vector; calculates the attention of the latent vector and the text vector by using the reference sub-network to obtain a first vector; inputs the reference image and the text prompt into the decomposition sub-network for decomposition processing to obtain a number of key stage texts; inputs the number of key stage texts into the second encoder for encoding processing to obtain a number of key vectors; denoises the number of key vectors and the first vector by using the denoising sub-network to obtain a number of second vectors; inputs the number of second vectors into the decoder for decoding processing to obtain a number of key frames; converts the number of key frames by using the motion prediction sub-network to obtain a number of converted frames; and splices the number of converted frames to obtain a target video, so as to improve the generation efficiency of video generation. The present invention will be described in detail below through specific embodiments.
[0016] Figure 1 It is a schematic flowchart of the video generation method provided by the embodiment of the present invention. As Figure 1 shown, the method includes the following steps S110-S190.
[0017] S110. Input the obtained reference image into the first encoder for encoding processing to obtain a latent vector.
[0018] In this embodiment, the reference image refers to the image input by the user. Specifically, the reference image input by the user can be obtained from the terminal. For example, it can be an image of drugs, medical devices, etc. related to medical and health, or an image of snacks, beverages, etc. related to food; the terminal can be, but is not limited to, various personal computers, laptop computers, smartphones, tablet computers and portable wearable devices; the first encoder can be a VAE-Encoder; specifically, the first encoder is used to encode the reference image to obtain the latent vector.
[0019] In one embodiment, inputting the obtained reference image into the first encoder for encoding processing to obtain a latent vector includes: Inputting the reference image into the first encoder for encoding processing to obtain a mean vector and a variance vector; Using the mean vector and the variance vector as the latent vector.
[0020] In this embodiment, the function of encoding the reference image using the first encoder is to represent the reference image with a low-dimensional latent vector to save computing resources. For example, the data format of the reference image is 224*224*3 = 150528, and after encoding processing by the first encoder (VAE-Encoder), only a 256-dimensional mean vector and a 256-dimensional variance vector are needed to represent the reference image.
[0021] S120. Use the obtained text prompt to query in a preset dictionary library to obtain a text vector.
[0022] In this embodiment, the text prompt refers to a text description input by the user. For example: I have a guiding text for video generation, "Lift one end of the medicine box and then put it down". This prompt will guide the generation of a video from a given picture and describe this process with the states of different stages; where the picture refers to the reference image.
[0023] Specifically, use the text prompt as a query condition to query in the preset dictionary library to obtain the text vector.
[0024] In one embodiment, before using the obtained text prompt to query in a preset dictionary library to obtain a text vector, it further includes: Pre-training the dictionary library using the rule that each word corresponds to a vector to obtain the preset dictionary library.
[0025] In this embodiment, each word can refer to a single character or multiple characters; each word corresponds to a vector one by one; the dictionary library can refer to a database storing several characters or words.
[0026] Specifically, pre-train the dictionary library using the rule that each word corresponds to a vector to obtain the preset dictionary library. Therefore, after pre-training the preset dictionary library, the text prompt can use the preset dictionary library for fast query processing to obtain the text vector, realizing fast conversion of the text prompt and improving the accuracy and generation efficiency of the data.
[0027] S130. Use the reference sub-network to perform attention calculation on the latent vector and the text vector to obtain a first vector.
[0028] In this embodiment, the reference sub-network may be ReferenceNet, that is, use ReferenceNet to perform attention calculation on the latent vector and the text vector to obtain the first vector.
[0029] In one embodiment, the step of using the reference sub-network to perform attention calculation on the latent vector and the text vector to obtain a first vector includes: Input the latent vector into the first attention formula for calculation to obtain a first intermediate vector; Input the first intermediate vector and the text vector into the second attention formula for calculation to obtain a second intermediate vector; Use the first attention formula and the second attention formula to perform a preset number of calculations on the second intermediate vector and the text vector to obtain the first vector.
[0030] In this embodiment, when inputting the latent vector into the first attention formula for calculation to obtain a first intermediate vector, the first attention formula may be: ; Where, is the spatial attention score, that is, the first intermediate vector; represents the Sigmoid operation; , which is used to compress the output to between 0 and 1, representing the attention weight; represents a 7×7 convolution operation; represents performing an average pooling operation on the input to obtain global average information; represents performing a max pooling operation on the input data to obtain global maximum information; the objective of the first attention formula is to strengthen important spatial regions and suppress unimportant regions.
[0031] When inputting the first intermediate vector and the text vector into the second attention formula for calculation to obtain a second intermediate vector, the second attention formula may be: ; Where, is the second intermediate vector, , , ; Z is the latent vector of the reference image; is the text vector of the text prompt, are three matrices learned through training; d is the dimension of the key vector K.
[0032] Performing a preset number of calculation processes on the second intermediate vector and the text vector by using the first attention formula and the second attention formula to obtain the first vector. Specifically, the preset number can be a value such as 11 or 12. Specifically, inputting the second intermediate vector into the first attention formula for attention calculation processing to obtain an updated second intermediate vector, inputting the updated second intermediate vector and the text vector into the second attention formula for attention calculation processing to obtain a processed second intermediate vector, and inputting the processed second intermediate vector into the first attention formula for the next attention calculation processing, and so on, to complete the attention calculation processing of the preset number of times. Wherein, one attention calculation process of the preset number of times may include the calculation process of the first attention formula and the calculation process of the second attention formula.
[0033] Therefore, the latent vector and the text vector serve as references to each other in the attention calculation process using the first attention formula and the second attention formula to ensure the spatio-temporal consistency of the main body between the subsequent obtained key frames, thereby ensuring the accuracy of the data and improving the generation efficiency.
[0034] S140: Input the reference image and the text prompt into the decomposition sub-network for decomposition processing to obtain a number of key stage texts.
[0035] In this embodiment, the text prompt is used by the user to describe the requirements for the generated video content; the text prompt and the reference image can be used as control conditions for video generation. Wherein, the decomposition sub-network performs corresponding processing based on a multi-modal large language model.
[0036] In one embodiment, the step of inputting the reference image and the text prompt into the decomposition sub-network for decomposition processing to obtain a number of key stage texts includes: Inputting the reference image and the text prompt into the multi-modal large language model in the decomposition sub-network for inference processing to obtain a motion process text; Decomposing the motion process text based on the multi-modal large language model to obtain the number of key stage texts.
[0037] In this embodiment, the reference image and the text prompt word are input into the multimodal large language model in the decomposition subnetwork for inference processing to obtain the motion process text; the motion process text is decomposed based on the multimodal large language model to obtain the several key stage texts. The motion process text refers to the text of the entire motion process of the text prompt word for the reference model, and the several key stage texts refer to the decomposition of different stages of the motion process text.
[0038] Take the text prompt word as 'I have a guide text for video generation "Lift one end of the medicine box and then put it down", this prompt will guide from the given picture To generate a video, use the states of different stages to describe this process', the reference image is a picture , specifically, a picture of a medicine box with a solid color background can be taken as an example; the movement process text obtained by inputting the reference image and the text prompt word into the multimodal large language model in the decomposition subnetwork for inference processing can be: "The medicine box is lying flat, untouched, and stable on the black table; one end of the medicine box is lifted up to form an inclined position, and the other end is still in contact with the table; the lifted end of the medicine box is released and begins to fall back to the surface; the medicine box contacts the table and returns to its original position or bounces slightly when impacted"; the several key stage texts obtained by decomposing the movement process text based on the multimodal large language model can be "1. The medicine box is lying flat, untouched, and stable on the black table; 2. One end of the medicine box is lifted up to form an inclined position, and the other end is still in contact with the table; 3. The lifted end of the medicine box is released and begins to fall back to the surface; 4. The medicine box contacts the table and returns to its original position or bounces slightly when impacted".
[0039] In the decomposition process of the decomposition sub-network, the reference image The text prompt words are input into the multimodal large language model, and the entire complete and complex motion process corresponding to the text prompt words is inferred by the reasoning ability of the multimodal large language model to obtain the motion process text, and the motion process text is decomposed into several simple steps, that is, several key stage texts; the number of simple steps (the number of key stage texts) will change dynamically according to the complexity of the text prompt words and the reasoning results of the multimodal large language model, which can improve the authenticity and rationality of the subsequently generated video, and the transition of the video screen will be smoother, thereby improving the generation efficiency.
[0040] S150: Input the plurality of key stage texts into the second encoder for encoding to obtain a plurality of key vectors.
[0041] In this embodiment, the second encoder may be the text encoder of the CLIP model. The several key-stage texts are input into the second encoder (CLIP model) for encoding to obtain the several key vectors.
[0042] S160. Use the denoising sub-network to perform denoising processing on the several key vectors and the first vector to obtain several second vectors.
[0043] In this embodiment, the denoising sub-network may be a Denoising Unet model. Specifically, use the denoising sub-network (Denoising Unet model) to perform denoising processing on the several key vectors and the first vector to obtain the several second vectors.
[0044] In one embodiment, the use of the denoising sub-network to perform denoising processing on the several key vectors and the first vector to obtain several second vectors includes: For each key vector among the several key vectors, select other vectors except the key vector from the several key vectors as random vectors; Input the key vector and the random vector into a third attention formula for attention calculation processing to obtain a third intermediate vector; Input the third intermediate vector and the first vector into a fourth attention formula for attention calculation processing to obtain a fourth intermediate vector; Input the fourth intermediate vector into a fifth attention formula for attention calculation processing to obtain a second vector corresponding to the key vector; Integrate the second vectors corresponding to each key vector to obtain the several second vectors.
[0045] In this embodiment, for each key vector among the several key vectors, randomly select other vectors except the key vector from the several key vectors as the random vectors.
[0046] The input of the key vector and the random vector into the third attention formula for attention calculation processing to obtain a third intermediate vector, and the third attention formula may be: ; Where is the third intermediate vector, is equal to the product of the key vector and ; is equal to the product of the transpose of the random vector and ; is equal to the product of the random vector and ; are three matrices learned through training; d is the dimension of the key vector K; it can be seen that during the denoising process of the i-th key vector ( ), a representation of another key vector is randomly selected . This is to introduce information between key vectors during the generation process, thereby ensuring the consistency of the generated data and improving the generation efficiency.
[0047] Inputting the third intermediate vector and the first vector into the fourth attention formula for attention calculation processing to obtain a fourth intermediate vector. The fourth attention formula can be: ; where is the fourth intermediate vector, is equal to the product of the third intermediate vector and , is equal to the product of the transpose of the first vector and , is equal to the product of the first vector and , are three matrices learned through training; d is the dimension of the key vector K; it can be seen that this is to retain the original information in the reference image, ensure that there is no distortion when generating key frames subsequently, and improve the generation efficiency.
[0048] Inputting the fourth intermediate vector into the fifth attention formula for attention calculation processing to obtain a second vector corresponding to the key vector. The fifth attention formula can be: ; where is the second vector, , that is is equal to the product of the fourth intermediate vector and ; , that is is equal to the product of the fourth intermediate vector and ; , that is is equal to the product of the fourth intermediate vector and ; are also three matrices learned through training; d is the dimension of the key vector K.
[0049] Therefore, in the attention calculation process using the third attention formula, the fourth attention formula, and the fifth attention formula for the several key vectors and the first vector, except for the key vectors, random sampling is performed on the other key vectors among the several key vectors, which can increase the robustness of the model to improve the generation efficiency.
[0050] S170. Input the several second vectors into the decoder for decoding processing to obtain several key frames.
[0051] In this embodiment, the decoder may be a VAE - Decoder, that is, input the several second vectors into the decoder (VAE - Decoder) for decoding processing to obtain the several key frames.
[0052] S180. Use the motion prediction sub - network to perform conversion processing on the several key frames to obtain several converted frames.
[0053] In one embodiment, the use of the motion prediction sub - network to perform conversion processing on the several key frames to obtain several converted frames includes: For each key frame among the several key frames except the first key frame and the last key frame, obtain the two adjacent key frames of the key frame; Use the third encoder to perform compression processing on the two key frames to obtain two adjacent spatial vectors; Based on the structure predictor in the motion prediction sub - network, perform expansion processing on the two adjacent spatial vectors to obtain a sequence vector; Input the sequence vector into the conversion sub - block in the motion prediction sub - network for conversion processing to obtain a converted frame; Integrate the converted frames corresponding to each key frame to obtain the several converted frames.
[0054] In this embodiment, the third encoder is the image encoder of the CLIP model, that is, use the third encoder (CLIP model) to perform encoding processing on the several key frames to obtain the spatial vectors.
[0055] Use the third encoder to perform compression processing on the two key frames to obtain two adjacent spatial vectors; based on the structure predictor in the motion prediction sub - network, perform expansion processing on the two adjacent spatial vectors to obtain a sequence vector; input the sequence vector into the conversion sub - block in the motion prediction sub - network for conversion processing to obtain the converted frame.
[0056] For the first key frame and the last key frame among the several key frames, respectively obtain an adjacent key frame of the first key frame and the last key frame, and use the first key frame and the corresponding key frame as two key frames, and the last key frame and the corresponding key frame as two key frames; use the third encoder to compress the two key frames to obtain two adjacent spatial vectors; based on the structure predictor in the motion prediction sub-network, perform expansion processing on the two adjacent spatial vectors to obtain a sequence vector; input the sequence vector into the conversion sub-block in the motion prediction sub-network for conversion processing to obtain the conversion frame.
[0057] Therefore, the complete motion process and details of the reference image and the text prompt are improved through the motion prediction sub-network to improve the generation efficiency.
[0058] S190. Perform splicing processing on the several conversion frames to obtain a target video.
[0059] In this embodiment, the target video can be obtained by performing splicing processing on the several conversion frames converted from the several key frames. Therefore, by adopting multi-stage processing, the difficulty of generating a complete video is reduced, the video quality can be kept stable, artifacts can be reduced, and the generation efficiency can be improved.
[0060] In summary, the embodiment of the present invention can input the obtained reference image into the first encoder for encoding processing to obtain a latent vector; use the obtained text prompt to perform query processing in a preset dictionary library to obtain a text vector; use the reference sub-network to perform attention calculation on the latent vector and the text vector to obtain a first vector; input the reference image and the text prompt into the decomposition sub-network for decomposition processing to obtain several key stage texts; input the several key stage texts into the second encoder for encoding processing to obtain several key vectors; use the denoising sub-network to perform denoising processing on the several key vectors and the first vector to obtain several second vectors; input the several second vectors into the decoder for decoding processing to obtain several key frames; use the motion prediction sub-network to perform conversion processing on the several key frames to obtain several conversion frames; perform splicing processing on the several conversion frames to obtain a target video. Therefore, the embodiment of the present invention adopts multi-stage processing, ensures the spatio-temporal consistency of the video subject, increases the understanding ability of the real physical world, reduces the difficulty of generating a complete video, can keep the video quality stable, reduces artifacts, and thus improves the generation efficiency.
[0061] Figure 2 It is a schematic block diagram of a video generation device provided by an embodiment of the present invention. As Figure 2As shown above, corresponding to the above video generation method, the present invention also provides a video generation device, which is applied to a video generation network architecture. The video generation network architecture includes a first encoder, a reference sub-network, a decomposition sub-network, a second encoder, a denoising sub-network, a decoder, and a motion prediction sub-network. The video generation device and the video generation network architecture are specifically applicable to the application scenarios of digital marketing for video material generation in various fields. For example, it can be video material generation in manufacturing fields such as the medical field, the food field, or the automotive field, etc. Specifically, please refer to Figure 2 , the video generation device 700 includes: A first encoding unit 701, configured to input the obtained reference image into the first encoder for encoding processing to obtain a latent vector; A query unit 702, configured to perform query processing on the obtained text prompt in a preset dictionary library to obtain a text vector; A calculation unit 703, configured to perform attention calculation on the latent vector and the text vector by using the reference sub-network to obtain a first vector; A decomposition unit 704, configured to input the reference image and the text prompt into the decomposition sub-network for decomposition processing to obtain a plurality of key stage texts; A second encoding unit 705, configured to input the plurality of key stage texts into the second encoder for encoding processing to obtain a plurality of key vectors; A denoising unit 706, configured to perform denoising processing on the plurality of key vectors and the first vector by using the denoising sub-network to obtain a plurality of second vectors; A decoding unit 707, configured to input the plurality of second vectors into the decoder for decoding processing to obtain a plurality of key frames; A conversion unit 708, configured to perform conversion processing on the plurality of key frames by using the motion prediction sub-network to obtain a plurality of conversion frames; A splicing unit 709, configured to perform splicing processing on the plurality of conversion frames to obtain a target video.
[0062] In some embodiments, when the calculation unit 703 executes the step of performing attention calculation on the latent vector and the text vector by using the reference sub-network to obtain a first vector, it is specifically configured to: Input the latent vector into a first attention formula for calculation processing to obtain a first intermediate vector; Input the first intermediate vector and the text vector into a second attention formula for calculation processing to obtain a second intermediate vector; Using the first attention formula and the second attention formula, perform a preset number of calculation processes on the second intermediate vector and the text vector to obtain the first vector.
[0063] In some embodiments, when the decomposition unit 704 executes the step of inputting the reference image and the text prompt into the decomposition sub-network for decomposition processing to obtain a plurality of key stage texts, it is specifically configured to: Input the reference image and the text prompt into the multi-modal large language model in the decomposition sub-network for inference processing to obtain the motion process text; Based on the multi-modal large language model, perform decomposition processing on the motion process text to obtain the plurality of key stage texts.
[0064] In some embodiments, when the denoising unit 706 executes the step of using the denoising sub-network to perform denoising processing on the plurality of key vectors and the first vector to obtain a plurality of second vectors, it is specifically configured to: For each key vector among the plurality of key vectors, select other vectors except the key vector from the plurality of key vectors as random vectors; Input the key vector and the random vector into the third attention formula for attention calculation processing to obtain a third intermediate vector; Input the third intermediate vector and the first vector into the fourth attention formula for attention calculation processing to obtain a fourth intermediate vector; Input the fourth intermediate vector into the fifth attention formula for attention calculation processing to obtain a second vector corresponding to the key vector; Integrate the second vectors corresponding to each key vector to obtain the plurality of second vectors.
[0065] In some embodiments, when the conversion unit 708 executes the step of using the motion prediction sub-network to perform conversion processing on the plurality of key frames to obtain a plurality of converted frames, it is specifically configured to: For each key frame except the first key frame and the last key frame among the plurality of key frames, obtain the two adjacent key frames of the key frame; Use the third encoder to compress the two key frames to obtain two adjacent spatial vectors; Based on the structure predictor in the motion prediction sub-network, perform expansion processing on the two adjacent spatial vectors to obtain a sequence vector; Input the sequence vector into the conversion sub-block in the motion prediction sub-network for conversion processing to obtain a converted frame; Integrate the converted frames corresponding to each key frame to obtain the plurality of converted frames.
[0066] In some embodiments, when the encoding unit 701 executes the step of inputting the acquired reference image into the first encoder for encoding to obtain a latent vector, it is specifically configured to: Input the reference image into the first encoder for encoding to obtain a mean vector and a variance vector; Use the mean vector and the variance vector as the latent vector.
[0067] In some embodiments, before the query unit 702 executes the step of querying in a preset dictionary library with the acquired text prompt to obtain a text vector, it is further configured to: Perform pre-training processing on the dictionary library according to the rule that each word corresponds to a vector to obtain the preset dictionary library.
[0068] It should be noted that those skilled in the art can clearly understand that the specific implementation processes of the above video generation device and each unit can refer to the corresponding descriptions in the foregoing method embodiments. For the convenience and brevity of description, they will not be elaborated herein.
[0069] The above video generation device can be implemented in the form of a computer program, and this computer program can run on an electronic device as shown in Figure 3 Figure.
[0070] Please refer to Figure 3 , Figure 3 which is a schematic block diagram of an electronic device provided by an embodiment of the present invention. The electronic device 800 can be a terminal or a server. Among them, the terminal can be an electronic device with communication functions. The server can be an independent server or a server cluster composed of multiple servers.
[0071] Refer to Figure 3 , the electronic device 800 includes a processor 802, a memory, and a network interface 805 connected through a system bus 801. Among them, the memory can include a non-volatile storage medium 803 and an internal memory 804.
[0072] The non-volatile storage medium 803 can store an operating system 8031 and a computer program 8032. The computer program 8032 includes program instructions, and when these program instructions are executed, the processor 802 can execute a video generation method.
[0073] The processor 802 is used to provide computing and control capabilities to support the operation of the entire electronic device 800.
[0074] The internal memory 804 provides an environment for the operation of the computer program 8032 in the non-volatile storage medium 803. When the computer program 8032 is executed by the processor 802, the processor 802 can be caused to execute a video generation method.
[0075] The network interface 805 is used for network communication with other devices. Those skilled in the art can understand that Figure 3 the structure shown in is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the electronic device 800 to which the solution of the present invention is applied. The specific electronic device 800 may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0076] Wherein, the processor 802 is used to run the computer program 8032 stored in the memory to implement the following steps: Input the obtained reference image into the first encoder for encoding processing to obtain a latent vector; Use the obtained text prompt to query in the preset dictionary library to obtain a text vector; Use the reference sub-network to perform attention calculation on the latent vector and the text vector to obtain a first vector; Input the reference image and the text prompt into the decomposition sub-network for decomposition processing to obtain a number of key stage texts; Input the number of key stage texts into the second encoder for encoding processing to obtain a number of key vectors; Use the denoising sub-network to perform denoising processing on the number of key vectors and the first vector to obtain a number of second vectors; Input the number of second vectors into the decoder for decoding processing to obtain a number of key frames; Use the motion prediction sub-network to perform conversion processing on the number of key frames to obtain a number of conversion frames; Perform splicing processing on the number of conversion frames to obtain a target video.
[0077] In some embodiments, when the processor 802 implements the step of using the reference sub-network to perform attention calculation on the latent vector and the text vector to obtain a first vector, it is specifically used for: Input the latent vector into the first attention formula for calculation processing to obtain a first intermediate vector; Input the first intermediate vector and the text vector into the second attention formula for calculation processing to obtain a second intermediate vector; Performing a preset number of calculation processes on the second intermediate vector and the text vector using the first attention formula and the second attention formula to obtain the first vector.
[0078] In some embodiments, when the processor 802 implements the step of inputting the reference image and the text prompt into the decomposition sub-network for decomposition processing to obtain a plurality of key stage texts, it is specifically configured to: Input the reference image and the text prompt into the multi-modal large language model in the decomposition sub-network for inference processing to obtain the motion process text; Based on the multi-modal large language model, perform decomposition processing on the motion process text to obtain the plurality of key stage texts.
[0079] In some embodiments, when the processor 802 implements the step of using the denoising sub-network to perform denoising processing on the plurality of key vectors and the first vector to obtain a plurality of second vectors, it is specifically configured to: For each key vector among the plurality of key vectors, select other vectors except the key vector from the plurality of key vectors as random vectors; Input the key vector and the random vector into the third attention formula for attention calculation processing to obtain a third intermediate vector; Input the third intermediate vector and the first vector into the fourth attention formula for attention calculation processing to obtain a fourth intermediate vector; Input the fourth intermediate vector into the fifth attention formula for attention calculation processing to obtain a second vector corresponding to the key vector; Integrate the second vectors corresponding to each key vector to obtain the plurality of second vectors.
[0080] In some embodiments, when the processor 802 implements the step of using the motion prediction sub-network to perform conversion processing on the plurality of key frames to obtain a plurality of converted frames, it is specifically configured to: For each key frame among the plurality of key frames except the first key frame and the last key frame, obtain the two adjacent key frames of the key frame; Use the third encoder to perform compression processing on the two key frames to obtain two adjacent spatial vectors; Based on the structure predictor in the motion prediction sub-network, perform expansion processing on the two adjacent spatial vectors to obtain a sequence vector; Input the sequence vector into the conversion sub-block in the motion prediction sub-network for conversion processing to obtain a converted frame; Integrate the converted frames corresponding to each key frame to obtain the plurality of converted frames.
[0081] In some embodiments, when the processor 802 implements the step of inputting the acquired reference image into the first encoder for encoding to obtain a latent vector, it is specifically configured to: Input the reference image into the first encoder for encoding to obtain a mean vector and a variance vector; Use the mean vector and the variance vector as the latent vector.
[0082] In some embodiments, before the processor 802 implements the step of querying the acquired text prompt in a preset dictionary library to obtain a text vector, it is further configured to: Pre-train the dictionary library according to the rule that each word corresponds to a vector to obtain the preset dictionary library.
[0083] It should be understood that in the embodiments of the present invention, the processor 802 may be a central processing unit (CPU), and the processor 802 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0084] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0085] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor executes the following steps: Input the acquired reference image into the first encoder for encoding to obtain a latent vector; Query the acquired text prompt in a preset dictionary library to obtain a text vector; Use the reference sub-network to perform attention calculation on the latent vector and the text vector to obtain a first vector; Input the reference image and the text prompt into the decomposition sub-network for decomposition processing to obtain a number of key stage texts; Input the number of key stage texts into the second encoder for encoding processing to obtain a number of key vectors; Use the denoising sub-network to perform denoising processing on the number of key vectors and the first vector to obtain a number of second vectors; Input the number of second vectors into the decoder for decoding processing to obtain a number of key frames; Use the motion prediction sub-network to perform conversion processing on the number of key frames to obtain a number of conversion frames; Perform splicing processing on the number of conversion frames to obtain the target video.
[0086] In one embodiment, when the processor executes the program instructions to implement the step of using the reference sub-network to perform attention calculation on the latent vector and the text vector to obtain a first vector, it is specifically used for: Input the latent vector into the first attention formula for calculation processing to obtain a first intermediate vector; Input the first intermediate vector and the text vector into the second attention formula for calculation processing to obtain a second intermediate vector; Use the first attention formula and the second attention formula to perform a preset number of calculation processes on the second intermediate vector and the text vector to obtain the first vector.
[0087] In one embodiment, when the processor executes the program instructions to implement the step of inputting the reference image and the text prompt into the decomposition sub-network for decomposition processing to obtain a number of key stage texts, it is specifically used for: Input the reference image and the text prompt into the multi-modal large language model in the decomposition sub-network for inference processing to obtain the motion process text; Based on the multi-modal large language model, perform decomposition processing on the motion process text to obtain the number of key stage texts.
[0088] In one embodiment, when the processor executes the program instructions to implement the step of using the denoising sub-network to perform denoising processing on the number of key vectors and the first vector to obtain a number of second vectors, it is specifically used for: For each key vector among the number of key vectors, select other vectors except the key vector from the number of key vectors as random vectors; Input the key vector and the random vector into the third attention formula for attention calculation to obtain a third intermediate vector; Input the third intermediate vector and the first vector into the fourth attention formula for attention calculation to obtain a fourth intermediate vector; Input the fourth intermediate vector into the fifth attention formula for attention calculation to obtain a second vector corresponding to the key vector; Integrate the second vectors corresponding to each key vector to obtain the several second vectors.
[0089] In one embodiment, when the processor executes the program instructions to implement the step of using the motion prediction sub-network to perform conversion processing on the several key frames to obtain several converted frames, it is specifically configured to: For each key frame among the several key frames except the first key frame and the last key frame, obtain two adjacent key frames of the key frame; Use the third encoder to perform compression processing on the two key frames to obtain two adjacent spatial vectors; Based on the structure predictor in the motion prediction sub-network, perform expansion processing on the two adjacent spatial vectors to obtain a sequence vector; Input the sequence vector into the conversion sub-block in the motion prediction sub-network for conversion processing to obtain a converted frame; Integrate the converted frames corresponding to each key frame to obtain the several converted frames.
[0090] In one embodiment, when the processor executes the program instructions to implement the step of inputting the obtained reference image into the first encoder for encoding processing to obtain a latent vector, it is specifically configured to: Input the reference image into the first encoder for encoding processing to obtain a mean vector and a variance vector; Use the mean vector and the variance vector as the latent vector.
[0091] In one embodiment, before the processor executes the program instructions to implement the step of using the obtained text prompt to perform query processing in the preset dictionary library to obtain a text vector, it is further configured to: Use the rule that each word corresponds to a vector to perform pre-training processing on the dictionary library to obtain the preset dictionary library.
[0092] The storage medium can be various computer-readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes.
[0093] Those of ordinary skill in the art will appreciate that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described in terms of function in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0094] In several embodiments provided by the present invention, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0095] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the apparatus embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of the present invention can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0096] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention.
[0097] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A video generation method, characterized in that: Applied to a video generation network architecture, the video generation network architecture includes a first encoder, a reference sub-network, a decomposition sub-network, a second encoder, a denoising sub-network, a decoder and a motion prediction sub-network; the video generation method includes: Inputting the acquired reference image into the first encoder for encoding to obtain a latent vector; Using the acquired text prompt words to perform query processing in a preset dictionary library to obtain a text vector; Using the reference sub-network to perform attention calculation on the potential vector and the text vector to obtain a first vector; Inputting the reference image and the text prompt words into the decomposition subnetwork for decomposition processing to obtain a plurality of key stage texts; Inputting the plurality of key stage texts into the second encoder for encoding to obtain a plurality of key vectors; Using the denoising sub-network to perform denoising on the plurality of key vectors and the first vector to obtain a plurality of second vectors; Inputting the plurality of second vectors into the decoder for decoding to obtain a plurality of key frames; Using the motion prediction subnetwork to convert the plurality of key frames to obtain a plurality of conversion frames; The target video is obtained by splicing the plurality of conversion frames.
2. The video generation method according to claim 1, characterized in that: The using the reference sub-network to perform attention calculation on the potential vector and the text vector to obtain a first vector includes: Inputting the potential vector into a first attention formula for calculation and processing to obtain a first intermediate vector; Inputting the first intermediate vector and the text vector into a second attention formula for calculation and processing to obtain a second intermediate vector; The first attention formula and the second attention formula are used to perform calculations on the second intermediate vector and the text vector a preset number of times to obtain the first vector.
3. The video generation method according to claim 1, characterized in that: The step of inputting the reference image and the text prompt word into the decomposition subnetwork for decomposition processing to obtain a plurality of key stage texts includes: Inputting the reference image and the text prompt word into the multimodal large language model in the decomposition sub-network for inference processing to obtain the motion process text; The movement process text is decomposed based on the multimodal large language model to obtain the several key stage texts.
4. The video generation method according to claim 1, characterized in that: The step of using the denoising sub-network to perform denoising on the plurality of key vectors and the first vector to obtain a plurality of second vectors includes: For each key vector among the plurality of key vectors, selecting other vectors except the key vector from the plurality of key vectors as random vectors; Inputting the key vector and the random vector into a third attention formula to perform attention calculation processing to obtain a third intermediate vector; Inputting the third intermediate vector and the first vector into a fourth attention formula to perform attention calculation processing to obtain a fourth intermediate vector; Inputting the fourth intermediate vector into a fifth attention formula to perform attention calculation processing to obtain a second vector corresponding to the key vector; The second vectors corresponding to each key vector are integrated to obtain the plurality of second vectors.
5. The video generation method according to claim 1, characterized in that: The converting the plurality of key frames using the motion prediction subnetwork to obtain a plurality of converted frames includes: For each key frame except the first key frame and the last key frame among the plurality of key frames, obtaining two adjacent key frames of the key frame; Using the third encoder to compress the two key frames to obtain two adjacent space vectors; Performing expansion processing on the two adjacent space vectors based on the structure predictor in the motion prediction subnetwork to obtain a sequence vector; Inputting the sequence vector into the conversion sub-block in the motion prediction sub-network for conversion processing to obtain a conversion frame; The conversion frames corresponding to each key frame are integrated to obtain the plurality of conversion frames.
6. The video generation method according to claim 1, characterized in that: The step of inputting the acquired reference image into the first encoder for encoding to obtain a potential vector includes: Inputting the reference image into the first encoder for encoding to obtain a mean vector and a variance vector; The mean vector and the variance vector are used as the latent vector.
7. The video generation method according to claim 1, characterized in that: Before the method uses the acquired text prompt words to perform query processing in a preset dictionary library to obtain a text vector, the method further includes: The preset dictionary library is obtained by pre-training the dictionary library using the rule that each word corresponds to a vector.
8. A video generating device, characterized in that: Applied to a video generation network architecture, the video generation network architecture includes a first encoder, a reference sub-network, a decomposition sub-network, a second encoder, a denoising sub-network, a decoder and a motion prediction sub-network; the video generation device includes: A first encoding unit, configured to input the acquired reference image into the first encoder for encoding to obtain a latent vector; A query unit, used to query the acquired text prompt words in a preset dictionary library to obtain a text vector; A computing unit, configured to use the reference sub-network to perform attention calculation on the potential vector and the text vector to obtain a first vector; A decomposition unit, used for inputting the reference image and the text prompt word into the decomposition sub-network for decomposition processing to obtain a plurality of key stage texts; A second encoding unit, configured to input the plurality of key stage texts into the second encoder for encoding to obtain a plurality of key vectors; A denoising unit, configured to perform denoising processing on the plurality of key vectors and the first vector using the denoising sub-network to obtain a plurality of second vectors; A decoding unit, used for inputting the plurality of second vectors into the decoder for decoding to obtain a plurality of key frames; A conversion unit, configured to convert the plurality of key frames using the motion prediction subnetwork to obtain a plurality of conversion frames; The splicing unit is used to splice the plurality of conversion frames to obtain a target video.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the video generating method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the video generating method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Training method, device and equipment for digital human model
CN119031210A
Video generation method and device, equipment and medium
CN119383289A
Video generation method, and server
WO2024228676A1
Cited By
Video generation method and device based on hidden space decomposition, equipment and storage medium
CN121239921A