Video generation method, device, equipment and medium
Through the multi-stage video generation network architecture, the problems of low generation efficiency and unstable spatial and temporal consistency are solved, efficient generation of high-quality videos is achieved, and understanding of the real physical world is enhanced.
Patent Information
- Application Number
- CN202510534037.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The existing video generation methods have problems with low generation efficiency and lack the ability to understand the real physical world, which leads to unstable spatial and temporal consistency of generated videos.
The video generation network architecture is adopted, including the first encoder, reference subnet, decomposition subnet, second encoder, denoising subnet and motion prediction subnet. Through multi-stage processing, the reference image and text prompt words are converted into keyframes and spliced to ensure the spatial and temporal consistency of the video subject and the understanding of the real physical world.
Improve the efficiency of video generation, maintain the stability of video quality, reduce artifacts, and enhance the ability to understand the real physical world.
Smart Images

Figure CN120186433B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a video generation method, device, equipment and medium. Background Art
[0002] In recent years, conditionally controlled image-generated videos have gradually become a research hotspot in computer vision, and image-generated video methods have achieved great success. However, these methods cannot effectively solve the problem of unstable spatiotemporal consistency of the subjects in the generated videos, and the models lack the ability to understand the real physical world, which leads to low generation efficiency.
[0003] Therefore, existing video generation methods have the problem of low generation efficiency. Summary of the Invention
[0004] The embodiments of the present invention provide a video generation method, apparatus, device and medium, aiming to solve the problem of low generation efficiency in existing video generation methods.
[0005] In a first aspect, an embodiment of the present invention provides a video generation method, which is applied to a video generation network architecture, wherein the video generation network architecture includes a first encoder, a reference subnetwork, a decomposition subnetwork, a second encoder, a denoising subnetwork, a decoder, and a motion prediction subnetwork; the video generation method includes:
[0006] Inputting the obtained reference image into the first encoder for encoding to obtain a latent vector;
[0007] Using the acquired text prompt words to query the preset dictionary library to obtain the text vector;
[0008] Performing attention calculation on the potential vector and the text vector using the reference sub-network to obtain a first vector;
[0009] Inputting the reference image and the text prompt words into the decomposition sub-network for decomposition processing to obtain a plurality of key stage texts;
[0010] Inputting the plurality of key stage texts into the second encoder for encoding to obtain a plurality of key vectors;
[0011] Using the denoising sub-network to perform denoising on the plurality of key vectors and the first vector to obtain a plurality of second vectors;
[0012] Inputting the plurality of second vectors into the decoder for decoding to obtain a plurality of key frames;
[0013] Using the motion prediction subnetwork to convert the plurality of key frames to obtain a plurality of conversion frames;
[0014] The plurality of converted frames are spliced together to obtain a target video.
[0015] In a second aspect, an embodiment of the present invention further provides a video generation device, which is applied to a video generation network architecture, wherein the video generation network architecture includes a first encoder, a reference subnetwork, a decomposition subnetwork, a second encoder, a denoising subnetwork, a decoder, and a motion prediction subnetwork; the video generation device includes:
[0016] A first encoding unit, configured to input the acquired reference image into the first encoder for encoding to obtain a latent vector;
[0017] A query unit, configured to use the acquired text prompt words to perform query processing in a preset dictionary library to obtain a text vector;
[0018] a computing unit, configured to perform attention calculation on the potential vector and the text vector using the reference sub-network to obtain a first vector;
[0019] a decomposition unit, configured to input the reference image and the text prompt word into the decomposition sub-network for decomposition processing to obtain a plurality of key stage texts;
[0020] A second encoding unit, configured to input the plurality of key stage texts into the second encoder for encoding to obtain a plurality of key vectors;
[0021] a denoising unit, configured to perform denoising processing on the plurality of key vectors and the first vector using the denoising sub-network to obtain a plurality of second vectors;
[0022] a decoding unit, configured to input the plurality of second vectors into the decoder for decoding to obtain a plurality of key frames;
[0023] a conversion unit, configured to convert the plurality of key frames using the motion prediction subnetwork to obtain a plurality of conversion frames;
[0024] The splicing unit is used to splice the plurality of converted frames to obtain a target video.
[0025] In a third aspect, an embodiment of the present invention further provides an electronic device, which includes a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method described in the first aspect is implemented.
[0026] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the method described in the first aspect can be implemented.
[0027] The present invention provides a video generation method, apparatus, device and medium. The video generation method is applied to a video generation network architecture, which includes a first encoder, a reference subnetwork, a decomposition subnetwork, a second encoder, a denoising subnetwork, a decoder and a motion prediction subnetwork. The video generation method includes: inputting an acquired reference image into the first encoder for encoding processing to obtain a latent vector; using the acquired text prompt word to query a preset dictionary library to obtain a text vector; using the reference subnetwork to perform attention calculation on the latent vector and the text vector to obtain a first vector; inputting the reference image and the text prompt word into the decomposition subnetwork for decomposition processing to obtain a plurality of key stage texts; inputting the plurality of key stage texts into the second encoder for encoding processing to obtain a plurality of key vectors; using the denoising subnetwork to denoise the plurality of key vectors and the first vector to obtain a plurality of second vectors; inputting the plurality of second vectors into the decoder for decoding processing to obtain a plurality of key frames; using the motion prediction subnetwork to convert the plurality of key frames to obtain a plurality of converted frames; and splicing the plurality of converted frames to obtain a target video. It can be seen that the embodiment of the present invention uses a first encoder, a reference sub-network, a decomposition sub-network, a second encoder, a denoising sub-network, a decoder and a motion prediction sub-network to implement multi-stage processing of the reference image and text prompt words to ensure the spatiotemporal consistency of the video subject, increase the ability to understand the real physical world, reduce the difficulty of complete video generation, maintain the stability of video quality, reduce artifacts, and thus improve generation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0029] Figure 1 A schematic diagram of a flow chart of a video generation method provided by an embodiment of the present invention;
[0030] Figure 2 A schematic block diagram of a video generation device provided by an embodiment of the present invention;
[0031] Figure 3 A schematic block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0032] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0033] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0034] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0035] It should be further understood that the term "and / or" used in the present specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations. The embodiments of the present invention provide a video generation method, apparatus, device and medium, and the video generation method can be applied to a video generation network architecture, and the video generation network architecture includes a first encoder, a reference subnetwork, a decomposition subnetwork, a second encoder, a denoising subnetwork, a decoder and a motion prediction subnetwork. The video generation method and the video generation network architecture are specifically applicable to the application scenarios of digital marketing of video material generation in various fields, for example, it can be the generation of video material in the fields of medical health, food or automobile. The present invention obtains a latent vector by inputting the acquired reference image into the first encoder for encoding processing; obtains a text vector by querying the acquired text prompt word in a preset dictionary library; obtains a text vector by using the reference sub-network to perform attention calculation on the latent vector and the text vector; inputs the reference image and the text prompt word into the decomposition sub-network for decomposition processing to obtain several key stage texts; inputs the several key stage texts into the second encoder for encoding processing to obtain several key vectors; uses the denoising sub-network to perform denoising processing on the several key vectors and the first vector to obtain several second vectors; inputs the several second vectors into the decoder for decoding processing to obtain several key frames; uses the motion prediction sub-network to convert the several key frames to obtain several conversion frames; and splices the several conversion frames to obtain a target video, so as to improve the generation efficiency of video generation. The present invention is described in detail below through specific embodiments.
[0036] Figure 1 Schematic diagram of the process of video generation method provided by the embodiment of the present invention. Figure 1 As shown, the method includes the following steps S110-S190.
[0037] S110: Input the obtained reference image into the first encoder for encoding to obtain a latent vector.
[0038] In this embodiment, the reference image refers to an image input by a user, and specifically, the reference image input by the user can be obtained from a terminal, such as an image of medical and health-related medicines, instruments, etc., or an image of food-related snacks, drinks, etc.; the terminal can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices; the first encoder can be a VAE-Encoder; specifically, the reference image is encoded using the first encoder to obtain the latent vector.
[0039] In one embodiment, inputting the acquired reference image into the first encoder for encoding to obtain a latent vector includes:
[0040] Inputting the reference image into the first encoder for encoding to obtain a mean vector and a variance vector;
[0041] The mean vector and the variance vector are used as the latent vector.
[0042] In this embodiment, encoding the reference image using the first encoder is used to represent the reference image using a low-dimensional latent vector, thereby saving computational resources. For example, if the data format of the reference image is 224*224*3=150528, after encoding by the first encoder (VAE-Encoder), only a 256-dimensional mean vector and a 256-dimensional variance vector are required to represent the reference image.
[0043] S120: Query the acquired text prompt words in a preset dictionary to obtain a text vector.
[0044] In this embodiment, the text prompt word refers to a text description input by the user, for example: I have a guide text for video generation "lift one end of the medicine box and then put it down", this prompt will guide the user to To generate a video, use the states of different stages to describe the process; among them, the picture refers to the reference image.
[0045] Specifically, the text prompt word is used as a query condition, and query processing is performed in the preset dictionary library to obtain the text vector.
[0046] In one embodiment, before using the acquired text prompt words to perform query processing in a preset dictionary library to obtain a text vector, the method further includes:
[0047] The preset dictionary library is obtained by pre-training the dictionary library using the rule that each word corresponds to a vector.
[0048] In this embodiment, each word may refer to a single character or multiple characters; each word corresponds to a vector one-to-one; and the dictionary library may refer to a database storing a number of characters or words.
[0049] Specifically, the dictionary library is pre-trained using the rule that each word corresponds to a vector to obtain the preset dictionary library. Therefore, after the preset dictionary library is pre-trained, the text prompt word can be quickly queried using the preset dictionary library to obtain the text vector, achieving rapid conversion of the text prompt word and improving data accuracy and generation efficiency.
[0050] S130. Use the reference sub-network to perform attention calculation on the potential vector and the text vector to obtain a first vector.
[0051] In this embodiment, the reference sub-network may be ReferenceNet, that is, ReferenceNet is used to perform attention calculation on the potential vector and the text vector to obtain the first vector.
[0052] In one embodiment, the step of performing attention calculation on the potential vector and the text vector using the reference sub-network to obtain the first vector includes:
[0053] Inputting the potential vector into a first attention formula for calculation to obtain a first intermediate vector;
[0054] Inputting the first intermediate vector and the text vector into a second attention formula for calculation to obtain a second intermediate vector;
[0055] The first vector is obtained by performing calculations on the second intermediate vector and the text vector a preset number of times using the first attention formula and the second attention formula.
[0056] In this embodiment, the potential vector is input into the first attention formula for calculation to obtain the first intermediate vector, wherein the first attention formula can be:
[0057] ;
[0058] in, is the spatial attention score, that is, the first intermediate vector; Represents Sigmoid operation; , which compresses the output to between 0 and 1, indicating the weight of attention; Represents a 7×7 convolution operation; Indicates that the average pooling operation is performed on the input to obtain the global average information; Indicates that the maximum pooling operation is performed on the input data to obtain the global maximum information; the goal of the first attention formula is to strengthen important spatial areas and suppress unimportant areas.
[0059] The first intermediate vector and the text vector are input into a second attention formula for calculation to obtain a second intermediate vector, wherein the second attention formula may be:
[0060] ;
[0061] in, is the second intermediate vector, , , ; Z is the latent vector of the reference image; is the text vector of the text prompt word, are three matrices learned through training; d is the dimension of the key vector K.
[0062] The first vector is obtained by performing calculation processing on the second intermediate vector and the text vector using the first attention formula and the second attention formula for a preset number of times. Specifically, the preset number of times can be a value such as 11 or 12; specifically, the second intermediate vector is input into the first attention formula for attention calculation processing to obtain an updated second intermediate vector, the updated second intermediate vector and the text vector are input into the second attention formula for attention calculation processing to obtain a processed second intermediate vector, and the processed second intermediate vector is input into the first attention formula for the next attention calculation processing, and so on, to complete the preset number of attention calculation processing; wherein, the preset number of attention calculation processing can include the calculation processing of the first attention formula and the calculation processing of the second attention formula.
[0063] Therefore, the latent vector and the text vector serve as references to each other in the attention calculation process performed using the first attention formula and the second attention formula to ensure the subject spatiotemporal consistency between the subsequently obtained key frames, thereby ensuring the accuracy of the data and improving the generation efficiency.
[0064] S140: Input the reference image and the text prompt word into the decomposition sub-network for decomposition processing to obtain a plurality of key stage texts.
[0065] In this embodiment, the text prompt is used by the user to describe the requirements for the generated video content; the text prompt and the reference image can serve as control conditions for video generation. The decomposition sub-network performs corresponding processing based on a multimodal large language model.
[0066] In one embodiment, the step of inputting the reference image and the text prompt word into the decomposition sub-network for decomposition processing to obtain a plurality of key stage texts includes:
[0067] Inputting the reference image and the text prompt word into the multimodal large language model in the decomposition sub-network for inference processing to obtain the motion process text;
[0068] The movement process text is decomposed based on the multimodal large language model to obtain the several key stage texts.
[0069] In this embodiment, the reference image and the text prompt are input into the multimodal large language model in the decomposition subnetwork for inference processing to obtain the movement process text. Based on the multimodal large language model, the movement process text is decomposed to obtain the multiple key stage texts. The movement process text refers to the text prompt for the entire movement process of the reference model, and the multiple key stage texts refer to the decomposition of the movement process text at different stages.
[0070] The text prompt is 'I have a guide text for video generation "Lift up one end of the medicine box and then put it down", this prompt will guide the video from the given picture To generate a video, use the states of different stages to describe this process', the reference image is a picture , specifically, a picture of a medicine box with a solid color background can be taken as an example; the reference image and the text prompt word are input into the multimodal large language model in the decomposition sub-network for inference processing to obtain the movement process text which can be: "The medicine box is placed flat, untouched, and stable on the black table; one end of the medicine box is lifted up to form an inclined position, and the other end is still in contact with the table; the lifted end of the medicine box is released and begins to fall back to the surface; the medicine box contacts the table and returns to its original position or bounces slightly when hit"; the several key stage texts obtained by decomposing the movement process text based on the multimodal large language model can be "1. The medicine box is placed flat, untouched, and stable on the black table; 2. One end of the medicine box is lifted up to form an inclined position, and the other end is still in contact with the table; 3. The lifted end of the medicine box is released and begins to fall back to the surface; 4. The medicine box contacts the table and returns to its original position or bounces slightly when hit".
[0071] During the decomposition process of the decomposition sub-network, the reference image The text prompt word is input into the multimodal large language model, and the reasoning ability of the multimodal large language model is used to infer the entire complex motion process corresponding to the text prompt word to obtain the motion process text, and the motion process text is decomposed into several simple steps, that is, several key stage texts; the number of simple steps (the number of key stage texts) will dynamically change according to the complexity of the text prompt word and the reasoning result of the multimodal large language model, which can improve the authenticity and rationality of the subsequently generated video, and the transition of the video screen will also be smoother, thereby improving the generation efficiency.
[0072] S150: Input the plurality of key stage texts into the second encoder for encoding to obtain a plurality of key vectors.
[0073] In this embodiment, the second encoder may be a text encoder of a CLIP model, and the plurality of key stage texts are input into the second encoder (CLIP model) for encoding processing to obtain the plurality of key vectors.
[0074] S160: Utilize the denoising sub-network to perform denoising processing on the plurality of key vectors and the first vector to obtain a plurality of second vectors.
[0075] In this embodiment, the denoising sub-network may be a Denoising Unet model; specifically, the denoising sub-network (Denoising Unet model) is used to perform denoising processing on the plurality of key vectors and the first vector to obtain the plurality of second vectors.
[0076] In one embodiment, the denoising sub-network is used to perform denoising on the plurality of key vectors and the first vector to obtain a plurality of second vectors, including:
[0077] For each key vector among the plurality of key vectors, selecting other vectors except the key vector from the plurality of key vectors as random vectors;
[0078] Inputting the key vector and the random vector into a third attention formula to perform attention calculation processing to obtain a third intermediate vector;
[0079] Inputting the third intermediate vector and the first vector into a fourth attention formula to perform attention calculation processing to obtain a fourth intermediate vector;
[0080] Inputting the fourth intermediate vector into a fifth attention formula to perform attention calculation processing to obtain a second vector corresponding to the key vector;
[0081] The second vectors corresponding to each key vector are integrated to obtain the plurality of second vectors.
[0082] In this embodiment, for each key vector among the multiple key vectors, other vectors except the key vector are randomly selected from the multiple key vectors as the random vector.
[0083] The key vector and the random vector are input into a third attention formula for performing attention calculation processing to obtain a third intermediate vector. The third attention formula can be:
[0084] ;
[0085] in, is the third intermediate vector, Equal to the key vector and The product of is equal to the transpose and sum of the random vectors The product of Equal to the random vector and The product of are three matrices learned through training; d is the dimension of the key vector K; it can be seen that in the i-th key vector ( ) in the denoising process, a representation of another key vector will be randomly selected ,This is to introduce information between the key vectors during the ,generation process, thereby ensuring the consistency of the generated data and ,improving the generation efficiency.
[0086] The third intermediate vector and the first vector are input into a fourth attention formula for attention calculation to obtain a fourth intermediate vector. The fourth attention formula may be:
[0087] ;
[0088] in, is the fourth intermediate vector, Equal to the third intermediate vector and The product of is equal to the transpose of the first vector and The product of is equal to the first vector and The product of are three matrices learned through training; d is the dimension of the key vector K; it can be seen that this is to retain the original information in the reference image and ensure that there is no distortion when the key frame is generated subsequently, so as to improve the generation efficiency.
[0089] The fourth intermediate vector is input into the fifth attention formula for attention calculation to obtain a second vector corresponding to the key vector. The fifth attention formula can be:
[0090] ;
[0091] in, is the second vector, ,Right now Equal to the fourth intermediate vector and The product of ,Right now Equal to the fourth intermediate vector and The product of ,Right now Equal to the fourth intermediate vector and The product of It is also a matrix learned through three trainings; d is the dimension of the key vector K.
[0092] Therefore, in the attention calculation processing of the several key vectors and the first vector using the third attention formula, the fourth attention formula and the fifth attention formula, except for the key vector, other key vectors among the several key vectors except the key vector are randomly sampled, which can increase the robustness of the model to improve the generation efficiency.
[0093] S170: Input the plurality of second vectors into the decoder for decoding to obtain a plurality of key frames.
[0094] In this embodiment, the decoder may be a VAE-Decoder, that is, the plurality of second vectors are input into the decoder (VAE-Decoder) for decoding to obtain the plurality of key frames.
[0095] S180: Utilize the motion prediction subnetwork to perform conversion processing on the plurality of key frames to obtain a plurality of conversion frames.
[0096] In one embodiment, the converting the plurality of key frames using the motion prediction subnetwork to obtain a plurality of converted frames includes:
[0097] For each key frame except the first key frame and the last key frame among the plurality of key frames, obtaining two adjacent key frames of the key frame;
[0098] Using the third encoder to compress the two key frames to obtain two adjacent space vectors;
[0099] Performing expansion processing on the two adjacent space vectors based on the structure predictor in the motion prediction subnetwork to obtain a sequence vector;
[0100] Inputting the sequence vector into the conversion sub-block in the motion prediction sub-network for conversion processing to obtain a conversion frame;
[0101] The conversion frames corresponding to each key frame are integrated to obtain the plurality of conversion frames.
[0102] In this embodiment, the third encoder is an image encoder of a CLIP model, that is, the third encoder (CLIP model) is used to encode the plurality of key frames to obtain the spatial vectors.
[0103] The two key frames are compressed using the third encoder to obtain two adjacent spatial vectors; the two adjacent spatial vectors are expanded based on the structure predictor in the motion prediction subnetwork to obtain a sequence vector; and the sequence vector is input into the conversion sub-block in the motion prediction subnetwork for conversion processing to obtain the conversion frame.
[0104] For the first key frame and the last key frame among the several key frames, obtain a key frame adjacent to the first key frame and the last key frame respectively, and use the first key frame and the corresponding key frame as two key frames, and the last key frame and the corresponding key frame as two key frames; use the third encoder to compress the two key frames to obtain two adjacent spatial vectors; based on the structure predictor in the motion prediction subnetwork, expand the two adjacent spatial vectors to obtain a sequence vector; input the sequence vector into the conversion subblock in the motion prediction subnetwork for conversion processing to obtain the conversion frame.
[0105] Therefore, the complete motion process and details of the reference image and the text prompt word are improved through the motion prediction sub-network to improve the generation efficiency.
[0106] S190: Splicing the plurality of converted frames to obtain a target video.
[0107] In this embodiment, the target video can be obtained by splicing a plurality of converted frames converted from the plurality of key frames. Therefore, the multi-stage processing reduces the difficulty of generating a complete video, maintains stable video quality, reduces artifacts, and thus improves generation efficiency.
[0108] In summary, the embodiment of the present invention can obtain a latent vector by inputting the obtained reference image into the first encoder for encoding; obtain a text vector by querying the obtained text prompt word in a preset dictionary library; use the reference sub-network to perform attention calculation on the latent vector and the text vector to obtain a first vector; input the reference image and the text prompt word into the decomposition sub-network for decomposition processing to obtain several key stage texts; input the several key stage texts into the second encoder for encoding processing to obtain several key vectors; use the denoising sub-network to denoise the several key vectors and the first vector to obtain several second vectors; input the several second vectors into the decoder for decoding processing to obtain several key frames; use the motion prediction sub-network to convert the several key frames to obtain several converted frames; and splice the several converted frames to obtain the target video. Therefore, the embodiment of the present invention adopts multi-stage processing to ensure the spatiotemporal consistency of the video subject, increase the ability to understand the real physical world, reduce the difficulty of generating a complete video, maintain stable video quality, reduce artifacts, and thus improve generation efficiency.
[0109] Figure 2 Schematic block diagram of a video generation device provided by an embodiment of the present invention. Figure 2 As shown, corresponding to the above video generation method, the present invention also provides a video generation device, which is applied to a video generation network architecture, and the video generation network architecture includes a first encoder, a reference sub-network, a decomposition sub-network, a second encoder, a denoising sub-network, a decoder and a motion prediction sub-network. The video generation device and the video generation network architecture are specifically suitable for digital marketing application scenarios of video material generation in various fields, for example, it can be video material generation in manufacturing fields such as the medical field, the food field or the automotive field. Specifically, please refer to Figure 2 , the video generating device 700 includes:
[0110] A first encoding unit 701 is configured to input the acquired reference image into the first encoder for encoding to obtain a latent vector;
[0111] A query unit 702 is configured to query a preset dictionary using the acquired text prompt words to obtain a text vector;
[0112] A calculation unit 703 is configured to perform attention calculation on the potential vector and the text vector using the reference sub-network to obtain a first vector;
[0113] A decomposition unit 704 is configured to input the reference image and the text prompt word into the decomposition sub-network for decomposition processing to obtain a plurality of key stage texts;
[0114] The second encoding unit 705 is configured to input the plurality of key stage texts into the second encoder for encoding to obtain a plurality of key vectors;
[0115] a denoising unit 706, configured to perform denoising processing on the plurality of key vectors and the first vector using the denoising sub-network to obtain a plurality of second vectors;
[0116] A decoding unit 707, configured to input the plurality of second vectors into the decoder for decoding to obtain a plurality of key frames;
[0117] A conversion unit 708 is configured to convert the plurality of key frames using the motion prediction subnetwork to obtain a plurality of conversion frames;
[0118] The splicing unit 709 is configured to splice the plurality of converted frames to obtain a target video.
[0119] In some embodiments, when executing the step of performing attention calculation on the potential vector and the text vector using the reference sub-network to obtain the first vector, the calculation unit 703 is specifically configured to:
[0120] Inputting the potential vector into a first attention formula for calculation to obtain a first intermediate vector;
[0121] Inputting the first intermediate vector and the text vector into a second attention formula for calculation to obtain a second intermediate vector;
[0122] The first vector is obtained by performing calculations on the second intermediate vector and the text vector a preset number of times using the first attention formula and the second attention formula.
[0123] In some embodiments, when the decomposition unit 704 inputs the reference image and the text prompt word into the decomposition sub-network for decomposition processing to obtain a plurality of key stage texts, it is specifically configured to:
[0124] Inputting the reference image and the text prompt word into the multimodal large language model in the decomposition sub-network for inference processing to obtain the motion process text;
[0125] The movement process text is decomposed based on the multimodal large language model to obtain the several key stage texts.
[0126] In some embodiments, when executing the step of performing denoising processing on the plurality of key vectors and the first vector using the denoising sub-network to obtain a plurality of second vectors, the denoising unit 706 is specifically configured to:
[0127] For each key vector among the plurality of key vectors, selecting other vectors except the key vector from the plurality of key vectors as random vectors;
[0128] Inputting the key vector and the random vector into a third attention formula to perform attention calculation processing to obtain a third intermediate vector;
[0129] Inputting the third intermediate vector and the first vector into a fourth attention formula to perform attention calculation processing to obtain a fourth intermediate vector;
[0130] Inputting the fourth intermediate vector into a fifth attention formula to perform attention calculation processing to obtain a second vector corresponding to the key vector;
[0131] The second vectors corresponding to each key vector are integrated to obtain the plurality of second vectors.
[0132] In some embodiments, when executing the step of converting the plurality of key frames using the motion prediction subnetwork to obtain a plurality of converted frames, the conversion unit 708 is specifically configured to:
[0133] For each key frame except the first key frame and the last key frame among the plurality of key frames, obtaining two adjacent key frames of the key frame;
[0134] Using the third encoder to compress the two key frames to obtain two adjacent space vectors;
[0135] Performing expansion processing on the two adjacent space vectors based on the structure predictor in the motion prediction subnetwork to obtain a sequence vector;
[0136] Inputting the sequence vector into the conversion sub-block in the motion prediction sub-network for conversion processing to obtain a conversion frame;
[0137] The conversion frames corresponding to each key frame are integrated to obtain the plurality of conversion frames.
[0138] In some embodiments, when the encoding unit 701 performs the step of inputting the acquired reference image into the first encoder for encoding processing to obtain a latent vector, it is specifically configured to:
[0139] Inputting the reference image into the first encoder for encoding to obtain a mean vector and a variance vector;
[0140] The mean vector and the variance vector are used as the latent vector.
[0141] In some embodiments, before executing the step of using the acquired text prompt words to perform query processing in a preset dictionary library to obtain a text vector, the query unit 702 is further configured to:
[0142] The preset dictionary library is obtained by pre-training the dictionary library using the rule that each word corresponds to a vector.
[0143] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned video generation device and each unit can refer to the corresponding description in the aforementioned method embodiment. For the convenience and brevity of description, it will not be repeated here.
[0144] The video generation device can be implemented in the form of a computer program. Figure 3 Runs on the electronic devices shown.
[0145] See also Figure 3 , Figure 3 8 is a schematic block diagram of an electronic device provided by an embodiment of the present invention. The electronic device 800 can be a terminal or a server, wherein the terminal can be an electronic device with communication functions. The server can be a standalone server or a server cluster consisting of multiple servers.
[0146] See Figure 3 The electronic device 800 includes a processor 802 , a memory, and a network interface 805 connected via a system bus 801 , wherein the memory may include a non-volatile storage medium 803 and an internal memory 804 .
[0147] The non-volatile storage medium 803 may store an operating system 8031 and a computer program 8032. The computer program 8032 includes program instructions, which, when executed, may enable the processor 802 to perform a video generation method.
[0148] The processor 802 is used to provide computing and control capabilities to support the operation of the entire electronic device 800.
[0149] The internal memory 804 provides an environment for the operation of the computer program 8032 in the non-volatile storage medium 803. When the computer program 8032 is executed by the processor 802, the processor 802 can execute a video generation method.
[0150] The network interface 805 is used to communicate with other devices over the network. Figure 3The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention, and does not constitute a limitation on the electronic device 800 to which the solution of the present invention is applied. The specific electronic device 800 may include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0151] The processor 802 is configured to execute a computer program 8032 stored in the memory to implement the following steps:
[0152] Inputting the obtained reference image into the first encoder for encoding to obtain a latent vector;
[0153] Using the acquired text prompt words to query the preset dictionary library to obtain a text vector;
[0154] Performing attention calculation on the potential vector and the text vector using the reference sub-network to obtain a first vector;
[0155] Inputting the reference image and the text prompt words into the decomposition sub-network for decomposition processing to obtain a plurality of key stage texts;
[0156] Inputting the plurality of key stage texts into the second encoder for encoding to obtain a plurality of key vectors;
[0157] Using the denoising sub-network to perform denoising on the plurality of key vectors and the first vector to obtain a plurality of second vectors;
[0158] Inputting the plurality of second vectors into the decoder for decoding to obtain a plurality of key frames;
[0159] Using the motion prediction subnetwork to convert the plurality of key frames to obtain a plurality of conversion frames;
[0160] The plurality of converted frames are spliced together to obtain a target video.
[0161] In some embodiments, when implementing the step of using the reference sub-network to perform attention calculation on the potential vector and the text vector to obtain the first vector, the processor 802 is specifically configured to:
[0162] Inputting the potential vector into a first attention formula for calculation to obtain a first intermediate vector;
[0163] Inputting the first intermediate vector and the text vector into a second attention formula for calculation to obtain a second intermediate vector;
[0164] The first vector is obtained by performing calculations on the second intermediate vector and the text vector a preset number of times using the first attention formula and the second attention formula.
[0165] In some embodiments, when implementing the step of inputting the reference image and the text prompt word into the decomposition sub-network for decomposition processing to obtain a plurality of key stage texts, the processor 802 is specifically configured to:
[0166] Inputting the reference image and the text prompt word into the multimodal large language model in the decomposition sub-network for inference processing to obtain the motion process text;
[0167] The movement process text is decomposed based on the multimodal large language model to obtain the several key stage texts.
[0168] In some embodiments, when implementing the step of using the denoising sub-network to perform denoising on the plurality of key vectors and the first vector to obtain a plurality of second vectors, the processor 802 is specifically configured to:
[0169] For each key vector among the plurality of key vectors, selecting other vectors except the key vector from the plurality of key vectors as random vectors;
[0170] Inputting the key vector and the random vector into a third attention formula to perform attention calculation processing to obtain a third intermediate vector;
[0171] Inputting the third intermediate vector and the first vector into a fourth attention formula to perform attention calculation processing to obtain a fourth intermediate vector;
[0172] Inputting the fourth intermediate vector into a fifth attention formula to perform attention calculation processing to obtain a second vector corresponding to the key vector;
[0173] The second vectors corresponding to each key vector are integrated to obtain the plurality of second vectors.
[0174] In some embodiments, when implementing the step of converting the plurality of key frames using the motion prediction subnetwork to obtain a plurality of converted frames, the processor 802 is specifically configured to:
[0175] For each key frame except the first key frame and the last key frame among the plurality of key frames, obtaining two adjacent key frames of the key frame;
[0176] Using the third encoder to compress the two key frames to obtain two adjacent space vectors;
[0177] Performing expansion processing on the two adjacent space vectors based on the structure predictor in the motion prediction subnetwork to obtain a sequence vector;
[0178] Inputting the sequence vector into the conversion sub-block in the motion prediction sub-network for conversion processing to obtain a conversion frame;
[0179] The conversion frames corresponding to each key frame are integrated to obtain the plurality of conversion frames.
[0180] In some embodiments, when implementing the step of inputting the acquired reference image into the first encoder for encoding processing to obtain a latent vector, the processor 802 is specifically configured to:
[0181] Inputting the reference image into the first encoder for encoding to obtain a mean vector and a variance vector;
[0182] The mean vector and the variance vector are used as the latent vector.
[0183] In some embodiments, before implementing the step of using the acquired text prompt words to perform query processing in a preset dictionary library to obtain a text vector, the processor 802 is further configured to:
[0184] The preset dictionary library is obtained by pre-training the dictionary library using the rule that each word corresponds to a vector.
[0185] It should be understood that in the embodiment of the present invention, the processor 802 may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0186] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.
[0187] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor performs the following steps:
[0188] Inputting the obtained reference image into the first encoder for encoding to obtain a latent vector;
[0189] Using the acquired text prompt words to query the preset dictionary library to obtain a text vector;
[0190] Performing attention calculation on the potential vector and the text vector using the reference sub-network to obtain a first vector;
[0191] Inputting the reference image and the text prompt words into the decomposition sub-network for decomposition processing to obtain a plurality of key stage texts;
[0192] Inputting the plurality of key stage texts into the second encoder for encoding to obtain a plurality of key vectors;
[0193] Using the denoising sub-network to perform denoising on the plurality of key vectors and the first vector to obtain a plurality of second vectors;
[0194] Inputting the plurality of second vectors into the decoder for decoding to obtain a plurality of key frames;
[0195] Using the motion prediction subnetwork to convert the plurality of key frames to obtain a plurality of conversion frames;
[0196] The plurality of converted frames are spliced together to obtain a target video.
[0197] In one embodiment, when the processor executes the program instructions to implement the step of using the reference sub-network to perform attention calculation on the potential vector and the text vector to obtain the first vector, it is specifically configured to:
[0198] Inputting the potential vector into a first attention formula for calculation to obtain a first intermediate vector;
[0199] Inputting the first intermediate vector and the text vector into a second attention formula for calculation to obtain a second intermediate vector;
[0200] The first vector is obtained by performing calculations on the second intermediate vector and the text vector a preset number of times using the first attention formula and the second attention formula.
[0201] In one embodiment, when the processor executes the program instructions to implement the step of inputting the reference image and the text prompt word into the decomposition subnetwork for decomposition processing to obtain a plurality of key stage texts, the processor is specifically configured to:
[0202] Inputting the reference image and the text prompt word into the multimodal large language model in the decomposition sub-network for inference processing to obtain the motion process text;
[0203] The movement process text is decomposed based on the multimodal large language model to obtain the several key stage texts.
[0204] In one embodiment, when the processor executes the program instructions to implement the step of using the denoising sub-network to denoise the plurality of key vectors and the first vector to obtain a plurality of second vectors, the processor is specifically configured to:
[0205] For each key vector among the plurality of key vectors, selecting other vectors except the key vector from the plurality of key vectors as random vectors;
[0206] Inputting the key vector and the random vector into a third attention formula to perform attention calculation processing to obtain a third intermediate vector;
[0207] Inputting the third intermediate vector and the first vector into a fourth attention formula to perform attention calculation processing to obtain a fourth intermediate vector;
[0208] Inputting the fourth intermediate vector into a fifth attention formula to perform attention calculation processing to obtain a second vector corresponding to the key vector;
[0209] The second vectors corresponding to each key vector are integrated to obtain the plurality of second vectors.
[0210] In one embodiment, when the processor executes the program instructions to implement the step of converting the plurality of key frames using the motion prediction subnetwork to obtain a plurality of converted frames, the processor is specifically configured to:
[0211] For each key frame except the first key frame and the last key frame among the plurality of key frames, obtaining two adjacent key frames of the key frame;
[0212] Using the third encoder to compress the two key frames to obtain two adjacent space vectors;
[0213] Performing expansion processing on the two adjacent space vectors based on the structure predictor in the motion prediction subnetwork to obtain a sequence vector;
[0214] Inputting the sequence vector into the conversion sub-block in the motion prediction sub-network for conversion processing to obtain a conversion frame;
[0215] The conversion frames corresponding to each key frame are integrated to obtain the plurality of conversion frames.
[0216] In one embodiment, when the processor executes the program instructions to implement the step of inputting the acquired reference image into the first encoder for encoding processing to obtain a latent vector, the processor is specifically configured to:
[0217] Inputting the reference image into the first encoder for encoding to obtain a mean vector and a variance vector;
[0218] The mean vector and the variance vector are used as the latent vector.
[0219] In one embodiment, before executing the program instructions to implement the step of using the acquired text prompt words to perform query processing in a preset dictionary library to obtain a text vector, the processor is further configured to:
[0220] The preset dictionary library is obtained by pre-training the dictionary library using the rule that each word corresponds to a vector.
[0221] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.
[0222] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0223] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the various units is merely a logical functional division, and actual implementation may employ other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.
[0224] The steps in the methods of the embodiments of the present invention may be adjusted in order, combined, or deleted as needed. The units in the devices of the embodiments of the present invention may be combined, divided, or deleted as needed. Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.
[0225] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling an electronic device (such as a personal computer, terminal, or network device) to perform all or part of the steps of the method described in various embodiments of the present invention.
[0226] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A video generation method, characterized in that: Applied to a video generation network architecture, the video generation network architecture includes a first encoder, a reference sub-network, a decomposition sub-network, a second encoder, a denoising sub-network, a decoder and a motion prediction sub-network; the video generation method includes: Inputting the obtained reference image into the first encoder for encoding to obtain a latent vector; Using the acquired text prompt words to query the preset dictionary library to obtain the text vector; Performing attention calculation on the potential vector and the text vector using the reference sub-network to obtain a first vector; Inputting the reference image and the text prompt words into the decomposition sub-network for decomposition processing to obtain a plurality of key stage texts; Inputting the plurality of key stage texts into the second encoder for encoding to obtain a plurality of key vectors; Using the denoising sub-network to perform denoising on the plurality of key vectors and the first vector to obtain a plurality of second vectors; Inputting the plurality of second vectors into the decoder for decoding to obtain a plurality of key frames; Using the motion prediction subnetwork to convert the plurality of key frames to obtain a plurality of conversion frames; Splicing the plurality of converted frames to obtain a target video; Inputting the acquired reference image into the first encoder for encoding to obtain a latent vector includes: Inputting the reference image into the first encoder for encoding to obtain a mean vector and a variance vector; The mean vector and the variance vector are used as the latent vector.
2. The video generation method according to claim 1, wherein: The using the reference sub-network to perform attention calculation on the potential vector and the text vector to obtain a first vector includes: Inputting the potential vector into a first attention formula for calculation to obtain a first intermediate vector; Inputting the first intermediate vector and the text vector into a second attention formula for calculation to obtain a second intermediate vector; The first vector is obtained by performing calculations on the second intermediate vector and the text vector a preset number of times using the first attention formula and the second attention formula.
3. The video generation method according to claim 1, wherein: The reference image and the text prompt word are input into the decomposition sub-network for decomposition processing to obtain several key stage texts, including: Inputting the reference image and the text prompt word into the multimodal large language model in the decomposition sub-network for inference processing to obtain the motion process text; The movement process text is decomposed based on the multimodal large language model to obtain the several key stage texts.
4. The video generation method according to claim 1, wherein: The step of performing denoising on the plurality of key vectors and the first vector using the denoising sub-network to obtain a plurality of second vectors includes: For each key vector among the plurality of key vectors, selecting other vectors except the key vector from the plurality of key vectors as random vectors; Inputting the key vector and the random vector into a third attention formula to perform attention calculation processing to obtain a third intermediate vector; Inputting the third intermediate vector and the first vector into a fourth attention formula to perform attention calculation processing to obtain a fourth intermediate vector; Inputting the fourth intermediate vector into a fifth attention formula to perform attention calculation processing to obtain a second vector corresponding to the key vector; The second vectors corresponding to each key vector are integrated to obtain the plurality of second vectors.
5. The video generation method according to claim 1, wherein: The converting process of the plurality of key frames using the motion prediction subnetwork to obtain a plurality of converted frames includes: For each key frame except the first key frame and the last key frame among the plurality of key frames, obtaining two adjacent key frames of the key frame; Using a third encoder to compress the two key frames to obtain two adjacent space vectors; Performing expansion processing on the two adjacent space vectors based on the structure predictor in the motion prediction subnetwork to obtain a sequence vector; Inputting the sequence vector into the conversion sub-block in the motion prediction sub-network for conversion processing to obtain a conversion frame; The conversion frames corresponding to each key frame are integrated to obtain the plurality of conversion frames.
6. The video generation method according to claim 1, wherein: Before the obtained text prompt words are used to perform query processing in a preset dictionary library to obtain a text vector, the method further includes: The preset dictionary library is obtained by pre-training the dictionary library using the rule that each word corresponds to a vector.
7. A video generating device, characterized in that: Applied to a video generation network architecture, the video generation network architecture includes a first encoder, a reference sub-network, a decomposition sub-network, a second encoder, a denoising sub-network, a decoder and a motion prediction sub-network; the video generation device includes: A first encoding unit, configured to input the acquired reference image into the first encoder for encoding to obtain a latent vector; A query unit, configured to use the acquired text prompt words to perform query processing in a preset dictionary library to obtain a text vector; a computing unit, configured to perform attention calculation on the potential vector and the text vector using the reference sub-network to obtain a first vector; a decomposition unit, configured to input the reference image and the text prompt word into the decomposition sub-network for decomposition processing to obtain a plurality of key stage texts; A second encoding unit, configured to input the plurality of key stage texts into the second encoder for encoding to obtain a plurality of key vectors; a denoising unit, configured to perform denoising processing on the plurality of key vectors and the first vector using the denoising sub-network to obtain a plurality of second vectors; A decoding unit, configured to input the plurality of second vectors into the decoder for decoding to obtain a plurality of key frames; a conversion unit, configured to convert the plurality of key frames using the motion prediction subnetwork to obtain a plurality of conversion frames; a splicing unit, configured to splice the plurality of converted frames to obtain a target video; Inputting the acquired reference image into the first encoder for encoding to obtain a latent vector includes: Inputting the reference image into the first encoder for encoding to obtain a mean vector and a variance vector; The mean vector and the variance vector are used as the latent vector.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the video generation method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that The storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the video generation method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Training method, device and equipment for digital human model
CN119031210A
Video generation method and device, equipment and medium
CN119383289A