Reference image-based video generation method, device, equipment, medium and product
By obtaining the text information and multi-frame reference images of the target video, extracting and splicing features, and using the video generation model to generate the target video, the problem of the inability to generate videos of specific protagonists in existing technologies is solved, and the quality of video generation and user experience are improved.
Patent Information
- Application Number
- CN202510953667.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-10
AI Technical Summary
Existing video generation models are unable to generate videos with specific protagonists, resulting in a low match between the generated videos and user needs, affecting the user experience.
By obtaining the text information and multiple frames of reference images of the target video, extracting text features and image features, and using a pre-trained video generation model to perform splicing and denoising processing, the target video is generated.
The quality of generated videos and user experience are improved, ensuring that the generated videos match user needs.
Smart Images

Figure CN120455807B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of video generation, and in particular to a video generation method and device based on reference images, equipment, medium and product. BACKGROUND
[0002] In recent years, with the rapid development of video models, video model AIGC (Artificial Intelligence Generated Content) technology has made great achievements. In related technologies, a text-to-video model usually generates a corresponding video according to a text description of the video. However, in actual applications, sometimes a video of a specific main character needs to be generated, and the text-to-video model cannot generate a video of a specific main character according to a text description, thereby reducing the matching degree between the generated video and the video required by the user and affecting the user experience. SUMMARY
[0003] To solve the above technical problems, the present disclosure provides a video generation method and device based on reference images, equipment, medium and product.
[0004] In one aspect of the present disclosure, a video generation method based on reference images is provided, including: obtaining text information corresponding to a target video and a plurality of reference images, wherein each of the plurality of reference images includes a main character object of the target video; performing encoding processing on the text information to obtain text features; performing image feature extraction on each of the reference images to obtain image features of the reference images; performing splicing processing on each image feature to obtain spliced features; and performing denoising processing on the spliced features, the text features and a preset noise for a preset number of time steps by using a pre-trained video generation model to generate a target video for the main character object.
[0005] In another aspect of the present disclosure, a video generation device based on reference images is provided, including: a first data acquisition module configured to obtain text information corresponding to a target video and a plurality of reference images, wherein each of the plurality of reference images includes a main character object of the target video; an encoding module configured to perform encoding processing on the text information to obtain text features; a first feature extraction module configured to perform image feature extraction on each of the reference images to obtain image features of the reference images; a first feature splicing module configured to perform splicing processing on each image feature to obtain spliced features; and a video generation module configured to perform denoising processing on the spliced features, the text features and a preset noise for a preset number of time steps by using a pre-trained video generation model to generate a target video for the main character object.
[0006] In still another aspect of the embodiments of the present disclosure, an electronic device is provided, comprising: a memory for storing a computer program; and a processor for executing the computer program stored in the memory, and the computer program, when executed, implements the method described above.
[0007] In still another aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, having a computer program stored thereon, and the computer program, when executed by a processor, implements the method described above.
[0008] In still another aspect of the embodiments of the present disclosure, a computer program product is provided, comprising computer program instructions, and the computer program instructions, when executed by a processor, implement the method described above.
[0009] In the embodiments of the present disclosure, by splicing the image features of the multiple frames of reference images including the main character object in the target video, the spliced features containing the information of the main character object in all aspects are obtained, so that when the target video is generated, the video generation model can pay attention to the text features and the spliced features at the same time, and can better learn the main character object and the text information to generate the target video about the main character object, thereby improving the quality of the generated target video and enhancing the user experience.
[0010] The technical solutions of the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0011] The accompanying drawings, which form a part of the specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0012] The present disclosure can be more clearly understood and appreciated from the following detailed description, taken in conjunction with the accompanying drawings, in which:
[0013] Figure 1 is a flowchart of a video generation method based on reference images provided by an exemplary embodiment of the present disclosure;
[0014] Figure 2 is a schematic diagram of a video generation method based on reference images provided by an application example of the present disclosure;
[0015] Figure 3 is a flowchart of a video generation method based on reference images provided by another exemplary embodiment of the present disclosure;
[0016] Figure 4 is a structural schematic diagram of an embodiment of a video generation device based on reference images of the present disclosure;
[0017] Figure 5 is a structural schematic diagram of an application embodiment of an electronic device of the present disclosure. DETAILED DESCRIPTION
[0018] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. Note that the relative arrangement, numerical expressions, and numerical values of components and steps set forth in these embodiments are not limiting to the scope of the present disclosure unless specifically stated otherwise.
[0019] Those skilled in the art can understand that the terms "first", "second", and the like in the embodiments of the present disclosure are only used to distinguish different steps, devices, or modules, and do not represent any specific technical meaning, nor do they represent a necessary logical sequence between them.
[0020] It should also be understood that in the embodiments of the present disclosure, "multiple" can mean two or more, and "at least one" can mean one, two, or more.
[0021] It should also be understood that for any component, data, or structure mentioned in the embodiments of the present disclosure, unless specifically limited or the context gives a contrary implication, it can be understood as one or more in general.
[0022] In addition, the term "and / or" in the present disclosure is only a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can represent the existence of A alone, the existence of A and B together, and the existence of B alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the front and rear associated objects.
[0023] It should also be understood that the description of various embodiments of the present disclosure focuses on the differences between various embodiments, and the same or similar parts can be referred to each other, and for the sake of brevity, will not be repeated.
[0024] At the same time, it should be understood that, for the sake of brevity, the size of each part shown in the drawings is not drawn in accordance with the actual proportional relationship.
[0025] The following description of at least one exemplary embodiment is merely illustrative in nature and does not in any way limit the disclosure and its application or uses.
[0026] Techniques, methods, and devices known to those of ordinary skill in the relevant art can not be discussed in detail, but should be considered part of the specification where appropriate.
[0027] It should be noted that similar reference numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0028] The embodiments of the present disclosure can be applied to terminal devices, computer systems, servers, and other electronic devices, which can operate with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with terminal devices, computer systems, servers, and other electronic devices include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, and the like.
[0029] The terminal devices, computer systems, servers, and other electronic devices can be described in the general context of computer system executable instructions, such as program modules, which are executed by the computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, and the like, which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment, in which tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules can be located on local or remote computer system storage media including storage devices.
[0030] In the process of implementing the present disclosure, the inventors found that in practical applications, in some video generation application scenarios, it is necessary to generate a video including a specific character. For example, a video creator wants to create a video with himself as the main character, or an animation creator draws a specific main character and needs to create a complete animation video of the specific main character. However, the existing video model cannot generate a video including a specific main character, which leads to a low matching degree between the video generated by the text generation video model and the video required by the user, affecting the user experience.
[0031] Figure 1 FIG. 1 is a flowchart of a video generation method based on reference images provided by an example embodiment of the present disclosure. The present embodiment can be applied to an electronic device, such as a terminal device, a computer system, a server, and the like. Figure 1 As shown in FIG. 1, the video generation method based on reference images can include the following steps:
[0032] Step S100, obtaining text information corresponding to a target video and a plurality of reference images.
[0033] The plurality of reference images each include a main character object in the target video.
[0034] In an embodiment, the main character object in the target video can be understood as a main character in the target video, etc. Illustratively, the main character object can include at least one of the following: a cartoon character, an animal, a real person, a mythological or legendary character, a building, a plant, etc. The multi-frame reference image can include information of at least one aspect of the main character object. For example, the main character object is Diaochan, and the multi-frame reference image can include Diaochan in different perspectives, respectively. For example, the first frame reference image is a front view of Diaochan, the second frame reference image can be a left side view of Diaochan, the third frame reference image can be a right side view of Diaochan, and the fourth frame reference image can include a back view of Diaochan, etc. The text information includes information describing the video content of the target video. In an implementation, the reference image can include multiple main character objects of the target video.
[0035] The user can manually draw the multi-frame reference image including the main character object, or generate the multi-frame reference image including the main character object through image generation software such as Midjourney, Adobe Firefly, Bing Image Creator, etc.
[0036] In step S110, the text information is encoded to obtain text features.
[0037] The text features are vector representations extracted from the text information to indicate the content information of the target video. The text information can be encoded by a text encoder to obtain the text features of the text information. The text encoder can use a T5 (Text To Text Transfer Transformer) model, etc.
[0038] In step S120, image features of each reference image are extracted to obtain image features of each reference image.
[0039] The image features of each reference image are vector representations extracted from the reference image to represent its visual content, spatial structure, and semantic information. The reference image can be feature-extracted based on a pre-trained deep learning model to obtain the image features of the reference image. The deep learning model can use a convolutional neural network (CNN) or a recurrent neural network (RNN), etc.
[0040] In step S130, the image features are spliced to obtain spliced features.
[0041] The image features can be spliced to obtain the spliced features. Illustratively, the image features can be spliced in the channel dimension to obtain the spliced features.
[0042] In step S140, based on the spliced feature, the text feature, and the preset noise, a preset number of time steps of denoising processing is performed by using a pre-trained video generation model to generate a target video for the main character object.
[0043] The target video includes the main character object. That is, the target video is a video about the main character object. The preset noise can include Gaussian noise. For example, a Gaussian noise matrix can be randomly generated in advance, and the Gaussian noise matrix is determined as the preset noise.
[0044] In an embodiment, the video generation model can adopt a diffusion model. The spliced feature, the text feature, and the preset noise can be input into the video generation model, and the video generation model performs the preset number of time steps of denoising processing to output the target video.
[0045] In the embodiments of the present disclosure, the image features of the multiple frames of reference images including the main character object in the target video are spliced to obtain the spliced feature containing the information of the main character object. Thus, when generating the target video, the video generation model can simultaneously focus on the text feature and the spliced feature, and can better learn the main character object and the text information to generate the target video about the main character object. Thus, the quality of the generated target video is improved, and the user experience is improved.
[0046] In some optional embodiments, step S120 in the embodiments of the present disclosure can include: for each reference image, performing image feature extraction on the reference image by using a pre-trained multi-modal model to obtain the image feature of the reference image.
[0047] The image feature of the reference image includes a classification token of the reference image.
[0048] The classification token (CLS token) in image processing is used to summarize the global features of the image. The multi-modal model can adopt a contrastive language-image pretraining (CLIP) model. Each reference image can be input into the CLIP model, and the CLIP model outputs the CLS token of the reference image.
[0049] Correspondingly, in the embodiments of the present disclosure, step S130 can include: performing splicing processing on each classification token to obtain the spliced feature.
[0050] In the embodiments of the present disclosure, the classification labels of the reference images are spliced to obtain the spliced features, thereby reducing the size of the spliced features while ensuring that the spliced features can include the features of the main character object, improving the efficiency of processing the spliced features, and further improving the efficiency of generating the target video.
[0051] In some optional embodiments, in the embodiments of the present disclosure, the video generation model includes a cross-attention module. The cross-attention module includes a query feature mapping layer, a text key feature mapping layer, a text value feature mapping layer, an image key feature mapping layer, and an image value feature mapping layer. The cross-attention module is configured to perform cross-attention calculation on data. The query feature mapping layer, the text key feature mapping layer, the text value feature mapping layer, the image key feature mapping layer, and the image value feature mapping layer are configured to perform feature mapping processing.
[0052] In one embodiment, the video generation model is a diffusion model based on a Transformer structure (Diffusion Transformer, DiT). Compared with a traditional deep neural network model structure, the video generation model in the embodiment has stronger cross-modal information alignment capability and complex semantic control capability in a video generation task. Specifically, the video generation model includes a cross-attention module, and the cross-attention module includes a query feature mapping layer, a text key feature mapping layer, a text value feature mapping layer, an image key feature mapping layer, and an image value feature mapping layer. In the denoising processing at each time step: the cross-attention module performs cross-attention calculation on the spliced features, the text features, and the preset noise to obtain cross-attention output, so as to determine the target video based on the denoising processing at each time step and the cross-attention output at each time step. The cross-attention module is configured to perform cross-attention calculation on data. The query feature mapping layer, the text key feature mapping layer, the text value feature mapping layer, the image key feature mapping layer, and the image value feature mapping layer are configured to perform feature mapping processing. Through the video generation model with the above model structure, cross-model fusion between the spliced features and the text features can be achieved.
[0053] Specifically, the preset noise is input into the pre-trained video generation model, and the spliced features and the text features are used as conditions of the pre-trained video generation model. Through denoising processing of the preset number of time steps, a target video that conforms to the spliced image and the text information semantics can be obtained. In the denoising processing at each time step, cross-attention calculation is performed by the cross-attention module in the video generation model, and the output of the video generation model at each time step T can be used as the input of the next time step T+1. Based on the processing of the last time step, the video generation model can output the generated target video.
[0054] Cross-Attention is a deep learning technique that allows dynamic association between different input feature sequences. By interacting query features (Query) with key-value features (Key-Value), it realizes cross-modal or cross-sequence information fusion. In this embodiment, based on the query feature mapping layer in the video generation model and the preset noise, the query feature can be obtained. Based on the text key feature mapping layer, the text value feature mapping layer, the image key feature mapping layer and the image value feature mapping layer, the text feature and the image feature can be mapped into text key-value features and image key-value features respectively, and then the cross-modal semantic dependency relationship can be established.
[0055] In one embodiment, the structure of the video generation model can adopt the structure of the diffusion model. The video generation model can include a plurality of network layers, each network layer adopting a Transformer network, and each network layer being provided with a cross-attention module.
[0056] In one embodiment, in the denoising process at each time step, the cross-attention module in the Lth network layer performs cross-attention calculation on the spliced features, the text features and the preset noise to obtain cross-attention output, and then the cross-attention output is taken as the input of the (L+1)th network layer. Then the cross-attention module in the (L+1)th network layer repeats the operation of the cross-attention module in the Lth network layer, and the cross-attention output output by the cross-attention module of the last network layer is determined as the output of the time step.
[0057] In the embodiments of the present disclosure, by introducing spliced features in the target video generation process, and setting the cross-attention module in the video generation model to include a query feature mapping layer, a text key feature mapping layer, a text value feature mapping layer, an image key feature mapping layer and an image value feature mapping layer. Thereby, the main character object in each reference image can be better focused during the video generation process, and the reference image and the text information can be better fused, so as to generate a target video that conforms to the text information, the video content is continuous and is about the main character object, thereby improving the quality of the generated target video, ensuring the matching degree between the target video and the user needs, and improving the user experience.
[0058] In some optional embodiments, in the embodiments of the present disclosure, the cross-attention calculation on the spliced feature, the text feature and the preset noise by the cross-attention module can include: performing feature mapping on the query feature based on the preset noise by using a query feature mapping layer to obtain a query feature (Quary), performing feature mapping on the text feature by using a text key feature mapping layer and a text value feature mapping layer respectively to obtain a first key feature (Key1) and a first value feature (Value1), performing feature mapping on the spliced feature by using an image key feature mapping layer and an image value feature mapping layer respectively to obtain a second key feature (Key2) and a second value feature (Value2), and performing cross-attention mechanism calculation on the query feature, the first key feature, the first value feature, the second key feature and the second value feature to obtain a cross-attention output.
[0059] In one embodiment, the query feature can represent the state in the current video generation process. The first key feature and the first value feature can represent the key information of the text information. The second key feature and the second value feature represent the key information of the multi-frame reference image.
[0060] Correspondingly, in the embodiments of the present disclosure, the cross-attention mechanism calculation on the query feature, the first key feature, the first value feature, the second key feature and the second value feature can include:
[0061] The cross-attention calculation on the query feature, the first key feature and the first value feature obtains a first sub-fusion feature, the cross-attention calculation on the query feature, the second key feature and the second value feature obtains a second sub-fusion feature, and the cross-attention output is obtained based on the first sub-fusion feature and the second sub-fusion feature.
[0062] In one embodiment, the matrix multiplication operation is performed on the query feature and the first key feature to obtain a first attention feature map (Attention Map), and then the multiplication operation is performed on the first attention feature map and the first value feature to obtain the first sub-fusion feature. The matrix multiplication operation is performed on the query feature and the second key feature to obtain a second attention feature map (CilpAttention Map), and then the multiplication operation is performed on the second attention feature map and the second value feature sequence to obtain the second sub-fusion feature. The cross-attention output can be obtained by multiplying or weighting and then multiplying the first sub-fusion feature and the second sub-fusion feature.
[0063] Correspondingly, in one embodiment, the cross-attention output based on the first sub-fusion feature and the second sub-fusion feature can include: performing linear calculation on the first sub-fusion feature and the second sub-fusion feature to obtain the cross-attention output.
[0064] For example, the first sub-fusion feature and the second sub-fusion feature can be added to obtain the cross-attention output.
[0065] In one embodiment, the preset noise is initialized and the initialized preset noise x is input into a to_q layer (query feature mapping layer), and the to_q layer outputs a query matrix Q (query feature); the spliced features are respectively input into a to_k layer (text key feature mapping layer) and a to_v layer (text value feature mapping layer), the to_k layer outputs K (first key feature), and the to_v layer outputs V (first value feature); the image features are respectively input into a to_k_clip layer (image key feature mapping layer) and a to_v_clip layer (image value feature mapping layer), the to_k_clip layer outputs K_clip (second key feature), and the to_v_clip layer outputs V_clip (second value feature). First, based on Q, K and V, cross attention calculation is performed by using formula (1) to obtain a first sub fusion feature, that is, Q and K matrices are multiplied to obtain an Attention Map (first attention feature map), and then multiplied with V; then, based on Q, K_clip and V_clip, cross attention calculation is performed by using formula (2) to obtain a second sub fusion feature, that is, Q and K_clip matrices are multiplied to obtain a Cilp Attention Map (second attention feature map), and then multiplied with V_clip, the first sub fusion feature and the second sub fusion feature are added to obtain a cross attention output.
[0066] Formula (1)
[0067] Formula (2)
[0068] In formulas (1) and (2), Attention(Q,K,V) represents the first sub fusion feature, represents the vector dimension (depth) of K, and T represents transposition calculation, represents the second sub fusion feature, represents the vector dimension (depth) of K.
[0069] In the embodiments of the present disclosure, the query feature corresponding to the preset noise, the first key feature and the first value feature corresponding to the text feature, and the second key feature and the second value feature corresponding to the splicing feature are respectively generated by querying the feature mapping layer, the text key feature mapping layer, the text value feature mapping layer, the image key feature mapping layer and the image value feature mapping layer, and the cross-attention calculation processing is performed on the query feature, the first key feature, the first value feature, the second key feature and the second value feature, so that the video generation model can effectively pay attention to and align the text information and the reference image from different sources, thereby helping the video generation model to better capture the correlation between the main character object in the text information and the reference image, so that the generated target video can meet the requirements of matching the text information and being about the main character object.
[0070] Exemplary, Figure 2 is a schematic diagram of a video generation method based on a reference image provided by an application example of the present disclosure. As Figure 2 shown, the video generation model includes n network layers, wherein at least one network layer includes a cross-attention module. In this example, the network layer adopts a Transformer network. The cross-attention module includes a query feature mapping layer (not shown in the figure), a text key feature mapping layer (not shown in the figure), a text value feature mapping layer (not shown in the figure), an image key feature mapping layer (not shown in the figure) and an image value feature mapping layer (not shown in the figure).
[0071] The text information, the plurality of reference images and the preset noise are obtained, and the plurality of reference images each include a main character object of the target video. The text information is encoded to obtain a text feature. Each reference image is input into a CLIP model, and the CLIP model outputs a CLS token of each reference image. The CLS tokens are spliced to obtain a splicing feature. The splicing feature and the preset noise are input into an encoder of a Variational Auto-Encoder (VAE) for encoding, and then the splicing feature, the text feature and the preset noise are input into a video generation model for denoising processing at m time steps (a preset number of time steps), to determine the target video.
[0072] In the denoising processing of the i-th (0 < i < m, i is an integer) time step: in the cross-attention module in the f-th (0 < f < n, f is an integer) network layer, the query feature is obtained by performing feature mapping on the preset noise based on the query feature mapping layer, the first key feature and the first value feature are obtained by performing feature mapping on the text feature based on the text key feature mapping layer and the text value feature mapping layer respectively, the second key feature and the second value feature are obtained by performing feature mapping on the spliced feature based on the image key feature mapping layer and the image value feature mapping layer respectively, the first attention feature map is obtained by performing matrix multiplication operation on the query feature and the first key feature, then the first sub-fusion feature is obtained by performing multiplication operation on the first attention feature map and the first value feature; the second attention feature map is obtained by performing matrix multiplication operation on the query feature and the second key feature, then the second sub-fusion feature is obtained by performing multiplication operation on the second attention feature map and the second value feature sequence, the cross-attention output is obtained by adding the first sub-fusion feature and the second sub-fusion feature, the cross-attention output is taken as the input of the cross-attention module in the (f+1)-th network layer, the above operation is repeated, the cross-attention output output by the cross-attention module of the n-th network layer is determined as the cross-attention output of the i-th time step, the cross-attention output of the i-th time step is taken as the preset noise in the denoising processing of the (i+1)-th time step, and the target video is determined based on the cross-attention output of the denoising processing of the m-th time step. For example, the cross-attention output of the denoising processing of the m-th time step can be input into the decoder of the VAE for decoding processing to obtain the target video.
[0073] Figure 3 is a flowchart of a video generation method based on a reference image provided by another exemplary embodiment of the present disclosure. In some optional embodiments, as shown in Figure 3 , the video generation model is obtained by the following way:
[0074] In step S200, a to-be-trained model and sample data are obtained.
[0075] The sample data includes: a plurality of sample videos, and a label text, a plurality of frames of label images and a label noise corresponding to any sample video, the plurality of frames of label images corresponding to any sample video all include a main character object in the sample video, the label noise corresponding to any sample video is generated by adding noise to the sample video, and the to-be-trained model includes a cross-attention module, and the cross-attention module includes a query feature mapping layer, a text key feature mapping layer, a text value feature mapping layer, an image key feature mapping layer and an image value feature mapping layer.
[0076] In an embodiment, a plurality of video clips can be acquired, and the plurality of video clips can be cropped to a fixed resolution and length to obtain a plurality of sample videos. The sample videos can be input into a vision-language model (VLM), and the VLM can output label texts corresponding to the sample videos. The sample videos can be processed by adding random Gaussian noise to obtain label noise corresponding to the sample videos. A plurality of video images including a main character object can be selected from the sample videos as label images, or the plurality of video images including the main character object can be generated by using image generation software such as Midjourney, Adobe Firefly, and Bing Image Creator.
[0077] In step S210, for a plurality of sample videos, a plurality of label images and label texts corresponding to the sample videos are subjected to feature extraction to obtain a plurality of label image features and label text features corresponding to the sample videos.
[0078] In the above embodiment, the CLIP model can be used to input each label image, and the CLIP model can output a CLS token of each label image as an image feature of each label image. A text encoder can be used to encode the label texts to obtain label text features.
[0079] In step S220, a plurality of label image features are spliced to obtain a label splicing feature corresponding to the sample video.
[0080] In the above embodiment, the CLS tokens of the plurality of label images can be spliced to obtain the label splicing feature corresponding to the sample video.
[0081] In step S230, the sample video, the label noise corresponding to the sample video, the label splicing feature, and the label text feature are input into a to-be-trained model, and the to-be-trained model is subjected to noise addition processing for a preset number of time steps to output a predicted noise.
[0082] In the above embodiment, in the noise addition processing at each time step, a cross-attention module performs cross-attention calculation on the label splicing feature, the label text feature, and the sample video to obtain a label cross-attention output, so as to determine the predicted noise based on the noise addition processing at each time step and the label cross-attention output at each time step.
[0083] Specifically, based on the sample video, feature mapping is performed using the query feature mapping layer to obtain the label query feature, based on the label text feature, feature mapping is performed using the text key feature mapping layer and the text value feature mapping layer respectively to obtain the label first key feature and the label first value feature, based on the label splicing feature, feature mapping is performed using the image key feature mapping layer and the image value feature mapping layer respectively to obtain the label second key feature and the label second value feature, cross-attention calculation is performed on the label query feature, the label first key feature, and the label first value feature to obtain the label first sub-fusion feature, cross-attention calculation is performed on the label query feature, the label second key feature, and the label second value feature to obtain the label second sub-fusion feature, linear calculation is performed on the label first sub-fusion feature and the label second sub-fusion feature to obtain the label cross-attention output.
[0084] For example, Figure 2 The corresponding video generation model is used as an example to illustrate. In the noise addition processing at the i-th time step: in the cross attention module in the f-th network layer, based on the sample video, the query feature mapping layer is used to perform feature mapping to obtain the label query feature, based on the label text feature, the text key feature mapping layer and the text value feature mapping layer are used to perform feature mapping respectively to obtain the label first key feature and the label first value feature, based on the label splicing feature, the image key feature mapping layer and the image value feature mapping layer are used to perform feature mapping respectively to obtain the label second key feature and the label second value feature, the label query feature and the label first key feature are matrix multiplied to obtain the label first attention feature map, and then the label first attention feature map and the label first value feature are multiplied to obtain the label first sub-fusion feature; the label query feature and the label The second key feature of the label is matrix multiplied to obtain the second attention feature map of the label, and then the second attention feature map of the label is multiplied with the second value feature sequence of the label to obtain the second sub-fusion feature of the label, and the first sub-fusion feature of the label and the second sub-fusion feature of the label are added to obtain the label cross-attention output, and the label cross-attention output is used as the input of the cross-attention module in the f+1th network layer. The above operation is repeated, and the label cross-attention output output of the cross-attention module of the nth network layer is determined as the label cross-attention output of the i-th time step, and the label cross-attention output of the i-th time step is used as the sample video in the noise addition processing of the i+1th time step. The predicted noise is determined based on the label cross-attention output of the m-th time step. Exemplarily, the label cross-attention output of the m-th time step can be determined as the predicted noise.
[0085] In step S240 , the parameters of the model to be trained are adjusted according to the predicted noises and the corresponding label noises until a preset training end condition is met, and a video generation model is obtained from the model to be trained.
[0086] A loss function value can be determined based on the difference between the label noise and the predicted noise corresponding to each sample video using a preset loss function. Examples of the preset loss function include, but are not limited to, a cross-entropy error function or a mean square error function. The loss function value determination operation can be iteratively performed, and the parameters of the model to be trained can be iteratively adjusted to continuously reduce the loss function value until the loss function value converges. Preset training termination conditions are then determined to be met, and the training of the model to be trained is completed. The trained model to be trained is then used as the video generation model. Parameter optimizers such as stochastic gradient descent (SGD), adaptive gradient algorithm (Adagrad), adaptive moment estimation (Adam), and root mean square prop (RMSprop) can be used to adjust the parameters of the model to be trained. For example, the parameter optimizer can be used to calculate the gradient of each parameter of the model to be trained, and the parameters can be adjusted along the gradient direction, where the gradient represents the direction of the greatest loss function value reduction. The loss function value determination operation is iteratively performed until the loss function value no longer decreases, thus completing the training of the model to be trained and obtaining the video generation model.
[0087] In the embodiment of the present disclosure, the training model is trained using label splicing features including features of the protagonist object, label text features, and sample videos. This not only enables efficient training of the training model, but also enables the trained video generation model to generate videos about the target object.
[0088] Figure 4 FIG. 1 is a schematic diagram of a structure of an embodiment of a video generation device based on a reference image disclosed in the present invention. Figure 4 As shown, the device of this embodiment may include:
[0089] A first data acquisition module 300 is configured to acquire text information and multiple reference images corresponding to a target video, wherein each of the multiple reference images includes a main character object of the target video;
[0090] The encoding module 310 is used to encode the text information to obtain text features;
[0091] A first feature extraction module 320 is configured to extract image features from each reference image to obtain image features of each reference image;
[0092] A first feature stitching module 330 is used to stitch the image features to obtain stitching features;
[0093] The video generation module 340 is configured to perform preset time step denoising processing on the spliced feature, the text feature and preset noise by using a pre-trained video generation model, and generate a target video for the main character object.
[0094] In some possible implementation manners of the present disclosure, in the embodiments of the present disclosure,
[0095] The first feature extraction module 320 is configured to perform image feature extraction on the reference images by using a pre-trained multi-modal model to obtain image features of the reference images, wherein the image features of the reference images include classification labels of the reference images.
[0096] The first feature splicing module 330 is configured to splice the classification labels to obtain the spliced feature.
[0097] In some possible implementation manners of the present disclosure, the video generation model in the embodiments of the present disclosure includes a cross-attention module, the cross-attention module includes a query feature mapping layer, a text key feature mapping layer, a text value feature mapping layer, an image key feature mapping layer and an image value feature mapping layer; in each time step of denoising processing: the cross-attention module performs cross-attention calculation on the spliced feature, the text feature and the preset noise to obtain cross-attention output, so as to determine the target video based on the denoising processing of each time step and the cross-attention output of each time step.
[0098] In some possible implementation manners of the present disclosure, the cross-attention module in the embodiments of the present disclosure performs cross-attention calculation on the spliced feature, the text feature and the preset noise, and is further configured to:
[0099] perform feature mapping on the preset noise by using the query feature mapping layer to obtain query features;
[0100] perform feature mapping on the text feature by using the text key feature mapping layer and the text value feature mapping layer respectively to obtain first key features and first value features;
[0101] perform feature mapping on the spliced feature by using the image key feature mapping layer and the image value feature mapping layer respectively to obtain second key features and second value features;
[0102] perform cross-attention mechanism calculation on the query features, the first key features, the first value features, the second key features and the second value features to obtain the cross-attention output.
[0103] In some possible implementation manners of the present disclosure, the cross-attention mechanism calculation on the query feature, the first key feature, the first value feature, the second key feature, and the second value feature in the embodiments of the present disclosure is further used for:
[0104] performing cross-attention calculation on the query feature, the first key feature, and the first value feature to obtain a first sub-fusion feature;
[0105] performing cross-attention calculation on the query feature, the second key feature, and the second value feature to obtain a second sub-fusion feature;
[0106] obtaining the cross-attention output based on the first sub-fusion feature and the second sub-fusion feature.
[0107] In some possible implementation manners of the present disclosure, the obtaining of the cross-attention output based on the first sub-fusion feature and the second sub-fusion feature in the embodiments of the present disclosure is further used for:
[0108] performing linear calculation on the first sub-fusion feature and the second sub-fusion feature to obtain the cross-attention output.
[0109] In some possible implementation manners of the present disclosure, the video generation apparatus based on a reference image in the embodiments of the present disclosure further includes:
[0110] a second data acquisition module configured to acquire a to-be-trained model and sample data, wherein the sample data includes a plurality of sample videos, and label text, a plurality of frames of label images, and label noise corresponding to any sample video, the plurality of frames of label images corresponding to the any sample video each include a main character object in the any sample video, the label noise is generated by adding noise to the any sample video, and the to-be-trained model includes a cross-attention module, the cross-attention module includes a query feature mapping layer, a text key feature mapping layer, a text value feature mapping layer, an image key feature mapping layer, and an image value feature mapping layer;
[0111] a second feature extraction module configured to perform feature extraction on a plurality of label images and label text corresponding to the sample video for the plurality of sample videos to obtain a plurality of label image features and label text features corresponding to the sample video;
[0112] a second feature splicing module configured to splice the plurality of label image features to obtain label splicing features corresponding to the sample video;
[0113] A first training module is configured to input the sample video, and the label noise, label splicing features, and label text features corresponding to the sample video, into the to-be-trained model, and the to-be-trained model performs a noise addition process for a preset number of time steps and outputs a predicted noise, wherein, in the noise addition process at each time step, the cross-attention module performs a cross-attention calculation on the label splicing features, the label text features, and the sample video to obtain a label cross-attention output, so as to determine the predicted noise based on the noise addition process at each time step and the label cross-attention output at each time step;
[0114] The second training module is used to adjust the parameters of the model to be trained according to each predicted noise and the corresponding label noise until a preset training end condition is met, and the video generation model is obtained from the model to be trained.
[0115] The video generation device based on reference images in the embodiment of the present disclosure corresponds to the embodiment of the video generation method based on reference images in the present disclosure. The relevant contents can be referenced to each other and will not be repeated here.
[0116] The beneficial technical effects corresponding to the exemplary embodiment of the reference image-based video generation device of the embodiment of the present disclosure can be found in the corresponding beneficial technical effects of the above-mentioned corresponding exemplary method part, which will not be repeated here.
[0117] In addition, an embodiment of the present disclosure further provides an electronic device, including:
[0118] memory for storing computer programs;
[0119] The processor is configured to execute the computer program stored in the memory, and when the computer program is executed, the method for generating a video based on a reference image as described in any one of the above embodiments of the present disclosure is implemented.
[0120] Figure 5 This is a schematic diagram of the structure of an application embodiment of an electronic device disclosed herein. Below, referring to Figure 5, an electronic device according to an embodiment of the present disclosure is described. The electronic device can be either or both of the first and second devices, or a standalone device independent of them. The standalone device can communicate with the first and second devices to receive collected input signals from them.
[0121] like Figure 5 As shown, the electronic device includes one or more processors and memory.
[0122] The processor may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.
[0123] The memory can include one or more computer program products that can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory, for example, can include random access memory (RAM), and / or a cache, etc. The non-volatile memory, for example, can include read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage media, and the processor can execute the program instructions to implement the reference image-based video generation method of various embodiments of the disclosure described above and / or other desired functions.
[0124] In one example, the electronic device can further include an input device and an output device, and these components are interconnected through a bus system and / or other forms of connection mechanism (not shown).
[0125] In addition, the input device can further include, for example, a keyboard, a mouse, etc.
[0126] The output device can output various information, including the determined distance information, direction information, etc., to the outside. The output device can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, etc.
[0127] Of course, in order to simplify, Figure 5 In FIG. 1, only some of the components related to the present disclosure among the components in the electronic device are illustrated, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device can further include any other appropriate components according to a specific application.
[0128] In addition to the above-described method and device, an embodiment of the present disclosure can be a computer program product including computer program instructions that, when executed by a processor, cause the processor to perform steps of the reference image-based video generation method according to various embodiments of the present disclosure described in the above parts of the specification.
[0129] The computer program product can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, C++, etc., and conventional procedural programming languages, such as the "C" programming language, or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device, or entirely on the remote computing device or server.
[0130] Furthermore, an embodiment of the present disclosure can also be a computer readable storage medium having stored thereon computer program instructions which, when executed by a processor, cause the processor to perform the steps described in the foregoing detailed description of the reference image based video generation method according to various embodiments of the present disclosure.
[0131] The computer readable storage medium can be any combination of one or more computer readable medium(s). The computer readable medium can be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0132] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by program instruction related hardware, and the foregoing program can be stored in a computer readable storage medium. When the program is executed, the steps of the method embodiments are executed. The foregoing storage medium includes ROM, RAM, magnetic disc or optical disc and various storage media that can store program codes.
[0133] The above describes the basic principles of the present disclosure in combination with specific embodiments. However, it should be noted that the advantages, advantages, effects and the like mentioned in the present disclosure are only examples and are not limiting. These advantages, advantages, effects and the like cannot be considered as the must-have of each embodiment of the present disclosure. In addition, the above specific details are only for the purpose of example and understanding, and the above details do not limit the present disclosure to the must-use specific details.
[0134] Each embodiment in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same or similar parts of each embodiment can be referred to each other. For system embodiments, since they basically correspond to method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.
[0135] The block diagrams of devices, apparatuses, equipment, systems referred to in this disclosure are merely illustrative examples and are not intended to require or imply that the connection, arrangement, configuration must be as shown in the block diagrams. These devices, apparatuses, equipment, systems can be connected, arranged, configured in any manner as will be appreciated by those skilled in the art. Words such as "include," "contain," "have," and the like are open-ended words that are to be interpreted to mean "including but not limited to," and are not to be interpreted as limiting the described embodiment to features, elements, and / or steps disclosed herein. The words "or" and "and" as used herein are to be interpreted as the word "and / or," and are not to be interpreted as requiring both features, elements, and / or steps disclosed herein. The word "such as" as used herein is to be interpreted as the phrase "such as but not limited to," and is not to be interpreted as limiting the described embodiment to features, elements, and / or steps disclosed herein.
[0136] The methods and apparatuses of this disclosure can be implemented in a number of ways. For example, the methods and apparatuses of this disclosure can be implemented using software, hardware, firmware, or any combination of these. The above described order of steps for the methods is merely illustrative, and the steps of the methods of this disclosure are not limited to the order specifically described above unless otherwise specifically stated. Furthermore, in some embodiments, the disclosure can also be implemented as a program recorded in a recording medium, which includes machine readable instructions for implementing the methods according to the disclosure. Thus, the disclosure also covers a recording medium storing a program for executing the methods according to the disclosure.
[0137] It is also important to note that the devices, equipment, and methods of this disclosure can be embodied in a variety of ways. These variations are contemplated as being within the scope of the present disclosure.
[0138] The above description of the disclosed aspects is given for illustrative purposes and is not intended to limit the scope of the disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0139] The above description has been given for illustrative and descriptive purposes. In addition, this description is not intended to limit embodiments of the disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those of skill in the art will recognize certain modifications, permutations, additions, and sub-combinations thereof.
Claims
1. A video generation method based on a reference image, characterized in that: include: Acquire text information and multiple reference images corresponding to a target video, wherein the multiple reference images each include a protagonist object of the target video; Encoding the text information to obtain text features; Performing image feature extraction on each reference image to obtain image features of each reference image; Perform splicing processing on each image feature to obtain a splicing feature; Based on the splicing features, the text features and the preset noise, a pre-trained video generation model is used to perform denoising processing for a preset number of time steps to generate a target video for the protagonist object, including: the video generation model includes a cross-attention module, and the cross-attention module includes a query feature mapping layer, a text key feature mapping layer, a text value feature mapping layer, an image key feature mapping layer and an image value feature mapping layer; in the denoising processing of each time step: based on the preset noise, feature mapping is performed using the query feature mapping layer to obtain query features; based on the text features, feature mapping is performed using the text key feature mapping layer and the text value feature mapping layer respectively to obtain to the first key feature and the first value feature; based on the splicing feature, use the image key feature mapping layer and the image value feature mapping layer to perform feature mapping respectively to obtain the second key feature and the second value feature; perform cross-attention calculation on the query feature, the first key feature, and the first value feature to obtain the first sub-fusion feature; perform cross-attention calculation on the query feature, the second key feature, and the second value feature to obtain the second sub-fusion feature; obtain the cross-attention output based on the first sub-fusion feature and the second sub-fusion feature; determine the target video based on the denoising processing of each time step and the cross-attention output of each time step.
2. The method according to claim 1, characterized in that The extracting image features from each reference image includes: For each reference image, extract image features of the reference image using a pre-trained multimodal model to obtain image features of the reference image, wherein the image features of the reference image include a classification label of the reference image; The stitching process of each image feature includes: Each classification mark is spliced to obtain the splicing feature.
3. The method according to claim 1 or 2, characterized in that In the denoising process of each time step: the cross-attention module performs cross-attention calculation on the splicing features, the text features, and the preset noise to obtain a cross-attention output, so as to determine the target video based on the denoising process of each time step and the cross-attention output of each time step.
4. The method according to claim 3, characterized in that The cross attention module performs cross attention calculation on the splicing feature, the text feature, and the preset noise, including: A cross-attention mechanism is performed on the query feature, the first key feature, the first value feature, the second key feature, and the second value feature to obtain the cross-attention output.
5. The method according to claim 1, wherein The obtaining the cross attention output based on the first sub-fusion feature and the second sub-fusion feature includes: Performing linear calculation on the first sub-fusion feature and the second sub-fusion feature to obtain the cross attention output.
6. The method according to claim 1, characterized in that The video generation model is obtained in the following way: Obtaining a model to be trained and sample data, wherein the sample data includes multiple sample videos, and label text, multiple frame label images, and label noise corresponding to any sample video, the multiple frame label images corresponding to any sample video all include the protagonist object in the any sample video, the label noise is generated by adding noise to the any sample video, and the model to be trained includes a cross-attention module, and the cross-attention module includes a query feature mapping layer, a text key feature mapping layer, a text value feature mapping layer, an image key feature mapping layer, and an image value feature mapping layer; For the multiple sample videos, feature extraction is performed on the multiple label images and label texts corresponding to the sample videos to obtain multiple label image features and label text features corresponding to the sample videos; Splicing the multiple label image features to obtain label splicing features corresponding to the sample video; The sample video, as well as the label noise, label splicing features, and label text features corresponding to the sample video, are input into the to-be-trained model, and the to-be-trained model performs noise addition processing for a preset number of time steps and outputs predicted noise, wherein, in the noise addition processing at each time step: the cross-attention module performs cross-attention calculation on the label splicing features, the label text features, and the sample video to obtain a label cross-attention output, so as to determine the predicted noise based on the noise addition processing at each time step and the label cross-attention output at each time step; The parameters of the model to be trained are adjusted according to each predicted noise and the corresponding label noise until a preset training end condition is met, and the video generation model is obtained from the model to be trained.
7. A video generation device based on a reference image, characterized in that: include: A first data acquisition module is configured to acquire text information and multiple reference images corresponding to a target video, wherein each of the multiple reference images includes a main character object of the target video; An encoding module, configured to encode the text information to obtain text features; A first feature extraction module is used to extract image features from each reference image to obtain image features of each reference image; A first feature stitching module is used to stitch the features of each image to obtain a stitching feature; The video generation module is used to perform denoising processing for a preset number of time steps based on the splicing features, the text features and the preset noise using a pre-trained video generation model to generate a target video for the protagonist object, including: the video generation model includes a cross-attention module, the cross-attention module includes a query feature mapping layer, a text key feature mapping layer, a text value feature mapping layer, an image key feature mapping layer and an image value feature mapping layer; in the denoising processing of each time step: based on the preset noise, feature mapping is performed using the query feature mapping layer to obtain query features; based on the text features, feature mapping is performed using the text key feature mapping layer and the text value feature mapping layer respectively. Mapping to obtain a first key feature and a first value feature; based on the splicing feature, use the image key feature mapping layer and the image value feature mapping layer to perform feature mapping respectively to obtain a second key feature and a second value feature; perform cross-attention calculation on the query feature, the first key feature, and the first value feature to obtain a first sub-fusion feature; perform cross-attention calculation on the query feature, the second key feature, and the second value feature to obtain a second sub-fusion feature; obtain the cross-attention output based on the first sub-fusion feature and the second sub-fusion feature; determine the target video based on the denoising processing of each time step and the cross-attention output of each time step.
8. An electronic device, characterized in that: include: memory for storing computer programs; A processor is configured to execute a computer program stored in the memory, and when the computer program is executed, implements the method described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising computer program instructions, characterized in that When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Image generation method and device and electronic equipment
CN116363262A
Role video generation method and device, electronic equipment and storage medium
CN118015159A
Data processing method and device
CN119094814A