Video generation method and device based on reference image, equipment, medium and product
By obtaining the text information of the target video and multi-frame reference images, extracting and splicing features, the video generation model is used to generate the video of the protagonist object, which solves the problem of low matching in the existing technology and improves the video generation quality and user experience.
Patent Information
- Application Number
- CN202510953667.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-07-10
AI Technical Summary
The existing video generation model cannot generate videos of specific protagonists, resulting in low matching between the generated video and user needs and affecting the user experience.
By obtaining the text information of the target video and multi-frame reference images, text features and image features are extracted, and after performing stitching processing, the pre-trained video generation model is used for denoising to generate the target video for the protagonist object.
Improve the quality and user experience of generated videos, ensure that the generated videos match user needs, and better learn the protagonist objects and text information.
Smart Images

Figure CN120455807A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of video generation technology, and in particular to a method, apparatus, device, medium, and product for video generation based on a reference image. Background Art
[0002] In recent years, the rapid development of video models has led to a surge in the field of AIGC (Artificial Intelligence Generated Content) video model technology. In related technologies, text-based video models typically generate videos based on the corresponding text descriptions. However, in practical applications, sometimes videos featuring specific protagonists need to be generated, but text-based video models are unable to generate videos based on text descriptions. This reduces the match between the generated videos and the user's desired videos, impacting the user experience. Summary of the Invention
[0003] In order to solve the above technical problems, the embodiments of the present disclosure provide a method, apparatus, device, medium and product for generating a video based on a reference image.
[0004] One aspect of an embodiment of the present disclosure provides a video generation method based on reference images, comprising: obtaining text information and multiple reference image frames corresponding to a target video, wherein the multiple reference image frames each include a protagonist object of the target video; encoding the text information to obtain text features; extracting image features from the reference images to obtain image features of the reference images; splicing the image features to obtain splicing features; and performing denoising for a preset number of time steps using a pre-trained video generation model based on the splicing features, the text features, and preset noise to generate a target video for the protagonist object.
[0005] Another aspect of the embodiments of the present disclosure provides a video generation device based on reference images, including: a first data acquisition module, used to acquire text information and multiple frames of reference images corresponding to a target video, wherein the multiple frames of reference images all include the protagonist object of the target video; an encoding module, used to encode the text information to obtain text features; a first feature extraction module, used to extract image features of the reference images to obtain image features of the reference images; a first feature splicing module, used to splice the image features to obtain splicing features; a video generation module, used to perform denoising for a preset number of time steps based on the splicing features, the text features and preset noise, using a pre-trained video generation model to generate a target video for the protagonist object.
[0006] Another aspect of the embodiments of the present disclosure provides an electronic device, including: a memory for storing a computer program; a processor for executing the computer program stored in the memory, and implementing the above method when the computer program is executed.
[0007] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented.
[0008] According to another aspect of the embodiments of the present disclosure, a computer program product is provided, including computer program instructions, which implement the above method when executed by a processor.
[0009] In an embodiment of the present disclosure, image features of multiple frames of reference images including the protagonist object in the target video are spliced together to obtain splicing features that include all aspects of information about the protagonist object. When generating the target video, the video generation model can pay attention to both text features and splicing features, and can better learn the protagonist object and text information, and generate a target video about the protagonist object, thereby improving the quality of the generated target video and enhancing the user experience.
[0010] The technical solution of the present disclosure is further described in detail below through the accompanying drawings and examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0012] The present disclosure can be more clearly understood from the following detailed description with reference to the accompanying drawings, in which: Figure 1 is a flowchart of a method for generating a video based on a reference image provided by an exemplary embodiment of the present disclosure; Figure 2 is a schematic diagram of a video generation method based on a reference image provided by an application example of the present disclosure; Figure 3 is a flowchart of a method for generating a video based on a reference image provided by another exemplary embodiment of the present disclosure; Figure 4 This is a structural diagram of an embodiment of a video generation device based on a reference image disclosed herein; Figure 5 The figure is a schematic structural diagram of an application embodiment of the electronic device disclosed herein. DETAILED DESCRIPTION
[0013] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure.
[0014] Those skilled in the art will understand that the terms "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, and do not represent any specific technical meanings, nor do they indicate a necessary logical order between them.
[0015] It should also be understood that in the embodiments of the present disclosure, “a plurality of” may refer to two or more than two, and “at least one” may refer to one, two, or more than two.
[0016] It should also be understood that any component, data or structure mentioned in the embodiments of the present disclosure can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.
[0017] In addition, the term "and / or" in this disclosure is merely a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this disclosure generally indicates that the related objects are in an "or" relationship.
[0018] It should also be understood that the description of the various embodiments in this disclosure focuses on the differences between the various embodiments, and the same or similar aspects thereof can be referenced with each other. For the sake of brevity, they will not be described one by one.
[0019] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0020] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.
[0021] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.
[0022] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0023] The embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate in conjunction with numerous other general-purpose or specialized computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with terminal devices, computer systems, servers, and other electronic devices include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, among others.
[0024] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system-executable instructions (such as program modules) executed by the computer system. Generally, program modules can include routines, programs, object programs, components, logic, data structures, and the like that perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in distributed cloud computing environments, where tasks are performed by remote processing devices linked via a communications network. In distributed cloud computing environments, program modules can be located on local or remote computer system storage media, including storage devices.
[0025] During the implementation of the present disclosure, the inventors discovered that in practical applications, some video generation scenarios require the generation of videos featuring specific characters. For example, a video creator might want to create a video featuring themselves, or an animation creator might draw a specific character and need to create a complete animated video featuring that character. However, existing video models are unable to generate videos featuring specific characters. This results in a poor match between the videos generated by the text-generated video model and the videos desired by users, impacting the user experience.
[0026] Figure 1 FIG is a flow chart of a method for generating a video based on a reference image provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to electronic devices, such as Figure 1 As shown, the video generation method based on the reference image may include the following steps: Step S100: Acquire text information and multiple reference image frames corresponding to the target video.
[0027] The multiple reference image frames all include the main character object in the target video.
[0028] In one embodiment, the protagonist object in the target video can be understood as the protagonist in the target video, etc. Exemplarily, the protagonist object may include at least one of the following: cartoon characters, animals, real people, mythical or legendary characters, buildings, plants, etc. Multiple frames of reference images may include information on at least one aspect of the protagonist object. For example, the protagonist object is Diao Chan, and multiple frames of reference images may include Diao Chan from different perspectives. For example, the first frame of reference image is a frontal image of Diao Chan, the second frame of reference image may be a left side image of Diao Chan, the third frame of reference image may be a right side image of Diao Chan, and the fourth frame of reference image may include a back image of Diao Chan, etc. The text information includes information describing the video content of the target video. In one implementation, the reference image may include multiple protagonist objects of the target video.
[0029] Users can manually draw multiple frames of reference images including the main object, or generate multiple frames of reference images including the main object through image generation software such as Midjourney, Adobe Firefly, Bing Image Creator, etc.
[0030] Step S110: Encode the text information to obtain text features.
[0031] Text features are vector representations extracted from text information that indicate the content of the target video. A text encoder can be used to encode the text information to obtain text features. For example, the T5 (Text to Text Transfer Transformer) model can be used as the text encoder.
[0032] Step S120 , performing image feature extraction on each reference image to obtain image features of each reference image.
[0033] The image features of each reference image are vector representations extracted from the reference image that characterize its visual content, spatial structure, and semantic information. Feature extraction can be performed on the reference image using a pre-trained deep learning model. Deep learning models can employ convolutional neural networks (CNNs) or recurrent neural networks (RNNs), for example.
[0034] Step S130 , performing splicing processing on each image feature to obtain a spliced feature.
[0035] The image features may be spliced together to obtain a spliced feature. For example, the image features may be spliced together in a channel dimension to obtain a spliced feature.
[0036] Step S140 , based on the splicing features, text features and preset noise, a pre-trained video generation model is used to perform denoising processing for a preset number of time steps to generate a target video for the main character object.
[0037] The target video includes a protagonist object. That is, the target video is a video about the protagonist object. The preset noise may include Gaussian noise. Exemplarily, a Gaussian noise matrix may be randomly generated in advance and determined as the preset noise.
[0038] In one embodiment, the video generation model may adopt a diffusion model. The splicing features, text features, and preset noise may be input into the video generation model, which then performs denoising for a preset number of time steps and outputs a target video.
[0039] In an embodiment of the present disclosure, image features of multiple frames of reference images including the protagonist object in the target video are spliced together to obtain splicing features that include all aspects of information about the protagonist object. When generating the target video, the video generation model can pay attention to both text features and splicing features, and can better learn the protagonist object and text information, and generate a target video about the protagonist object, thereby improving the quality of the generated target video and enhancing the user experience.
[0040] In some optional implementations, step S120 in the embodiment of the present disclosure may include: for each reference image, using a pre-trained multimodal model to extract image features of the reference image to obtain image features of the reference image.
[0041] The image feature of the reference image includes a classification label of the reference image.
[0042] Classification tokens (CLS tokens) are used in image processing to summarize global image features. Multimodal models can employ the Contrastive Language-Image Pretraining (CLIP) model. Each reference image is input into the CLIP model, which then outputs a CLS token for that reference image.
[0043] Accordingly, in the embodiment of the present disclosure, step S130 may include: performing splicing processing on each classification mark to obtain a splicing feature.
[0044] In the embodiment of the present disclosure, the classification labels of each reference image are spliced to obtain a spliced feature, thereby ensuring that the spliced feature can include the features of the protagonist object while reducing the size of the spliced feature, thereby improving the efficiency of processing the spliced feature and further improving the efficiency of generating the target video.
[0045] In some optional embodiments, in the disclosed embodiments, the video generation model includes a cross-attention module. The cross-attention module includes a query feature map layer, a text key feature map layer, a text value feature map layer, an image key feature map layer, and an image value feature map layer. The cross-attention module is configured to perform cross-attention calculations on the data. The query feature map layer, the text key feature map layer, the text value feature map layer, the image key feature map layer, and the image value feature map layer are configured to perform feature mapping processing.
[0046] In one embodiment, the video generation model is a Diffusion Transformer (DiT) model based on a Transformer structure. Compared to traditional deep neural network model structures, the video generation model in this embodiment has stronger cross-modal information alignment capabilities and complex semantic control capabilities in video generation tasks. Specifically, the video generation model includes a cross-attention module, which includes a query feature map layer, a text key feature map layer, a text value feature map layer, an image key feature map layer, and an image value feature map layer. During the denoising process at each time step, the cross-attention module performs cross-attention calculations on the splicing features, text features, and preset noise to obtain a cross-attention output. The target video is determined based on the denoising process and the cross-attention output at each time step. The cross-attention module is used to perform cross-attention calculations on the data. The query feature map layer, the text key feature map layer, the text value feature map layer, the image key feature map layer, and the image value feature map layer are used to perform feature mapping processing. Through the video generation model with the above-described model structure, cross-model fusion between splicing features and text features can be achieved.
[0047] Specifically, preset noise is input into a pre-trained video generation model, with the splicing features and text features serving as the pre-trained video generation model's conditions. De-noising is performed over a preset number of time steps to produce a target video that conforms to the semantics of the spliced image and text information. During the denoising process at each time step, cross-attention calculations are performed using the cross-attention module in the video generation model. The output of the video generation model at each time step T serves as the input for the next time step T+1. Based on the processing at the last time step, the video generation model outputs the generated target video.
[0048] Cross-Attention is a deep learning technique that allows dynamic associations between different input feature sequences. It enables cross-modal or cross-sequence information fusion by interacting query features with key-value features. In this embodiment, query features are derived based on the query feature mapping layer and preset noise in the video generation model. Text features and image features are mapped to text key-value features and image key-value features, respectively, based on the text key feature mapping layer, text value feature mapping layer, image key feature mapping layer, and image value feature mapping layer, thereby establishing cross-modal semantic dependencies.
[0049] In one embodiment, the structure of the video generation model may adopt the structure of a diffusion model. The video generation model may include multiple network layers, each of which adopts a Transformer network, and each network layer is provided with a cross attention module.
[0050] In one embodiment, in the denoising process of each time step, the cross-attention module in the Lth network layer performs cross-attention calculations on the splicing features, text features, and preset noise to obtain a cross-attention output, which is then used as the input of the L+1th network layer. The cross-attention module in the L+1th network layer then repeats the operation of the cross-attention module in the Lth network layer, and the cross-attention output output by the cross-attention module of the last network layer is determined as the output of this time step.
[0051] In the disclosed embodiment, by introducing splicing features during the target video generation process and configuring the cross-attention module in the video generation model to include a query feature mapping layer, a text key feature mapping layer, a text value feature mapping layer, an image key feature mapping layer, and an image value feature mapping layer, the video generation process can better focus on the protagonist object in each reference image, and better integrate the reference image and text information, thereby generating a target video that is consistent with the text information, has continuous video content, and is about the protagonist object. This improves the quality of the generated target video, ensures the match between the target video and user needs, and enhances the user experience.
[0052] In some optional embodiments, in the embodiments of the present disclosure, the cross-attention module performs cross-attention calculations on the splicing features, text features, and preset noise, including: performing feature mapping based on the preset noise using the query feature mapping layer to obtain the query feature (Quary), performing feature mapping based on the text feature using the text key feature mapping layer and the text value feature mapping layer to obtain the first key feature (Key1) and the first value feature (Value1), performing feature mapping based on the splicing feature using the image key feature mapping layer and the image value feature mapping layer to obtain the second key feature (Key2) and the second value feature (Value2), performing cross-attention mechanism calculations on the query feature, the first key feature, the first value feature, the second key feature, and the second value feature to obtain the cross-attention output.
[0053] In one embodiment, the query feature may represent the state of the current video generation process, the first key feature and the first value feature may represent key information of the text information, and the second key feature and the second value feature may represent key information of multiple reference frames.
[0054] Accordingly, in the embodiment of the present disclosure, performing cross-attention mechanism calculation on the query feature, the first key feature, the first value feature, the second key feature, and the second value feature may include: A cross-attention calculation is performed on the query feature, the first key feature, and the first value feature to obtain a first sub-fusion feature; a cross-attention calculation is performed on the query feature, the second key feature, and the second value feature to obtain a second sub-fusion feature; and a cross-attention output is obtained based on the first sub-fusion feature and the second sub-fusion feature.
[0055] In one embodiment, a matrix multiplication operation is performed on the query feature and the first key feature to obtain a first attention feature map (Attention Map), which is then multiplied by the first value feature to obtain a first sub-fusion feature. A matrix multiplication operation is performed on the query feature and the second key feature to obtain a second attention feature map (CilpAttention Map), which is then multiplied by the second attention feature map and the second value feature sequence to obtain a second sub-fusion feature. The cross-attention output can be obtained by multiplying the first sub-fusion feature and the second sub-fusion feature or by weighted multiplication.
[0056] Accordingly, in one embodiment, obtaining the cross-attention output based on the first sub-fusion feature and the second sub-fusion feature may include: performing linear calculation on the first sub-fusion feature and the second sub-fusion feature to obtain the cross-attention output.
[0057] Exemplarily, the first sub-fusion feature and the second sub-fusion feature can be added to obtain a cross-attention output.
[0058] In one embodiment, the preset noise is initialized and the initialized preset noise x is input into the to_q layer (query feature mapping layer), and the to_q layer outputs the query matrix Q (query feature); the splicing features are respectively input into the to_k layer (text key feature mapping layer) and the to_v layer (text value feature mapping layer), and the to_k layer outputs K (first key feature), and the to_v layer outputs V (first value feature); the image features are respectively input into the to_k_clip layer (image key feature mapping layer) and the to_v_clip layer (image value feature mapping layer), and the to_k_clip layer outputs K_clip (second key feature), and the to_v_clip layer outputs V_clip (second value feature). First, based on Q, K and V, use formula (1) to perform cross-attention calculation to obtain the first sub-fusion feature, that is, the Q and K matrices are multiplied to obtain the Attention Map (first attention feature map), and then multiplied by V; then, based on Q, K_clip and V_clip, use formula (2) to perform cross-attention calculation to obtain the second sub-fusion feature, that is, the Q and K_clip matrices are multiplied to obtain the Cilp Attention Map (second attention feature map), and then multiplied by V_clip. The first sub-fusion feature and the second sub-fusion feature are added to obtain the cross-attention output.
[0059] Formula (1) Formula (2) In formulas (1) and (2), Attention(Q,K,V) represents the first sub-fusion feature, represents the vector dimension (depth) of K, T represents the transpose calculation, represents the second sub-fusion feature, express The vector dimension (depth) of .
[0060] In the embodiment of the present disclosure, query features corresponding to preset noise, first key features and first value features corresponding to text features, and second key features and second value features corresponding to splicing features are generated respectively through the query feature mapping layer, text key feature mapping layer, text value feature mapping layer, image key feature mapping layer and image value feature mapping layer, and cross-attention calculation processing is performed on the query features, first key features, first value features, second key features and second value features, so that the video generation model can effectively focus on and align text information and reference images from different sources, thereby helping the video generation model to better capture the correlation between the text information and the protagonist object in the reference image, so that the generated target video can simultaneously meet the requirements of matching the text information and being a video about the protagonist object.
[0061] For example, Figure 2 FIG is a schematic diagram of a video generation method based on a reference image provided by an application example of the present disclosure. Figure 2 As shown, the video generation model includes n network layers, at least one of which includes a cross-attention module. In this example, the network layer uses a Transformer network. The cross-attention module includes: a query feature map layer (not shown in the figure), a text key feature map layer (not shown in the figure), a text value feature map layer (not shown in the figure), an image key feature map layer (not shown in the figure), and an image value feature map layer (not shown in the figure).
[0062] Obtain text information, multiple reference images, and preset noise. Each reference image includes the main subject of the target video. Encode the text information to obtain text features. Input each reference image into the CLIP model, which outputs a CLS token for each reference image. These CLS tokens are concatenated to obtain concatenated features. The concatenated features and the preset noise are then input into the encoder of a variational auto-encoder (VAE) for encoding. The concatenated features, text features, and preset noise are then input into a video generation model for denoising over m time steps (preset number of time steps) to determine the target video.
[0063] In the denoising process of the i-th (0<i<m, i is an integer) time step: in the cross attention module in the f-th (0<f<n, f is an integer) network layer, feature mapping is performed using the query feature mapping layer based on the preset noise to obtain a query feature, feature mapping is performed using the text key feature mapping layer and the text value feature mapping layer based on the text feature to obtain a first key feature and a first value feature, feature mapping is performed using the image key feature mapping layer and the image value feature mapping layer based on the splicing feature to obtain a second key feature and a second value feature, matrix multiplication operation is performed on the query feature and the first key feature to obtain a first attention feature map, and then the first attention feature map and the first value feature are multiplied to obtain a first sub-fusion feature; Perform a matrix multiplication operation on the query feature and the second key feature to obtain a second attention feature map, then perform a multiplication operation on the second attention feature map and the second value feature sequence to obtain a second sub-fusion feature, add the first sub-fusion feature and the second sub-fusion feature to obtain a cross-attention output, use the cross-attention output as the input of the cross-attention module in the f+1th network layer, repeat the above operation, determine the cross-attention output output of the cross-attention module of the nth network layer as the cross-attention output of the i-th time step, use the cross-attention output of the i-th time step as the preset noise in the denoising process of the i+1th time step, and determine the target video based on the cross-attention output of the denoising process of the m-th time step. Exemplarily, the cross-attention output of the denoising process of the m-th time step can be input into the decoder of the VAE for decoding to obtain the target video.
[0064] Figure 3 FIG. 1 is a flow chart of a method for generating a video based on a reference image provided by another exemplary embodiment of the present disclosure. Figure 3 As shown in Figure 2, the video generation model is obtained as follows: Step S200: Obtain the model to be trained and sample data.
[0065] Among them, the sample data includes: multiple sample videos, and the label text, multi-frame label images and label noise corresponding to any sample video. The multi-frame label images corresponding to any sample video include the protagonist object in any sample video. The label noise corresponding to any sample video is generated by adding noise to any sample video. The model to be trained includes: a cross-attention module, and the cross-attention module includes: a query feature mapping layer, a text key feature mapping layer, a text value feature mapping layer, an image key feature mapping layer and an image value feature mapping layer.
[0066] In one embodiment, multiple video clips can be obtained and cropped to a fixed resolution and duration to produce multiple sample videos. The sample videos can be input into a visual language model (VLM), which then outputs label text corresponding to the sample videos. Random Gaussian noise can be added to the sample videos to produce label noise corresponding to the sample videos. Multiple frames of video images containing the main subject can be selected from the sample videos as label images. Alternatively, multiple frames of label images containing the main subject can be generated using image generation software such as Midjourney, Adobe Firefly, or BingImageCreator.
[0067] Step S210 : For a plurality of sample videos, feature extraction is performed on a plurality of label images and label texts corresponding to the sample videos to obtain a plurality of label image features and label text features corresponding to the sample videos.
[0068] Each label image can be input into the CLIP model, and the CLIP model outputs the CLStoken of each label image as the image feature of each label image. A text encoder can be used to encode the label text to obtain the label text feature.
[0069] Step S220 , stitching multiple label image features to obtain label stitching features corresponding to the sample video.
[0070] Among them, the CLS tokens of each label image are spliced to obtain the label splicing features corresponding to the sample video.
[0071] In step S230, the sample video, as well as the label noise, label splicing features, and label text features corresponding to the sample video are input into the model to be trained. The model to be trained performs noise addition processing for a preset number of time steps and outputs predicted noise.
[0072] Among them, in the noise addition processing of each time step: the cross-attention module performs cross-attention calculation on the label splicing features, the label text features, and the sample video to obtain the label cross-attention output, so as to determine the predicted noise based on the noise addition processing of each time step and the label cross-attention output of each time step.
[0073] Specifically, based on the sample video, feature mapping is performed using the query feature mapping layer to obtain the label query feature, based on the label text feature, feature mapping is performed using the text key feature mapping layer and the text value feature mapping layer respectively to obtain the label first key feature and the label first value feature, based on the label splicing feature, feature mapping is performed using the image key feature mapping layer and the image value feature mapping layer respectively to obtain the label second key feature and the label second value feature, cross-attention calculation is performed on the label query feature, the label first key feature, and the label first value feature to obtain the label first sub-fusion feature, cross-attention calculation is performed on the label query feature, the label second key feature, and the label second value feature to obtain the label second sub-fusion feature, linear calculation is performed on the label first sub-fusion feature and the label second sub-fusion feature to obtain the label cross-attention output.
[0074] For example, Figure 2 The corresponding video generation model is used as an example to illustrate. In the noise addition processing at the i-th time step: in the cross attention module in the f-th network layer, based on the sample video, the query feature mapping layer is used to perform feature mapping to obtain the label query feature, based on the label text feature, the text key feature mapping layer and the text value feature mapping layer are used to perform feature mapping respectively to obtain the label first key feature and the label first value feature, based on the label splicing feature, the image key feature mapping layer and the image value feature mapping layer are used to perform feature mapping respectively to obtain the label second key feature and the label second value feature, the label query feature and the label first key feature are matrix multiplied to obtain the label first attention feature map, and then the label first attention feature map and the label first value feature are multiplied to obtain the label first sub-fusion feature; the label query feature and the label The second key feature of the label is matrix multiplied to obtain the second attention feature map of the label, and then the second attention feature map of the label is multiplied with the second value feature sequence of the label to obtain the second sub-fusion feature of the label, and the first sub-fusion feature of the label and the second sub-fusion feature of the label are added to obtain the label cross-attention output, and the label cross-attention output is used as the input of the cross-attention module in the f+1th network layer. The above operation is repeated, and the label cross-attention output output of the cross-attention module of the nth network layer is determined as the label cross-attention output of the i-th time step, and the label cross-attention output of the i-th time step is used as the sample video in the noise addition processing of the i+1th time step. The predicted noise is determined based on the label cross-attention output of the m-th time step. Exemplarily, the label cross-attention output of the m-th time step can be determined as the predicted noise.
[0075] In step S240 , the parameters of the model to be trained are adjusted according to the predicted noises and the corresponding label noises until a preset training end condition is met, and a video generation model is obtained from the model to be trained.
[0076] A loss function value can be determined based on the difference between the label noise and the predicted noise corresponding to each sample video using a preset loss function. Examples of the preset loss function include, but are not limited to, a cross-entropy error function or a mean square error function. The loss function value determination operation can be iteratively performed, and the parameters of the model to be trained can be iteratively adjusted to continuously reduce the loss function value until the loss function value converges. Preset training termination conditions are then determined to be met, and the training of the model to be trained is completed. The trained model to be trained is then used as the video generation model. Parameter optimizers such as stochastic gradient descent (SGD), adaptive gradient algorithm (Adagrad), adaptive moment estimation (Adam), and root mean square prop (RMSprop) can be used to adjust the parameters of the model to be trained. For example, the parameter optimizer can be used to calculate the gradient of each parameter of the model to be trained, and the parameters can be adjusted along the gradient direction, where the gradient represents the direction of the greatest loss function value reduction. The loss function value determination operation is iteratively performed until the loss function value no longer decreases, thus completing the training of the model to be trained and obtaining the video generation model.
[0077] In the embodiment of the present disclosure, the training model is trained using label splicing features including features of the protagonist object, label text features, and sample videos. This not only enables efficient training of the training model, but also enables the trained video generation model to generate videos about the target object.
[0078] Figure 4 FIG. 1 is a schematic diagram of a structure of an embodiment of a video generation device based on a reference image disclosed in the present invention. Figure 4 As shown, the device of this embodiment may include: A first data acquisition module 300 is configured to acquire text information and multiple reference images corresponding to a target video, wherein each of the multiple reference images includes a main character object of the target video; The encoding module 310 is used to encode the text information to obtain text features; A first feature extraction module 320 is configured to extract image features from each reference image to obtain image features of each reference image; A first feature stitching module 330 is used to stitch the image features to obtain stitching features; The video generation module 340 is used to perform denoising processing for a preset number of time steps using a pre-trained video generation model based on the splicing features, the text features and the preset noise, so as to generate a target video for the protagonist object.
[0079] Among some possible implementations of the present disclosure, in the embodiments of the present disclosure: A first feature extraction module 320 is specifically configured to extract image features from each reference image using a pre-trained multimodal model to obtain image features of the reference image, wherein the image features of the reference image include a classification label of the reference image; The first feature splicing module 330 is specifically configured to perform splicing processing on each classification mark to obtain the splicing feature.
[0080] In some possible implementations of the present disclosure, the video generation model in the embodiment of the present disclosure includes a cross-attention module, and the cross-attention module includes a query feature mapping layer, a text key feature mapping layer, a text value feature mapping layer, an image key feature mapping layer and an image value feature mapping layer; in the denoising process of each time step: the cross-attention module performs cross-attention calculation on the splicing features, the text features, and the preset noise to obtain a cross-attention output, so as to determine the target video based on the denoising process of each time step and the cross-attention output of each time step.
[0081] In some possible implementations of the present disclosure, the cross attention module in the embodiment of the present disclosure performs cross attention calculation on the splicing feature, the text feature, and the preset noise, and is further used to: Based on the preset noise, feature mapping is performed using the query feature mapping layer to obtain query features; Based on the text feature, feature mapping is performed using the text key feature mapping layer and the text value feature mapping layer to obtain a first key feature and a first value feature; Based on the splicing features, feature mapping is performed using the image key feature mapping layer and the image value feature mapping layer to obtain a second key feature and a second value feature; A cross-attention mechanism is performed on the query feature, the first key feature, the first value feature, the second key feature, and the second value feature to obtain the cross-attention output.
[0082] In some possible implementations of the present disclosure, the cross-attention mechanism calculation is performed on the query feature, the first key feature, the first value feature, the second key feature, and the second value feature in the embodiment of the present disclosure, and is further used to: Performing a cross attention calculation on the query feature, the first key feature, and the first value feature to obtain a first sub-fusion feature; Performing a cross attention calculation on the query feature, the second key feature, and the second value feature to obtain a second sub-fusion feature; The cross attention output is obtained based on the first sub-fusion feature and the second sub-fusion feature.
[0083] In some possible implementations of the present disclosure, the cross attention output is obtained based on the first sub-fusion feature and the second sub-fusion feature in the embodiment of the present disclosure, and is further used to: Performing linear calculation on the first sub-fusion feature and the second sub-fusion feature to obtain the cross attention output.
[0084] In some possible implementations of the present disclosure, the apparatus for generating a video based on a reference image in an embodiment of the present disclosure further includes: A second data acquisition module is used to acquire a model to be trained and sample data, wherein the sample data includes multiple sample videos, and label text, multiple frame label images and label noise corresponding to any sample video, the multiple frame label images corresponding to any sample video all include the protagonist object in any sample video, the label noise is generated by adding noise to any sample video, and the model to be trained includes a cross-attention module, and the cross-attention module includes a query feature mapping layer, a text key feature mapping layer, a text value feature mapping layer, an image key feature mapping layer and an image value feature mapping layer; A second feature extraction module is configured to perform feature extraction on the plurality of label images and label texts corresponding to the plurality of sample videos to obtain a plurality of label image features and label text features corresponding to the plurality of sample videos; A second feature splicing module is used to splice the multiple label image features to obtain a label splicing feature corresponding to the sample video; A first training module is configured to input the sample video, and the label noise, label splicing features, and label text features corresponding to the sample video, into the to-be-trained model, and the to-be-trained model performs a noise addition process for a preset number of time steps and outputs a predicted noise, wherein, in the noise addition process at each time step, the cross-attention module performs a cross-attention calculation on the label splicing features, the label text features, and the sample video to obtain a label cross-attention output, so as to determine the predicted noise based on the noise addition process at each time step and the label cross-attention output at each time step; The second training module is used to adjust the parameters of the model to be trained according to each predicted noise and the corresponding label noise until a preset training end condition is met, and the video generation model is obtained from the model to be trained.
[0085] The video generation device based on reference images in the embodiment of the present disclosure corresponds to the embodiment of the video generation method based on reference images in the present disclosure. The relevant contents can be referenced to each other and will not be repeated here.
[0086] The beneficial technical effects corresponding to the exemplary embodiment of the reference image-based video generation device of the embodiment of the present disclosure can be found in the corresponding beneficial technical effects of the above-mentioned corresponding exemplary method part, which will not be repeated here.
[0087] In addition, an embodiment of the present disclosure further provides an electronic device, including: memory for storing computer programs; The processor is configured to execute the computer program stored in the memory, and when the computer program is executed, the method for generating a video based on a reference image as described in any one of the above embodiments of the present disclosure is implemented.
[0088] Figure 5 This is a schematic diagram of the structure of an application embodiment of an electronic device disclosed herein. Below, referring to Figure 5, an electronic device according to an embodiment of the present disclosure is described. The electronic device can be either or both of the first and second devices, or a standalone device independent of them. The standalone device can communicate with the first and second devices to receive collected input signals from them.
[0089] like Figure 5 As shown, the electronic device includes one or more processors and memory.
[0090] The processor may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.
[0091] The memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor may execute the program instructions to implement the reference image-based video generation method of the various embodiments of the present disclosure described above and / or other desired functions.
[0092] In one example, the electronic device may further include an input device and an output device, and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0093] In addition, the input device may also include, for example, a keyboard, a mouse, and the like.
[0094] The output device can output various information to the outside, including determined distance information, direction information, etc. The output device can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, and the like.
[0095] Of course, to simplify, Figure 5 Only some of the components related to the present disclosure in the electronic device are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, the electronic device may further include any other appropriate components according to specific application scenarios.
[0096] In addition to the above-mentioned methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the reference image-based video generation method according to various embodiments of the present disclosure described in the above part of this specification.
[0097] The computer program product may be written in any combination of one or more programming languages to implement the operations of the disclosed embodiments, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0098] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enable the processor to execute the steps of the reference image-based video generation method according to various embodiments of the present disclosure described in the above part of this specification.
[0099] The computer-readable storage medium may be any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0100] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk, etc. Various media that can store program codes.
[0101] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.
[0102] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For system embodiments, since they are essentially identical to the method embodiments, their description is relatively simple. For relevant parts, refer to the descriptions of the method embodiments.
[0103] The block diagrams of the devices, devices, equipment, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0104] The methods and apparatus of the present disclosure may be implemented in many ways. For example, the methods and apparatus of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above unless otherwise specified. In addition, in some embodiments, the present disclosure may also be implemented as programs recorded in a recording medium, which include machine-readable instructions for implementing the methods according to the present disclosure. Thus, the present disclosure also covers recording media that store programs for executing the methods according to the present disclosure.
[0105] It should also be noted that in the apparatus, device, and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.
[0106] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0107] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A video generation method based on a reference image, characterized in that: include: Acquire text information and multiple reference images corresponding to a target video, wherein the multiple reference images each include a protagonist object of the target video; Encoding the text information to obtain text features; Performing image feature extraction on each reference image to obtain image features of each reference image; Perform splicing processing on each image feature to obtain a splicing feature; Based on the splicing features, the text features and the preset noise, a pre-trained video generation model is used to perform denoising processing for a preset number of time steps to generate a target video for the protagonist object.
2. The method according to claim 1, characterized in that The extracting image features from each reference image includes: For each reference image, extract image features of the reference image using a pre-trained multimodal model to obtain image features of the reference image, wherein the image features of the reference image include a classification label of the reference image; The stitching process of each image feature includes: Each classification mark is spliced to obtain the splicing feature.
3. The method according to claim 1 or 2, characterized in that The video generation model includes a cross-attention module, which includes a query feature map layer, a text key feature map layer, a text value feature map layer, an image key feature map layer, and an image value feature map layer; In the denoising process of each time step: the cross-attention module performs cross-attention calculation on the splicing features, the text features, and the preset noise to obtain a cross-attention output, so as to determine the target video based on the denoising process of each time step and the cross-attention output of each time step.
4. The method according to claim 3, characterized in that The cross attention module performs cross attention calculation on the splicing feature, the text feature, and the preset noise, including: Based on the preset noise, feature mapping is performed using the query feature mapping layer to obtain query features; Based on the text feature, feature mapping is performed using the text key feature mapping layer and the text value feature mapping layer to obtain a first key feature and a first value feature; Based on the splicing features, feature mapping is performed using the image key feature mapping layer and the image value feature mapping layer to obtain a second key feature and a second value feature; A cross-attention mechanism is performed on the query feature, the first key feature, the first value feature, the second key feature, and the second value feature to obtain the cross-attention output.
5. The method according to claim 4, characterized in that The performing a cross attention mechanism calculation on the query feature, the first key feature, the first value feature, the second key feature, and the second value feature includes: Performing a cross attention calculation on the query feature, the first key feature, and the first value feature to obtain a first sub-fusion feature; Performing a cross attention calculation on the query feature, the second key feature, and the second value feature to obtain a second sub-fusion feature; The cross attention output is obtained based on the first sub-fusion feature and the second sub-fusion feature.
6. The method according to claim 5, characterized in that The obtaining the cross attention output based on the first sub-fusion feature and the second sub-fusion feature includes: Performing linear calculation on the first sub-fusion feature and the second sub-fusion feature to obtain the cross attention output.
7. The method according to claim 1, characterized in that The video generation model is obtained in the following way: Obtaining a model to be trained and sample data, wherein the sample data includes multiple sample videos, and label text, multiple frame label images, and label noise corresponding to any sample video, the multiple frame label images corresponding to any sample video all include the protagonist object in the any sample video, the label noise is generated by adding noise to the any sample video, and the model to be trained includes a cross-attention module, and the cross-attention module includes a query feature mapping layer, a text key feature mapping layer, a text value feature mapping layer, an image key feature mapping layer, and an image value feature mapping layer; For the multiple sample videos, feature extraction is performed on the multiple label images and label texts corresponding to the sample videos to obtain multiple label image features and label text features corresponding to the sample videos; Splicing the multiple label image features to obtain label splicing features corresponding to the sample video; The sample video, as well as the label noise, label splicing features, and label text features corresponding to the sample video, are input into the to-be-trained model, and the to-be-trained model performs noise addition processing for a preset number of time steps and outputs predicted noise, wherein, in the noise addition processing at each time step: the cross-attention module performs cross-attention calculation on the label splicing features, the label text features, and the sample video to obtain a label cross-attention output, so as to determine the predicted noise based on the noise addition processing at each time step and the label cross-attention output at each time step; The parameters of the model to be trained are adjusted according to each predicted noise and the corresponding label noise until a preset training end condition is met, and the video generation model is obtained from the model to be trained.
8. A video generation device based on a reference image, characterized in that: include: A first data acquisition module is configured to acquire text information and multiple reference images corresponding to a target video, wherein each of the multiple reference images includes a main character object of the target video; An encoding module, configured to encode the text information to obtain text features; A first feature extraction module is used to extract image features from each reference image to obtain image features of each reference image; A first feature stitching module is used to stitch the features of each image to obtain a stitching feature; The video generation module is used to perform denoising processing for a preset number of time steps using a pre-trained video generation model based on the splicing features, the text features and the preset noise, so as to generate a target video for the protagonist object.
9. An electronic device, characterized in that: include: memory for storing computer programs; A processor is configured to execute a computer program stored in the memory, and when the computer program is executed, implements the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
11. A computer program product comprising computer program instructions, characterized in that When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Video generation method and server
CN116233491A
Image generation method and device and electronic equipment
CN116363262A
Training method and application method of multi-view image generation model
CN117372631A
Role video generation method and device, electronic equipment and storage medium
CN118015159A
Data processing method and device
CN119094814A