Method, apparatus, device, storage medium and program product for video generation
Patent Information
- Application Number
- CN202510346361.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2026-09-22
AI Technical Summary
[0008] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description.
Smart Images

Figure CN122802747A_ABST
Abstract
Description
Technical Field
[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatus, devices, computer-readable storage media, and computer program products for video generation. Background Technology
[0002] In the field of computer vision (CV), various machine learning-based visual content processing techniques have seen significant development and widespread application. Machine learning-based visual content processing techniques can be used in scenarios that enhance user experience. In some example applications, the goal is to generate visual content (such as images or videos) that matches the user's input, such as textual descriptions. Summary of the Invention
[0003] In a first aspect of this disclosure, a video generation method is provided. The method includes: acquiring a video generation instruction, the video generation instruction including an input video, at least one reference image, and first descriptive information of the video to be generated, wherein the at least one reference image contains at least one subject; determining a visual feature representation set based on a video feature representation corresponding to the input video and an image feature representation corresponding to the at least one reference image; determining a text feature representation set based on a semantic feature representation corresponding to the at least one reference image and a text feature representation corresponding to the first descriptive information; and generating a target video containing at least one subject using an attention mechanism based on the visual feature representation set and the text feature representation set.
[0004] In a second aspect of this disclosure, an apparatus for video generation is provided. The apparatus includes: an acquisition module configured to acquire a video generation instruction, the video generation instruction including an input video, at least one reference image, and first descriptive information of a video to be generated, the at least one reference image containing at least one subject; a first determination module configured to determine a set of visual feature representations based on a video feature representation corresponding to the input video and an image feature representation corresponding to the at least one reference image; a second determination module configured to determine a set of text feature representations based on a semantic feature representation corresponding to the at least one reference image and a text feature representation corresponding to the first descriptive information; and a generation module configured to generate a target video containing at least one subject using an attention mechanism based on the set of visual feature representations and the set of text feature representations.
[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of this disclosure, a computer program product is provided, including computer-executable instructions, wherein when executed by a processor, the computer-executable instructions implement the method according to a first aspect of this disclosure.
[0008] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0010] Figure 1 A schematic diagram is shown of an example environment in which embodiments of the present disclosure may be implemented;
[0011] Figure 2 A schematic diagram of an example architecture for video generation according to some embodiments of the present disclosure is shown;
[0012] Figure 3 A schematic diagram of an example structure of a machine learning model for video generation according to some embodiments of the present disclosure is shown;
[0013] Figure 4 A schematic diagram of an example architecture of a diffusion model during the training phase according to some embodiments of the present disclosure is shown;
[0014] Figure 5 A block diagram illustrating an example process for video generation according to some embodiments of the present disclosure is shown;
[0015] Figure 6 A schematic structural block diagram of an example apparatus for video generation according to some embodiments of the present disclosure is shown; and
[0016] Figure 7 A block diagram of an electronic device in which one or more embodiments of the present disclosure can be implemented is shown. Detailed Implementation
[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0018] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.
[0019] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.
[0020] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0021] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.
[0022] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information, thereby enabling the user to choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers or storage media that perform the operation of the technical solution disclosed herein, based on the prompt message.
[0023] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, such as a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0024] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0025] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.
[0026] A neural network is a machine learning network based on deep learning. A neural network processes input and provides a corresponding output, typically consisting of an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications often include many hidden layers, thus increasing the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer is provided as the input to the next layer. The input layer receives the input to the neural network, while the output layer's output serves as the final output. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each node processing the input from the layer above.
[0027] Machine learning typically comprises three phases: training, testing, and application (also known as inference). In the training phase, a given model is trained using a large amount of training data, iteratively updating parameter values until the model can consistently generate inferences that meet the expected goals from the training data. Through training, the model can be considered to have learned the relationship between inputs and outputs (also known as the input-output mapping) from the training data. The parameter values of the trained model are determined. In the testing phase, test inputs are applied to the trained model to test whether it can provide the correct output, thus determining the model's performance. In the application phase, the model can be used to process actual model inputs based on the trained parameter values to determine the corresponding model output.
[0028] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. Figure 1As shown, environment 100 may include electronic device 110. In this example environment 100, electronic device 110 may acquire video generation instruction 112 to indicate the requirements for video generation. Video generation instruction 112 may include, for example, prompts, reference visual content such as videos, images, etc. Electronic device 110 may generate a target video 116 corresponding to video generation instruction 112 using machine learning model 105 based on video generation instruction 112. In some embodiments, target video 116 may include a subject. The subject may include, for example, a person, an animal, an object, etc.
[0029] The machine learning model 105 can be a single model or a combination of multiple models. In some embodiments, the machine learning model 105 can apply an attention mechanism to generate the target video 116. In some embodiments, the machine learning model 105 can be constructed, for example, from one or more diffusion models.
[0030] Diffusion models, also known as diffusion probability models, are a type of generative model. These models generate data by simulating a diffusion process. This process is inspired by physical processes such as thermal diffusion. Diffusion models include forward diffusion processes and reverse diffusion processes. A diffusion model generates new data samples by simulating a forward diffusion process with progressively added noise and then learning how to reverse this process.
[0031] In the forward diffusion process, noise is gradually added to the data, making it increasingly random through a series of steps until it resembles pure noise. This process can be viewed as a Markov chain, where Gaussian noise is added to the data at each step. The forward diffusion process can be represented as: Where x t This is the noise data at step t, α t This is used to control the amount of noise added. The forward diffusion process is performed during model training, and the data used to add noise are the training samples.
[0032] In the reverse diffusion process (or reverse denoising), the model learns how to reverse the steps of adding noise. Starting with pure noise, the diffusion model progressively removes the noise, generating data that matches the training distribution. The reverse diffusion process is typically simulated using a neural network that predicts the amount of noise added at each step. Where u θ and σ θThese are the learned model parameters. After model training is complete, the model performing the backdiffusion process can first sample from the noise distribution and then iteratively denoise it until the desired data is obtained.
[0033] In diffusion models, a time step refers to the number of steps in the forward diffusion process where noise is added. The total number of steps, T, is usually a preset value representing how many steps are required to transform the original data into pure noise. At each time step t, Gaussian noise is added to the data according to a predetermined noise scheme. This process is continuous, and each step depends on the result of the previous step.
[0034] In data generation, the inference step in a diffusion model refers to the number of steps required to recover the original data from pure noise during the backdiffusion process. The number of inference steps directly affects the quality and speed of the generated data. Generally, more inference steps result in higher quality data, but also increase computational cost and time. In practical applications, the number of inference steps can be adjusted to balance generation quality and efficiency. In some embodiments, inference steps correspond to time steps, and each inference step can correspond to one or more time steps. For example, if the total number of time steps in the diffusion model is 1000, and the inference steps are set to 50, then each inference step can correspond to 20 time steps.
[0035] In environment 100, electronic device 110 can be any type of computing device, including terminal devices or server devices. Terminal devices can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. Server devices may include, for example, computing systems / servers, such as mainframes, edge computing nodes, computing devices in cloud environments, and so on.
[0036] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0037] Currently, basic video generation models focus on two main tasks: text-to-video and image-to-video. Text-to-video (T2V) utilizes a language model to understand the input text instructions and generate visual content describing the expected characters, actions, and background. While it allows for creative and imaginative content combinations, it often struggles to produce consistent and predictable results due to its inherent randomness. On the other hand, image-to-video (I2V) typically provides the first frame of an image along with an optional text description to convert a static image into a dynamic video. Although it is more controllable, the richness of content is often limited by the strict "copy-and-paste" nature of the first frame. There is also a need to generate subject-consistent videos, which can be called subject-to-video (S2V). It involves capturing the subject from an image and flexibly generating video based on text prompts, while combining the diversity and controllability of joint image and text inputs. The essence of S2V lies in balancing bimodal prompts of text and images, requiring the model to align text instructions and image content simultaneously.
[0038] However, research on subject consistency in video generation tasks still lags behind that in image generation scenarios. With the maturity of text-to-image (T2I) base models, subject-to-image (S2I) has achieved good results, evolving from parameter optimization methods to adapter-based training methods and then to unified image editing methods. The most direct way to implement S2V is to combine S2I with I2V, but this has two main limitations. First, compared to S2V, S2I is more difficult to learn subject consistency because S2V training data naturally contains multi-view dynamic changes, allowing for a better understanding of the subject. Second, the transition from S2I to I2V can lead to information loss. For example, when generating back-to-foreground view motion, the subject's ID information may be lost due to the lack of ID information in the first frame, hindering I2V from maintaining ID consistency. Therefore, subject consistency generation requires a dedicated video model for unified processing.
[0039] The rise of diffusion models is rapidly reshaping the field of generative modeling. The advancements in video generation brought about by diffusion models are particularly significant. In the visual domain, video generation, compared to image generation, requires greater attention to the continuity and consistency of multiple frames, which presents additional challenges.
[0040] Significant progress has been made in subject-consistent generation for image tasks in recent years. Optimized training methods bind image content to specific identifiers during text-to-image generation. A notable work in the training and inference paradigm is IP-Adapter (a text-compatible image cue adapter capable of generating images based on image cues), which freezes the weights of the base model while training only the additional adapter to achieve subject-consistent generation. This approach is also widely used for tasks requiring facial ID consistency. However, these solutions are dependent on extracting image semantics, leading to a difficult trade-off between detailed reconstruction and flexible textual responses.
[0041] Currently, advancements in video generation capabilities and algorithmic innovation often lag behind those in image tasks. Similar to image consistency techniques, there is an optimization-based method for generating videos with consistent face IDs, which requires uploading multiple videos of the same person for optimization, resulting in significant computational costs. Adapter-based methods have also been explored for video ID consistency tasks. However, these works have been validated on small datasets, limiting their ability to perfectly align facial information with textual descriptions.
[0042] Therefore, a solution is needed to generate high-fidelity, subject-consistent videos. Furthermore, such a solution needs to balance the detail of video generation with flexible text responses, and also reduce computational costs.
[0043] In view of this, embodiments of the present disclosure propose a video generation scheme. In this scheme, a video generation instruction is obtained, which includes an input video, at least one reference image, and first descriptive information of the video to be generated, wherein the at least one reference image contains at least one subject. A visual feature representation set is determined based on the video feature representation corresponding to the input video and the image feature representation corresponding to the at least one reference image. A text feature representation set is determined based on the semantic feature representation corresponding to the at least one reference image and the text feature representation corresponding to the first descriptive information. Then, based on the visual feature representation set and the text feature representation set, an attention mechanism is used to generate a target video containing at least one subject.
[0044] In this manner, embodiments of this disclosure achieve unified single-subject or multi-subject generation by redesigning the text-image joint injection mechanism and utilizing dynamic feature integration. This solution enables high-fidelity subject-consistent video generation. Furthermore, it strikes a balance between the detail of video generation and flexible text responses, while reducing computational costs.
[0045] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.
[0046] Figure 2A schematic diagram of an example architecture 200 for video generation according to some embodiments of the present disclosure is shown. For ease of discussion, reference will be made to... Figure 1 These embodiments are described in the context of environment 100. These embodiments can be implemented in... Figure 1 In electronic device 110.
[0047] like Figure 2 As shown, the electronic device 110 acquires a video generation instruction 210. The video generation instruction 210 includes an input video 211, at least one reference image 212, and first descriptive information 213 for the video to be generated. The at least one reference image 212 contains at least one subject.
[0048] As an example, reference image 212 can be an image referenced when the video is generated. The subject can be the main object represented and focused on in reference image 212. As an example, the type of subject can include, but is not limited to, people, cartoon characters, animals, objects, etc. It should be understood that a subject can be any suitable object that can be focused on in visual content such as images or videos, without limitation. It is understood that if at least one subject includes one subject, that subject can be included in one reference image 212. If at least one subject includes multiple subjects, those multiple subjects can be included in one reference image 212 or multiple reference images 212. For example, a video generation instruction may include one reference image 212 that may contain two girls as subjects. Another video generation instruction may include two reference images 212, each containing one girl.
[0049] As an example, the first descriptive information 213 can be used to describe the content of the video to be generated. The described content may include, for example, the scene of the video to be generated, the time when the scene occurs, a description of one or more subjects in the video to be generated, etc. The first descriptive information 213 may sometimes also be referred to as a cue word.
[0050] As an example, the input video 211 can have different roles in the training and inference phases of the machine learning model 105, respectively. As an example, the machine learning model 105 can be configured to generate visual content such as images and videos based on text and images. It should be understood that the machine learning model 105 can include any model with text-to-image and image-to-image capabilities. In some embodiments, the machine learning model 105 can include a diffusion model. It should be understood that the generation of visual content based on text and images can be achieved through a single diffusion model or a combination of multiple diffusion models.
[0051] During the model training phase, the input video 211 can be used as the ground truth for constructing the training loss. For example, during the model training phase, after generating the predicted video based on the video generation instruction 210, the machine learning model 105 can define the corresponding training loss based on the difference between the predicted video and the input video 211, and then train the machine learning model 105 based on the training loss.
[0052] During the model inference phase (also known as the application phase), the input video 211 can be, for example, a noisy video. That is, during the model inference phase, the target video can be generated based on at least one reference image 212 and the first descriptive information 213 of the video to be generated.
[0053] Furthermore, a video feature representation 225 corresponding to the input video 211, an image feature representation 232 and a semantic feature representation 234 corresponding to at least one reference image 212, and a text feature representation 245 corresponding to the first descriptive information 213 can be determined. In the disclosed embodiments discussed herein, a feature representation can refer to a representation of the features of the original data that can be understood and processed by the machine learning model 105 after transformation of the original data and can express the features of the original data. In the embodiments of this disclosure, at least one reference image 212 is transformed into an image feature representation 232 and a semantic feature representation 234 from the image dimension and semantic dimension, respectively.
[0054] As an example, a variational autoencoder (VAE encoder) 220 for processing three-dimensional (3D) data can be used to determine the video feature representation 225 of the input video 211. After passing through the 3D VAE encoder 220, noise can be added to the input video 211 to obtain the video feature representation 225. During iterative training of the model, noise can be added to the input video 211 in each training iteration to progressively train the model. During the model's inference phase, the input video 211, after being noised, can become white noise, and then the model's inference is progressively executed.
[0055] As an example, a visual encoder 230 can be used to determine an image feature representation 232 and a semantic feature representation 234 corresponding to at least one reference image 212. As an example, the visual encoder 230 may include, for example, a VAE encoder and a graph-text multimodal encoder. The VAE encoder can encode the reference image 212 to provide detailed information about the image. The graph-text multimodal encoder can encode the reference image 212 to provide high-level semantic information about the image.
[0056] As an example, language model 240 can be used to determine the text feature representation 245 corresponding to the first descriptive information 213. Language model 240 can be configured to understand semantic information such as natural language and generate the corresponding text feature representation.
[0057] It should be understood that the encoders and language models described above for generating the corresponding feature representations are merely examples and are not intended to be limiting. In practical applications, any appropriate type of encoder or model can be used to generate feature representations of the corresponding data.
[0058] Further, based on the video feature representation 225 corresponding to the input video 211 and the image feature representation 232 corresponding to at least one reference image 212, the electronic device 110 determines a visual feature representation set 250. Based on the semantic feature representation 234 corresponding to at least one reference image 212 and the text feature representation 245 corresponding to the first descriptive information 213, the electronic device 110 determines a text feature representation set 260.
[0059] like Figure 2 As shown, video feature representation 225 and image feature representation 232 can be concatenated to generate a visual feature representation set 250. Semantic feature representation 234 and text feature representation 245 can be concatenated to generate a text feature representation set 260.
[0060] Taking the machine learning model 105, which includes a diffusion model, as an example, the diffusion model may include multiple modules 290. As an example, in each diffusion model module 290, a visual feature representation set 250 can be input into a visual branch 255, and a text feature representation set 260 can be input into a text branch 265.
[0061] Furthermore, based on the visual feature representation set 250 and the text feature representation set 260, the electronic device 110 utilizes an attention mechanism 270 to generate a target video containing at least one subject. The attention mechanism can be applied to a machine learning model 105, such as a diffusion model, to generate the target video.
[0062] The following will combine Figure 3 Let's discuss in detail how to use the attention mechanism 270 to generate target videos.
[0063] Figure 3 A schematic diagram of an example structure 300 for a machine learning model for video generation according to some embodiments of the present disclosure is shown. Structure 300 can be implemented in Figure 1 In the machine learning model 105 within environment 100. Taking machine learning model 105 as an example of a diffusion model, structure 300 can be shown as follows: Figure 2 The specific structure of the diffusion model module 290 in architecture 200.
[0064] In some embodiments, in order to generate a target video containing at least one subject, the electronic device 110 may divide multiple frames of the input video 211 into multiple groups of frames. As an example, each group of frames may correspond to a data window. For example, the multiple groups of frames may correspond to windows 1, 2, ..., N in box 310, respectively.
[0065] As an example, each group of frames in a set of frames can correspond to a set of frame feature representations. A set of frame feature representations can include the feature representations in video feature representation 225 corresponding to each frame in that set of frames. The feature representations in video feature representation 225 corresponding to each frame can be referred to as frame feature representations.
[0066] Furthermore, for each of the multiple sets of frames, based on a set of frame feature representations and image feature representations 232 corresponding to that set of frames in the visual feature representation set 250, the electronic device 110 can determine the corresponding combined visual representations in the query dimension, key dimension, and value dimension. Each feature representation in this embodiment can be converted into a feature vector representation in the query dimension, key dimension, and value dimension.
[0067] In some embodiments, to determine the corresponding combined visual representations across the query dimension, key dimension, and value dimension, for each of the query dimension, key dimension, and value dimension: for each frame in the set of frames, based on the frame feature representation of that frame, the electronic device 110 can determine the frame visual representation 322 of that frame in that dimension. As shown in box 320, the feature vector representation Q of each frame in the query dimension can be determined. vid The feature vector representation K in the key dimension vid The eigenvector representation V in the value dimension vid Q vid K vid and V vid They can be referred to individually or collectively as frame visual representations.
[0068] In box 310, each data window may include nine (or any other suitable number) frames of visual representation 322. It should be understood that the number of frames included in each group of frames can be set according to the actual application and is not limited here.
[0069] Furthermore, for each of at least one subject, based on the image feature representation 232 of the reference image 212 containing that subject, the electronic device 110 can determine the subject visual representation 332 in that dimension. As shown in box 330, the feature vector representation Q of each subject in the query dimension can be determined based on the image feature representation 232. ref_v The feature vector representation K in the key dimension ref_v The eigenvector representation V in the value dimensionref_v Q ref_v K ref_v and V ref_v It can be referred to as the main visual representation, either individually or collectively.
[0070] As an example, if at least one reference image 212 includes a subject, then the image feature representation 232 can correspond to a subject visual representation 332 in each dimension. If at least one reference image 212 includes multiple subjects, then the image feature representation 232 can correspond to multiple subject visual representations 332 in each dimension.
[0071] Then, in a predetermined order, the frame visual representations 322 determined for the group of frames and the subject visual representations 332 determined for at least one subject are combined to obtain a combined visual representation of the group of frames in that dimension. In some embodiments, the predetermined order may be a sequence in which the frame visual representations 322 come before the subject visual representations 332 in each data window. It should be understood that the predetermined order may also be any other suitable order in which the frame visual representations 322 and the subject visual representations 332 are arranged.
[0072] In some embodiments, the number of frames in each group of multiple frames may be related to the number of at least one subject. As an example, the sum of the number of frames in each group and the number of at least one subject may be equal to the length of the data window. Again, taking a data window length of 9 as an example, as shown in box 340, for a single subject, the combined visual representation may include 8 frame visual representations 322 and 1 subject visual representation 332. For 2 subjects (or other numbers greater than 1), the combined visual representation may include 7 frame visual representations 322 and 2 subject visual representations 332. In other embodiments, the length of the data window may not be fixed but may dynamically change under certain conditions; this disclosure does not limit this.
[0073] Further, based on the text feature representation 245 and semantic feature representation 234 in the text feature representation set 260, the electronic device 110 can determine the corresponding combined semantic representations on the query dimension, key dimension, and value dimension. In some embodiments, in order to determine the corresponding combined semantic representations on the query dimension, key dimension, and value dimension, for each of the query dimension, key dimension, and value dimension, based on the text feature representation 245, the electronic device 110 can determine the descriptive semantic representation 375 of the first descriptive information 213 on that dimension. For each of the at least one subject, based on the semantic feature representation 234 of the reference image 212 containing that subject, the electronic device 110 can determine the subject semantic representation 365 of that subject on that dimension.
[0074] As shown in box 370, the feature vector representation Q of the first description information 213 in the query dimension can be determined. text The feature vector representation K in the key dimension text The eigenvector representation V in the value dimension text Q text K text and V text This can be referred to individually or collectively as the descriptive semantic representation 375. As shown in box 360, the feature vector representation Q of each subject in the query dimension can be determined based on the semantic feature representation 234. ref_c The feature vector representation K in the key dimension ref_c The eigenvector representation V in the value dimension ref_c Q ref_c K ref_c and V ref_c It can be referred to as the subject semantic representation 365, either individually or collectively.
[0075] As an example, if at least one reference image 212 includes a subject, then the semantic feature representation 234 can correspond to one subject semantic representation 365 in each dimension. If at least one reference image 212 includes multiple subjects, then the semantic feature representation 234 can correspond to multiple subject semantic representations 365 in each dimension.
[0076] Then, in a predetermined order, the electronic device 110 can combine the descriptive semantic representation 375 and the subject semantic representation 365, each determined for at least one subject, to obtain a combined semantic representation in that dimension. In some embodiments, the predetermined order may be a sequence in which the descriptive semantic representation 375 comes first, followed by the subject semantic representation 365. It should be understood that the predetermined order may also be any other suitable order in which the descriptive semantic representation 375 and the subject semantic representation 365 are arranged.
[0077] Furthermore, the electronic device 110 can generate a target video by applying attention mechanism 270 to the corresponding combined semantic representation and the corresponding combined visual representation determined for each of the multiple frames. As an example, attention mechanism 270 can be applied based on the combination of the corresponding combined visual representation and the corresponding combined semantic representation after positional encoding (345) to generate the target video. As an example, positional encoding may include, for example, 3D rope positional encoding or any other suitable positional encoding. Attention mechanism 270 can also be referred to as a window self-attention mechanism.
[0078] Thus, for each of the query dimension, key dimension, and value dimension, by combining the combined visual representation of that dimension with the combined semantic representation of that dimension after position encoding corresponding to each data window (e.g., the nth window), the attention mechanism 270 can be used to output the combined feature representation 380.
[0079] Return to reference Figure 2 As an example, the combined feature representation 380 can be input into the 3D VAE decoder 280 to output the target video. As an example, the 3D VAE decoder 280 can be a counterpart to a 3D VAE encoder.
[0080] Combination Figure 2 and Figure 3 As discussed above, the visual encoder 230 includes a variational autoencoder (VAE) and a graph-text multimodal model. The image is encoded in F... ref_v As a feature, with video latent F vid Connect them together to obtain the spliced feature F V This allows for the reuse of 3D VAEs to maintain consistency in visual branch input. Simultaneously, image-text multimodal features F... ref_c With text features F text Connect them together to obtain the spliced feature F T This provides high-level semantic information, compensating for the low-level features of the VAE. It should be understood that feature merging involves dimension alignment, that is, merging the corresponding feature representations according to the query dimension, key dimension, and value dimension. Then, the concatenated feature F can be... T and F V The inputs are fed into the visual branch 255 and the text branch 265 of the diffusion model module 290, and the model separates the injected features only when calculating attention.
[0081] Specifically, the diffusion model module 290 is based on and improved for the reference image input, mainly modifying the attention module. First, according to F vid Calculated Q vid K vid V vid The features were divided into windows of size 9. Then, according to F... ref_v Calculated Q ref_v K ref_v V ref_v Dynamic connections are made to the end of each window, while in-situ features are moved sequentially to the beginning of the next window. This approach preserves the window structure while ensuring interaction between video and subject features within each window, as well as adaptive input for single or multiple subjects. Meanwhile, according to F... text Calculated Q text K textV text Features and according to F ref_c Calculated Q ref_c K ref_c V ref_c Features are dynamically concatenated. After collecting all reference information, each window is computed. Then, dynamically injected reference image features (including ref_v and ref_c) and text features from each window are extracted from the output features and averaged. This process ensures that the dimensions of input and output features remain consistent within the current module, thus facilitating computation in subsequent modules.
[0082] In some embodiments, the implementation using architecture 200 may be performed during the training of a diffusion model based on attention mechanism 270. The following will combine... Figure 4 Let's discuss the training process of the diffusion model in detail.
[0083] Figure 4 A schematic diagram of an example architecture 400 of a diffusion model during the training phase, according to some embodiments of the present disclosure, is shown. Architecture 400 can be implemented in... Figure 1 In environment 100, the diffusion model can be trained locally on electronic device 110 or on other devices / systems.
[0084] In some embodiments, the video generation instruction 210 may include training samples acquired in the following manner: The electronic device 110 may divide the initial video 411 into an input video 211 and a reference video 412 that is different from the input video 211. The initial video 411 may be derived from the video source 410.
[0085] In some embodiments, the first scene corresponding to the reference video 412 may be different from the second scene corresponding to the input video 211. For example, the first scene corresponding to the reference video 412 may be a scene of two people talking. The second scene corresponding to the input video 211 may be a scene of two people running. By extracting video clips from different scenes, the effectiveness of scene-related text can be improved and content diversity can be increased.
[0086] As an example, filtering (430) can be performed on the selected input video 211 and reference video 412. During the filtering process, video clips with low aesthetic appeal, low quality, low motion intensity, or video subtitles can be removed. Video clips with low motion intensity may include, for example, clips where the subject's motion amplitude is small.
[0087] Furthermore, the electronic device 110 can extract at least one reference image from the reference video 412 by performing subject detection 450 on the input video 211 and the reference video 412. Then, based on the at least one reference image, the input video 211, and descriptive information for the input video 211, the electronic device 110 can construct training samples.
[0088] In some embodiments, in order to extract at least one reference image from a reference video, the electronic device 110 may utilize a multimodal model to generate first descriptive information for the input video 211 and second descriptive information for the reference video 412. As an example, the multimodal model may be configured to understand textual and visual content and generate at least corresponding textual descriptive information. The multimodal model may, for example, include a visual-language model, or any other suitable type of model capable of understanding textual and visual content and generating at least textual descriptive information.
[0089] In some embodiments, the first descriptive information and the second descriptive information may respectively include a description of the subject and a description of the scene in the corresponding video. As an example, the description of the subject may include, for example, a description of the subject's appearance and posture (or behavior).
[0090] Furthermore, the electronic device 110 can detect (450) whether there are matching subject words in the first and second description information. As an example, a language model can be used to determine whether there are matching subject words by extracting description information about the subject from the first and second description information. For example, in box 445, description information 1 can be extracted from the first description information and description information 2 can be extracted from the second description information. It can be detected whether the two subject words "a person wearing a red hat" and "a person wearing a red dress" in description information 1 match the two subject words "a person wearing a red hat" and "a person wearing a red dress" in description information 2.
[0091] In some embodiments, if a matching subject word is detected, the electronic device 110 can extract at least one reference image from the reference video 412 based on the matching subject word. As an example, subjects in the input video 211 and the reference video 412 can be located by detection boxes in their respective video frames. This allows the use of a multimodal model to align the subjects in the detection boxes with the subject words in the descriptive information. Then, the subjects in the input video 211 and the reference video 412 can be matched (460) based on the detection boxes. Furthermore, at least one subject can be determined from the reference video 412, and at least one reference image containing at least one subject can be extracted. As an example, to highlight the subject in the reference image, a corresponding reference image can be generated by cropping each subject from the video frame.
[0092] In some embodiments, the implementation using architecture 200 can be executed during the inference phase. The first descriptive information may originate from user input. In some embodiments, using a language model, electronic device 110 can determine whether the first descriptive information meets predetermined requirements. The predetermined requirements may at least indicate a description of the appearance and posture of at least one subject. For example, for a subject such as a person, the first descriptive information may include, for example, the color of the person's clothing, the person's body posture, etc. As another example, for a subject such as an animal, the first descriptive information may include, for example, the color of the animal's fur, the animal's body posture, etc.
[0093] Furthermore, if the first description information does not meet the predetermined requirements, the electronic device 110 can update the first description information based on the predetermined requirements. That is, if the first description information input by the user lacks a description of the appearance and posture of at least one subject, the first description information can be rewritten to make it meet the predetermined requirements. In this way, by ensuring that the first description information can accurately describe the appearance and posture of each subject, confusion between similar subjects can be avoided.
[0094] Through embodiments of this disclosure, an adaptive injection mechanism is developed to dynamically determine the priority of text or image conditions according to task requirements, thereby reducing content leakage and improving text responsiveness, enhancing cross-modal alignment. Furthermore, it better simulates multi-agent interactions and implements consistent relative scaling. Adversarial debiasing techniques and synthetic data augmentation for underrepresented subjects are used to mitigate dataset and model bias. This topic-consistent video generation method achieves cross-modal alignment through text-image-video triple learning. By redesigning the text-image joint injection mechanism and utilizing dynamic feature integration, it demonstrates competitive performance in unified single / multi-agent generation and face ID preservation tasks.
[0095] Example process
[0096] Figure 5 A flowchart of an example process 500 for video generation according to some embodiments of the present disclosure is shown. Process 500 can be implemented at electronic device 110. Reference is made below. Figure 1 To describe process 500.
[0097] like Figure 5 As shown, in block 510, electronic device 110 acquires a video generation instruction, which includes an input video, at least one reference image, and first descriptive information of the video to be generated, wherein the at least one reference image contains at least one subject.
[0098] In box 520, electronic device 110 determines a set of visual feature representations based on video feature representations corresponding to the input video and image feature representations corresponding to at least one reference image.
[0099] In box 530, electronic device 110 determines a set of text feature representations based on semantic feature representations corresponding to at least one reference image and text feature representations corresponding to first descriptive information.
[0100] In box 540, electronic device 110 generates a target video containing at least one subject by using an attention mechanism based on a visual feature representation set and a text feature representation set.
[0101] In some embodiments, generating a target video containing at least one subject includes: dividing multiple frames of an input video into multiple groups of frames; for each group of frames, determining a corresponding combined visual representation in the query dimension, key dimension, and value dimension based on a set of frame feature representations and image feature representations corresponding to that group of frames in a visual feature representation set; determining a corresponding combined semantic representation in the query dimension, key dimension, and value dimension based on text feature representations and semantic feature representations in a text feature representation set; and generating the target video by applying an attention mechanism to the corresponding combined semantic representation and the corresponding combined visual representations determined for each group of frames.
[0102] In some embodiments, determining the corresponding combined visual representations on the query dimension, key dimension, and value dimension includes, for each of the query dimension, key dimension, and value dimension: for each frame in the set of frames, determining the frame visual representation of the frame on that dimension based on the frame feature representation of the frame; for each of at least one subject, determining the subject visual representation of the subject on that dimension based on the image feature representation of a reference image containing the subject; and combining the frame visual representations determined for the set of frames and the subject visual representations determined for at least one subject in a predetermined order to obtain the combined visual representation of the set of frames on that dimension.
[0103] In some embodiments, determining the corresponding combined semantic representations on the query dimension, key dimension, and value dimension includes, for each of the query dimension, key dimension, and value dimension: determining a descriptive semantic representation of first descriptive information on that dimension based on text feature representations; determining a subject semantic representation of that subject on that dimension based on semantic feature representations of a reference image containing that subject for each of at least one subject; and combining the descriptive semantic representations and the subject semantic representations determined for each of the at least one subject in a predetermined order to obtain the combined semantic representation on that dimension.
[0104] In some embodiments, the number of frames in each of the multiple frames is related to the number of at least one subject.
[0105] In some embodiments, the method is performed during the training of an attention-based diffusion model, and the video generation instruction includes training samples obtained by: dividing an initial video into an input video and a reference video different from the input video; extracting at least one reference image from the reference video by performing subject detection on the input video and the reference video; and constructing training samples based on at least one reference image, the input video, and descriptive information for the input video.
[0106] In some embodiments, the first scene corresponding to the reference video is different from the second scene corresponding to the input video.
[0107] In some embodiments, extracting at least one reference image from a reference video includes: using a multimodal model to generate first descriptive information for an input video and second descriptive information for a reference video, the first descriptive information and the second descriptive information respectively including a description of a subject and a description of a scene in the corresponding video; detecting whether there is a matching subject word in the first descriptive information and the second descriptive information; and in response to detecting a matching subject word, extracting at least one reference image from the reference video based on the matching subject word.
[0108] In some embodiments, the method is performed during the inference phase, the first descriptive information is derived from user input, and the method further includes: using a language model to determine whether the first descriptive information meets a predetermined requirement, the predetermined requirement indicating at least a description of the appearance and posture of at least one subject; and in response to the first descriptive information not meeting the predetermined requirement, updating the first descriptive information based on the predetermined requirement.
[0109] Example devices and equipment
[0110] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 6 A schematic structural block diagram of an example device 600 for video generation according to certain embodiments of the present disclosure is shown. Device 600 may be implemented as or included in electronic device 110. Various modules / components in device 600 may be implemented by hardware, software, firmware, or any combination thereof.
[0111] As shown in the figure, the device 600 includes: an acquisition module 610 configured to acquire a video generation instruction, the video generation instruction including an input video, at least one reference image, and first descriptive information of the video to be generated, the at least one reference image containing at least one subject; a first determination module 620 configured to determine a visual feature representation set based on a video feature representation corresponding to the input video and an image feature representation corresponding to the at least one reference image; a second determination module 630 configured to determine a text feature representation set based on a semantic feature representation corresponding to the at least one reference image and a text feature representation corresponding to the first descriptive information; and a generation module 640 configured to generate a target video containing at least one subject using an attention mechanism based on the visual feature representation set and the text feature representation set.
[0112] In some embodiments, the generation module 640 is configured to: divide multiple frames of the input video into multiple groups of frames; for each group of frames, determine a corresponding combined visual representation in the query dimension, key dimension, and value dimension based on a set of frame feature representations and image feature representations corresponding to that group of frames in the visual feature representation set; determine a corresponding combined semantic representation in the query dimension, key dimension, and value dimension based on text feature representations and semantic feature representations in the text feature representation set; and generate a target video by applying an attention mechanism to the corresponding combined semantic representations and the corresponding combined visual representations determined for each group of frames.
[0113] In some embodiments, the generation module 640 is configured to: for each frame in the set of frames, determine a frame visual representation of the frame in that dimension based on the frame feature representation of the frame; for each of at least one subject, determine a subject visual representation of the subject in that dimension based on the image feature representation of a reference image containing the subject; and combine the frame visual representations determined for the set of frames and the subject visual representations determined for at least one subject in a predetermined order to obtain a combined visual representation of the set of frames in that dimension.
[0114] In some embodiments, the generation module 640 is configured to: determine a descriptive semantic representation of the first descriptive information in the dimension based on the text feature representation; determine a subject semantic representation of the subject in the dimension based on the semantic feature representation of a reference image containing the subject for each of at least one subject; and combine the descriptive semantic representation and the subject semantic representation determined for each of the at least one subject in a predetermined order to obtain a combined semantic representation in the dimension.
[0115] In some embodiments, the number of frames in each of the multiple frames is related to the number of at least one subject.
[0116] In some embodiments, the apparatus 600 is an apparatus for training an attention-based diffusion model, and the video generation instruction includes training samples obtained by: dividing an initial video into an input video and a reference video different from the input video; extracting at least one reference image from the reference video by performing subject detection on the input video and the reference video; and constructing training samples based on at least one reference image, the input video, and descriptive information for the input video.
[0117] In some embodiments, the first scene corresponding to the reference video is different from the second scene corresponding to the input video.
[0118] In some embodiments, extracting at least one reference image from a reference video includes: using a multimodal model to generate first descriptive information for an input video and second descriptive information for a reference video, the first descriptive information and the second descriptive information respectively including a description of a subject and a description of a scene in the corresponding video; detecting whether there is a matching subject word in the first descriptive information and the second descriptive information; and in response to detecting a matching subject word, extracting at least one reference image from the reference video based on the matching subject word.
[0119] In some embodiments, the apparatus 600 is used in the reasoning phase, the first description information is derived from user input, and the apparatus 600 is further configured to use a language model to determine whether the first description information meets a predetermined requirement, the predetermined requirement indicating at least a description of the appearance and posture of at least one subject; and to update the first description information based on the predetermined requirement in response to the first description information not meeting the predetermined requirement.
[0120] The units and / or modules included in device 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units and / or modules in device 600 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0121] Figure 7 A block diagram of an electronic device 700 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 7 The electronic device 700 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 7 The illustrated electronic device 700 may include or be implemented as Figure 1 Electronic devices 110, or Figure 6 Device 600.
[0122] like Figure 7 As shown, electronic device 700 is in the form of a general-purpose electronic device. Components of electronic device 700 may include, but are not limited to, one or more processors or processor 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. Processor 710 may be a physical or virtual processor and is capable of performing various processes according to executable instructions stored in memory 720. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 700.
[0123] Electronic device 700 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 730 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 700.
[0124] Electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 7 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 720 may include computer-executable instruction product 725 having one or more executable instruction modules configured to perform various methods or actions of various embodiments of this disclosure.
[0125] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 700 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0126] Input device 750 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 760 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 700 can also communicate with one or more external devices (not shown) via communication unit 740 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 700, or with any device that enables electronic device 700 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0127] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer-executable instruction product is also provided, which is tangibly stored on a non-transient computer-readable medium and includes computer-executable instructions that are executed by a processor to implement the methods described above.
[0128] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer-executable instruction products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable and executable instructions.
[0129] These computer-executable instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-executable instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0130] Computer-executable instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0131] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-executable instruction products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, executable instruction, or portion of instructions, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0132] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A video generation method, comprising: Obtain a video generation instruction, the video generation instruction including an input video, at least one reference image and first descriptive information of the video to be generated, the at least one reference image containing at least one subject; A visual feature representation set is determined based on the video feature representation corresponding to the input video and the image feature representation corresponding to the at least one reference image; A set of text feature representations is determined based on the semantic feature representations corresponding to the at least one reference image and the text feature representations corresponding to the first descriptive information; as well as Based on the visual feature representation set and the text feature representation set, an attention mechanism is used to generate a target video containing at least one subject.
2. The method of claim 1, wherein generating the target video containing the at least one subject comprises: The input video is divided into multiple groups of frames; For each of the multiple sets of frames, based on a set of frame feature representations corresponding to that set of frames in the visual feature representation set and the image feature representation, a corresponding combined visual representation in the query dimension, key dimension, and value dimension is determined; Based on the text feature representations and semantic feature representations in the text feature representation set, determine the corresponding combined semantic representations on the query dimension, the key dimension, and the value dimension; and The target video is generated by applying the attention mechanism to the corresponding combined semantic representation and the corresponding combined visual representation determined for the multiple sets of frames.
3. The method of claim 2, wherein determining the corresponding combined visual representation on the query dimension, the key dimension, and the value dimension includes, for each of the query dimension, the key dimension, and the value dimension: For each frame in the set of frames, the visual representation of the frame in that dimension is determined based on the frame feature representation of that frame; For each of the at least one subject, based on the image feature representation of a reference image containing that subject, a subject visual representation in that dimension is determined; and According to a predetermined order, the frame visual representations determined for the group of frames and the subject visual representations determined for the at least one subject are combined to obtain the combined visual representation of the group of frames in that dimension.
4. The method of claim 2, wherein determining the corresponding combined semantic representation on the query dimension, the key dimension, and the value dimension includes, for each of the query dimension, the key dimension, and the value dimension: Based on the text feature representation, determine the descriptive semantic representation of the first descriptive information in this dimension; For each of the at least one subject, based on the semantic feature representation of a reference image containing that subject, determine the subject semantic representation in that dimension; and According to a predetermined order, the descriptive semantic representation and the subject semantic representation determined for each of the at least one subject are combined to obtain the combined semantic representation on this dimension.
5. The method of claim 2, wherein the number of frames in each of the plurality of frames is related to the number of the at least one subject.
6. The method of claim 1, wherein the method is performed during the training of a diffusion model based on the attention mechanism, and the video generation instruction includes training samples obtained in the following manner: The initial video is divided into the input video and a reference video that is different from the input video; By performing subject detection on the input video and the reference video, the at least one reference image is extracted from the reference video; and The training samples are constructed based on the at least one reference image, the input video, and descriptive information for the input video.
7. The method according to claim 6, wherein the first scene corresponding to the reference video is different from the second scene corresponding to the input video.
8. The method of claim 6, wherein extracting the at least one reference image from the reference video comprises: Using a multimodal model, first description information for the input video and second description information for the reference video are generated, wherein the first description information and the second description information respectively include a description of the subject and a description of the scene in the corresponding video; Detect whether there are matching subject words in the first description information and the second description information; as well as In response to the detection of a matching subject word, the at least one reference image is extracted from the reference video based on the matching subject word.
9. The method of claim 1, wherein the method is performed during the inference phase, the first descriptive information is derived from user input, and the method further comprises: Using a language model, it is determined whether the first descriptive information meets predetermined requirements, which at least indicate a description of the appearance and posture of the at least one subject; as well as In response to the first description information not meeting the predetermined requirements, the first description information is updated based on the predetermined requirements.
10. An apparatus for video generation, comprising: The acquisition module is configured to acquire a video generation instruction, the video generation instruction including an input video, at least one reference image and first descriptive information of the video to be generated, the at least one reference image containing at least one subject; The first determining module is configured to determine a visual feature representation set based on the video feature representation corresponding to the input video and the image feature representation corresponding to the at least one reference image; The second determining module is configured to determine a set of text feature representations based on the semantic feature representations corresponding to the at least one reference image and the text feature representations corresponding to the first descriptive information. as well as The generation module is configured to generate a target video containing at least one subject based on the visual feature representation set and the text feature representation set, using an attention mechanism.
11. An electronic device, comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 9 when executed by the at least one processor.
12. A computer-readable storage medium having stored thereon computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 9.
13. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 9.