Graphics video generation method based on non-training strategy and multi-subject attention alignment

Through the graph-generated video generation method that is aligned with multi-subject attention without training strategy, the subject perception and attention decoding mechanism are used to solve the problem of subject recognition and distinction in multi-subject scenes, the independent motion and high-quality animation generation of each subject are realized, and the application scenarios of graph-generated video are expanded.

CN120050481AActive Publication Date: 2025-05-27ANHUI UNIVERSITY OF TECHNOLOGY

Patent Information

Application Number
CN202510195783.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-27
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

The existing video generation method of image generation is difficult to accurately identify and distinguish different subjects in multi-subject scenarios, resulting in confusion in content and movement, and lack of data sets generated for multi-subjects, limiting application scenarios.

Method used

A graph-generated video generation method that is aligned with multi-subject attention is adopted using a training-free strategy. Through subject perception and attention decoupling mechanisms, the input text and images are processed separately, the masks and features of each subject are extracted, and the cross attention is calculated in the U-Net network to achieve independent motion and high-quality animation generation of each subject.

Benefits of technology

It realizes accurate identification and distinction between different subjects in multi-subject scenarios, ensures that each subject performs different actions independently, enriches the application scenarios of image-generated video generation, and improves generation effect and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050481A_ABST
    Figure CN120050481A_ABST
Patent Text Reader

Abstract

The invention discloses a map video generation method based on a non-training strategy and multi-subject attention alignment, and belongs to the technical field of map videos. According to the method, different subjects are separated through subject perception and attention disaggregation processing, only information of one subject is aligned when attention is calculated, other subjects are shielded, and frame features and text features of each subject need to be aligned in attention calculation. The solved attention can extract each subject and align the subject with the corresponding text feature, and other objects do not participate in calculation during processing, so that each subject only expresses the action corresponding to the own text in the final result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image-to-video technology, and particularly to an image-to-video generation method based on a training-free strategy and multi-agent attention alignment. Background Art

[0002] Diffusion models have achieved remarkable success in the field of generation, especially in image generation, even surpassing generative adversarial networks (GANs). With the progress of text-to-image generation technology, video generation tasks have also developed rapidly. The Video Diffusion Model (VDM) first extended the 2D U-Net to a 3D U-Net structure to achieve joint training of images and videos. In addition, AnimateDiff trains a motion module to adapt it to different personalized text-to-image (T2I) models and combines other specialized content models to generate high-quality videos. Text2VideoZero proposed a sampling method without additional training, which can enhance motion dynamics while maintaining frame consistency, thereby generating expected video content.

[0003] Custom video generation aims to generate highly personalized video content according to specific user needs or input conditions, mainly focusing on the text-to-video field. This is achieved by fine-tuning the generation of video subjects using text with specific symbols and fine-tuning specific motions using text representing the motions in relevant videos. The resulting videos generated by fine-tuning the text have the desired subjects and motions, thus realizing custom video generation. Existing custom video generation adopts the fine-tuning method. For example, CustomVideo uses the DreamBooth fine-tuning method to modulate its model with specific text representing images. LAMP fine-tunes specific text actions and binds these actions to the actions in the training videos to achieve specific action fine-tuning. DreamVideo fine-tunes spatial subjects and temporal actions respectively to generate a video model with specific subjects and actions. Recently, there have been studies on multi-objective customization. DisenStudio inputs image data containing multiple subjects and performs stable diffusion, fine-tunes specific text prompts, and uses masks to distinguish different subjects to achieve text-driven motion of different subjects. Customized video generation has promoted the development of multi-objective control in terms of subject and motion control, enriching the field of video generation.

[0004] Recently, some progress has been made in the research on multi-agent generation. In the field of image generation, Be Yourself significantly alleviates the semantic confusion and misalignment problems of multiple agents in the text during StableDiffusion generation in a training-free manner. Mastering text-to-image diffusion uses large language models to automatically separate the agents in the text and form separate prompts. Each prompt enters the UNet network simultaneously, and the final latent results are stitched together according to specific dimensions to ensure that each agent does not interfere with each other, thus achieving multi-agent image generation. In the field of video, multi-agent generation is also developing rapidly, mainly focusing on personalized custom video generation. For example, both DisenStudio and CustomVideo train the model by stitching multiple agents into a single image. They use special placeholders to bind the agents and use masks to distinguish different agents. The difference is that the former uses the stitching method to combine the cross-attention of different agents based on different texts instead of the original attention and uses masks to distinguish one by one. The latter uses masks to distinguish throughout the cross-attention.

[0005] Recent research in the field of image-to-video has made progress in the controllable movement of multiple agents, such as Follow-Your-Pose v2, which animates multiple objects in an image based on the poses of multiple people. Image-to-video usually animates the input image according to given conditions to achieve dynamic content from the image. With the development of image-to-video, several branches have emerged. For example, text-guided image-to-video generation aims to generate videos that match both the text semantics and the image content. Human pose-guided models convert human images into videos with additional controls, such as dense poses, depth maps, etc. Image animation under optical flow conditions is similar to human pose control animation, mainly focusing on character animation. Models that control movement in the mask area achieve this by controlling the movement of selected areas in the image, while in the field of trajectory control, they control the movement direction. These models usually combine clean images with initial noise or inject image information into the initial noise, and then generate videos through the U-Net network, ensuring the consistency of spatial content while paying attention to the movement in the time dimension.

[0006] The above methods have achieved good performance, but there is no research on image-to-video that only uses text to drive, and some key problems need to be solved urgently.

[0007] (1) In the research on image-to-video generation that only relies on text, when the generated content involves multiple agents, the content and movement of these agents often become chaotic, making it difficult to accurately align the text with the corresponding agents or distinguish different agents.

[0008] (2) In real-world scenarios, the occurrence frequency of multi-subject images is much higher than that of single-subject images. However, existing methods generally ignore this point, which greatly limits the application of image-to-video generation in multi-subject scenarios. In addition, most existing video generation datasets are targeted at scene motion or single-subject actions, lacking pure multi-subject generation datasets, and the difficulty of dataset production and search is also relatively large. Summary of the Invention

[0009] The technical problem to be solved by the present invention is as follows: how to solve the deficiencies of existing image-to-video methods in multi-subject scenarios, and provide an image-to-video generation method based on a non-training strategy and multi-subject attention alignment. This method can not only accurately identify different subjects in the image and precisely correspond them to their respective text descriptions, but also enables each subject to independently perform different actions, enriching the application scenarios of image-to-video generation and expanding the technical boundaries of this field.

[0010] The present invention solves the above technical problems through the following technical solutions. The present invention includes the following steps:

[0011] S1: Obtain the input text and picture;

[0012] S2: Separate the input text and simultaneously segment the input picture to obtain the mask of each subject, and perform perceptual processing on the mask;

[0013] S3: Input the extracted mask, the separated text, and the corresponding input picture into a pre-trained image-to-video model based on a diffusion model;

[0014] S4: Encode the picture and then copy it multiple times, and add noise to make it noise;

[0015] S5: Send the separated text into the CLIP text encoder to extract the features of each subject respectively, and copy each subject feature to the same number of copies as the number of frames of the target generated video, and store these copied subject features in the list set p in sequence;

[0016] S6: Input the obtained set p, the masks of each subject region, and the noise into the U-Net network;

[0017] S7: When calculating cross-attention in the U-Net network, use the mask to obtain the feature q of each subject from multiple subject features q l , the subject feature q l and the matching text feature k l and v l First, calculate q l and k lThe similarity is used to obtain an attention weight, which represents the correlation between each input and other inputs. Then, the attention weights are normalized. Finally, these normalized attention weights are applied to the corresponding value vectors to generate an attention output, where k l and v l are obtained from the set p, and l represents the l-th subject;

[0018] S8: Concatenate multiple attention outputs and perform a normalization operation with the sum of all masks to obtain an attention DeAttention that matches the pre-trained image-to-video model based on the diffusion model, and replace the attention in the original attention module in the U-Net network. Then, obtain a denoised latent variable in one step;

[0019] S9: Generate the final latent variable in one-step denoising using the classifier-free guidance method and obtain the latent representation;

[0020] S10: Gradually eliminate noise by repeating denoising T times to obtain the target generated video.

[0021] Furthermore, in the step S2, the specific processing process is as follows:

[0022] S21: Through input text separation processing, separate the descriptive text containing multiple subjects, and store the description of each subject separately in a set;

[0023] S22: Using the separated text descriptions, segment each subject in the picture through visual segmentation processing to generate corresponding subject area masks;

[0024] S23: Then perform perception processing on each mask: in each mask, set the pixel values of the area corresponding to the subject to 1, and the pixel values of the areas of other subjects to 0, so that each mask accurately represents the area of its corresponding subject.

[0025] Furthermore, in the step S4, the specific processing process is as follows:

[0026] S41: After converting the input picture into a latent variable through a pre-trained encoder, add noise to it;

[0027] S42: Duplicate the latent variable multiple times so that its quantity is the same as the number of frames of the target generated video;

[0028] S43: Add noise with a frame number dimension to each copy of the latent variable;

[0029] S44: Concatenate the latent variables with noise along the frame number dimension to obtain a latent representation suitable for video generation.

[0030] Furthermore, in the step S7, the main feature q l is calculated as follows:

[0031] q l = q · mask;

[0032] The calculation formula of the attention output is as follows:

[0033]

[0034] where atten l is the attention output of the text feature of the l-th main body aligned with the main feature, and d is the dimension of q l and k l .

[0035] Furthermore, in the step S9, the acquisition formula of the latent representation is as follows:

[0036]

[0037] where ω is the weight, ε is the U-Net network, ε(z t , p) is the conditional latent variable, and ε(z t ) is the unconditional latent variable.

[0038] The present invention has the following advantages compared with the prior art:

[0039] (1), Distinguishing different main bodies

[0040] Subject perception and attention disentanglement separate different main bodies. When calculating attention, only the information of one main body itself is aligned, and the rest of the main bodies are masked. In the calculation of attention, it is necessary to align the frame features and text features of each main body. The disentangled attention we proposed can extract each main body and align it with the corresponding text feature, and the rest of the objects do not participate in the calculation during processing. In this way, each main body only shows the actions corresponding to its own text in the final result.

[0041] (2), Improvement of flexibility and efficiency

[0042] The present invention constructs the architecture based on training-free. Existing image-to-video generation can generate action animations with text semantics when processing a single subject. However, when dealing with multiple subjects, it cannot distinguish between subjects, easily leading to confusion in appearance and actions, resulting in the absence of subjects and chaotic actions. Training-free avoids the cumbersome training process based on these single subjects, thus significantly saving time. It is usually designed to be seamlessly integrated with existing models or systems to achieve plug-and-play functionality, which provides great convenience for users. In the face of the lack of multi-subject data, good results can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a schematic flowchart of the image-to-video generation method based on training-free strategy and multi-subject attention alignment in an embodiment of the present invention;

[0044] Figure 2 is a schematic structural diagram of the training-free sampling framework in an embodiment of the present invention;

[0045] Figure 3 are partial experimental results in an embodiment of the present invention, where (a) is the first experimental result and (b) is the second experimental result. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The embodiments of the present invention will be described in detail below. These embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.

[0047] Embodiment 1

[0048] The task objective of multi-subject image-to-video generation will be described below:

[0049] Image-to-video generation aims to realize a qualified dynamic animation of static pictures under the guidance of input conditions. There are many guiding conditions for image-to-video generation, such as trajectory, pose, optical flow, text, etc. However, most of these methods focus on a single subject, while multiple subjects are very common in reality. In the present invention, we propose a training-free decoupled diffusion model for multi-subject image-to-video generation. Specifically, we use subject perception to process the input text and images, match different texts with their corresponding subjects to obtain a series of individual texts and corresponding subjects. After the attention of each subject is untied, the matched text and subject are bound, the remaining subjects are masked, and finally the generated attention is fused to replace the original attention, so that each subject only focuses on its own area and ignores the areas of other subjects, effectively solving the problem of the appearance and movement intersection between subjects in the multi-subject scenario, ensuring the independent movement of each subject, and achieving high-quality animation effects.

[0050] We achieve the above task objectives through the following steps:

[0051] 1. Subject perception: To align different subjects with their respective text descriptions and avoid conflicts with other subjects in subsequent attention calculations, we use subject perception. After the input text and image, we use visual perception methods to segment the input image to obtain the masks of different subjects in the image. These masks are further processed to retain their respective parts while masking other subjects, which means that the areas of other subjects are set to 0 and the rest to 1. To more accurately extract the regions of each subject, we use text perception methods to divide the complete text prompt into multiple descriptions, each containing only one subject, and store them in a set. The prompts in the set correspond one-to-one with the subjects in the input image. Each prompt in the set is used as a segmentation guide to sequentially segment the subject in the image that matches the text prompt and accurately obtain the mask of each subject. Then, the list of masks obtained by inverting and masking other subjects is further adjusted. Finally, each mask shows only a single subject, masks other subjects, and aligns each subject with the corresponding text when calculating attention.

[0052] 2. Attention disentanglement: In existing models, when processing attention, the semantics of multiple entities may become entangled, resulting in the mixing of subject features. This confusion leads to inconsistent appearances of entities in the generated results and actions that do not match the semantic descriptions. To solve this problem, we disentangle different attentions and redesign the alignment between the image features of each subject and the corresponding text embeddings in the frame image when processing through the U-Net network. We use the Mask-base Query Generation (MQG) method to disentangle the features of different subjects and align these features with the corresponding text features. The disentangled attentions of each subject are recombined into a single attention consistent with the U-Net network parameters. We call the combined calculation of MQG and attention fusion subject-aware attention.

[0053] Embodiment 2

[0054] The present invention proposes a method for generating videos from images by combining a non-training strategy with multi-agent attention alignment, achieving multi-agent animation generation through agent perception and attention disentanglement, and further enriching the expressiveness of the task of generating videos from images. Specifically, agent perception is used to process the input text and images, accurately matching different text prompts with their corresponding agents, and generating a series of combinations of independent text descriptions and corresponding agents. Subsequently, after the agent attention of the original model is disentangled, the matched text and agents are bound, while other non-target agents are masked. Finally, the generated independent attention features are recombined to replace the attention mechanism in the original model. This method ensures that each agent focuses on its own area and ignores the interference of other agents, effectively avoiding the problem of action interleaving between agents in multi-agent scenarios, and achieving independent movement of each agent and high-quality animation generation effects.

[0055] In addition, the present invention uses image segmentation, text description segmentation, the alignment operation of agents therein, and the re-fusion operation of attention to align each agent and text description in the image together. The purpose is to enable all agents to move as much as possible according to the matching text semantics while masking the remaining objects. The core steps are as follows:

[0056] Step 1: Obtain the input text and pictures.

[0057] Step 2: Use text segmentation technology to independently extract the description of each agent, and generate a mask for each agent during the image segmentation process through these independent texts. To reduce the influence of other agents, each mask is processed: the pixel values of the area corresponding to the agent in the mask are set to 1, and the pixel values of the areas of other agents are set to 0. At the same time, the area representing the agent in the mask is set to 1 again to ensure that the displayed agent has greater freedom and less interference during its movement.

[0058] Step 3: Input the extracted masks, the separated texts, and the corresponding images into a pre-trained image-to-video (I2V) model based on the Diffusion model. In this way, the ability of the model can be utilized to combine the text description and mask information to generate a video result corresponding to the input content.

[0059] Step 4: After the above input pictures are passed through a pre-trained encoder to obtain latent variables and are added noise, copies are made in the same number as the number of frames of the generated video. The noise is increased by one frame dimension and all the noises are concatenated in the frame dimension.

[0060] Step 5: Use the pre-trained clip text encoder to extract the features of each agent respectively and copy the number of video frames, and store them in the list collection p in sequence.

[0061] Step 6: Feed the obtained set p, noise, and mask into the U-Net network.

[0062] Step 7: When calculating the attention in the U-Net, use the mask to obtain the features of each subject, and the subject feature q l and the matching text feature k l , v l First, calculate the similarity between q l and k l through dot product to obtain an attention score (weight), which represents the correlation between each input and other inputs. Then, normalize these weights (usually using softmax) to ensure that the sum of the weights is 1. Finally, apply these normalized weights to the corresponding value vectors to generate the attention output.

[0063] Step 8: Concatenate the multiple attentions obtained above into a single attention that finally matches the pre-trained model to replace the original attention, and then obtain a denoised latent variable.

[0064] Step 9: Use the classifier-free guidance method to generate the final hidden representation from the denoised latent variable.

[0065] Step 10: Repeat the denoising process T times to obtain a clean animated video.

[0066] Example 3

[0067] As Figure 1 shown, this example further illustrates the method in Example 2, and the specific steps are as follows:

[0068] Step 1: Obtain the input text text and image x 0 (artificial input).

[0069] Step 2: Through the text separation method (TP), separate the descriptive text containing multiple subjects, and store the description of each subject separately in a set. Then, using these separated text descriptions, adopt the visual segmentation (VP) method to more accurately segment each subject in the image, generating corresponding subject region masks (masks). Then perform perception processing on each mask: in each mask, set the pixel values of the region corresponding to the subject to 1, and the pixel values of the regions of other subjects to 0, ensuring that each mask accurately represents the region of its corresponding subject. The method for obtaining the subject region mask is as follows:

[0070] masks = perception(VP(x 0, TP (test)) ) (1)

[0071] Step 3: Input the extracted mask, the separated text, and the corresponding image into a pre-trained image-to-video (I2V) model based on the Diffusion model. In this way, the capabilities of the model can be utilized to combine the text description and the mask information to generate a video result corresponding to the input content.

[0072] Step 4: After converting the input image into latent variables through a pre-trained encoder, add noise to it. Then, duplicate this latent variable multiple times so that the number of copies is the same as the number of frames of the target generated video. Subsequently, add noise with a frame number dimension to each copy of the latent variable. Finally, concatenate these latent variables with noise along the frame number dimension to obtain a latent representation suitable for video generation. The specific processing method is as follows:

[0073] z 0 = E(x 0 ) (2)

[0074]

[0075] where E is the pre-trained encoder, and z 0 represents encoding the input image as the first frame, is the latent variable from the 0th frame to the (f - 1)th frame at the t-th moment.

[0076] Step 5: Use the CLIP text encoder to extract the features of each subject respectively, and duplicate each feature to the same number of copies as the number of video frames. Then, store these duplicated features in the list set p in sequence for subsequent use.

[0077] p = clip(TP(text)) (4)

[0078] Step 6: Input the obtained set p and the noise masks into the U-Net network.

[0079] Step 7: When calculating the cross-attention in the U-Net network, first use the mask to extract the feature q of each subject from the original multiple subject features q l . The subject feature q l performs a dot product operation with the matching text features k l and v l (k l and v l are obtained from the set p), and calculate q l and k lThe similarity between them is calculated to obtain attention scores (weights), which represent the correlation between each subject feature and other inputs. Then, these attention weights are normalized using softmax to ensure that the sum of the weights is 1. Finally, the normalized attention weights are applied to the corresponding value vector v to generate the final attention output. The specific processing steps are as follows:

[0080] q l = q · mask(5)

[0081]

[0082] where atten l is the cross-attention feature output of the l-th subject.

[0083] Step 8: The multiple attentions obtained above are simply concatenated and then normalized with the sum of all masks to obtain an attention DeAttention that matches the pre-trained image-to-video (I2V) model. This attention DeAttention replaces the attention in the original attention module in the U-Net network, and then a denoised latent variable is obtained:

[0084]

[0085] where mask l is the mask of the l-th subject, and n is the number of subjects.

[0086] Step 9: The final latent variable is generated by denoising in one step using the classifier-free guidance method. In this process, by combining unconditional and conditional latent variable predictions, the generation is guided to produce a latent representation that better conforms to the input conditions. The final latent representation is obtained after removing the noise and is a compact expression of the generated content. The way to obtain the latent representation is as follows:

[0087]

[0088] where ω is the weight, ε is the U-Net network, ε(z t , p) is the conditional latent variable, and ε(z t ) is the unconditional latent variable.

[0089] Step 10: By repeating the denoising process T times, the noise is gradually eliminated, and finally a clear animated video is generated. In each iteration, the noise gradually decreases, and the video frames become clearer until a complete denoised result is obtained, thus generating a smooth and clean animated video.

[0090] The present invention effectively solves the problem of action confusion between subjects in multi-subject generated animation by introducing subject perception and attention disentanglement mechanism. Existing I2V technology often cannot effectively distinguish different subjects when processing multiple subjects, resulting in the generated animation results, the subject's actions and features are intertwined with each other, resulting in inconsistent appearance or mismatched actions. In order to achieve clear distinction between subjects, the present invention independently segments each subject in the image at the input stage, generates a corresponding mask, and matches it one by one with the subject through text description. In this way, each subject only shows actions consistent with its text description in the animation. This method significantly improves the accuracy and coherence of the generated animation, allowing the audience to more intuitively understand the behavior and characteristics of each subject, enhancing the expressiveness and fun of the animation.

[0091] The design of the present invention adopts a training-free architecture, which greatly simplifies the complexity of model construction and integration. Existing I2V models usually need to rely on large and diverse multi-subject data sets for training, and such data sets are relatively scarce in practical applications, limiting the applicability and generalization ability of the model. At the same time, the process of traditional model training is not only time-consuming, but also consumes a lot of computing resources, which is a considerable burden for many users. In contrast, the present invention can be quickly integrated into the existing system without tedious training. This means that users only need to provide simple text and image inputs to achieve animation generation of multi-subject scenes, greatly reducing the technical threshold and improving the flexibility of application. This design not only makes it more convenient for users to use, but also effectively improves the application efficiency and practical value of the model.

[0092] The present invention realizes independent and diverse motion performance of the subject by combining text and mask information, thereby ensuring the richness and accuracy of the animation content. The existing I2V technology is often limited by the model structure and training data when generating motion, resulting in the generated motion performance being often single, mechanical, lacking in vividness and creativity. At the same time, traditional models usually require users to make complex adjustments to the generation parameters, making the operation process cumbersome and the experience poor. In order to solve these problems, the present invention not only realizes independent control of each subject through subject perception, but also through the attention disentanglement mechanism, so that each subject can move freely according to its own text description, showing richer emotions and dynamic changes. In the process of using the present invention, users can easily create high-quality animation content without having to deeply understand the complexity of the internal parameters of the model. This plug-and-play feature greatly improves the user's creative efficiency and satisfaction, allowing users to focus on the creation itself instead of being bothered by technical details.

[0093] Embodiment 4

[0094] Existing image-to-video (I2V) generation methods are mainly based on diffusion models, which have attracted much attention due to their advantages in generation controllability and high fidelity. With continuous exploration in aspects such as motion control, adjustment modules, and spatio-temporal attention mechanisms, these methods have gradually improved I2V generation technology and performed well on single-agent images. However, when dealing with multi-agent scenarios, these methods usually have difficulty effectively distinguishing different agents. Their image and text components interact in a coupled manner, which may lead to visual chaos in the generated videos, and the appearance and actions of the agents often do not match the text prompts. In fact, multi-agent images are far more common in real-world scenarios than single-agent images, and the deficiencies of existing methods in this regard severely limit the application of I2V technology.

[0095] To address this limitation, the present invention proposes a training-free diffusion method for multi-agent I2V generation. Figure 2 The training-free sampling framework designed by the present invention is shown. The VAE in the pre-trained video model encodes the video frames composed of images and noise into the latent space to generate latent variables. After the denoising process ends, the network converts the generated video back to the pixel space. The framework designed by the present invention consists of two parts: a pre-trained U-Net network and a subject perception and attention disentanglement part. The U-Net network aligns the received latent variables and text features and performs iterative denoising. Before sampling, subject perception processes the text description and image, identifies and matches each subject with its corresponding text description. During the sampling process, our attention disentanglement processes the video frames, disentangles different subjects, binds them to their respective text descriptions, and masks other irrelevant subjects. Subject perception combines visual and text information, while attention disentanglement replaces the original attention with subject perception attention, binding each subject in the video to its corresponding text prompt. In this way, the actions of each subject can be accurately aligned with their text descriptions and do not interfere with each other during the generation process.

[0096] Next, the attention unlocking part designed by the present invention will be further introduced in detail, explaining how to enable different subjects to only focus on themselves and perform independent actions. First, in the original model, the latent variable and the text prompt are input into the U-Net network to generate the original query of the video frame from the features of the input image. Next, different subjects in the original query are masked to obtain a query containing only the features of a single subject, which is stored in the q set. At the same time, the separated text is sequentially fed into CLIP to extract the text features corresponding to each subject. Based on different text inputs, the feature sets k and v are obtained. Since the separated text is used to segment the image and generate a mask through subject perception, the q obtained by mask segmentation naturally corresponds to k and v without additional matching. Attention fusion performs cross-attention calculations on q, k, and v to generate the attention values of multiple subjects. Finally, these attention values are added and normalized together with the merged mask to generate a decoupled single attention map (DeAttention, Decoupled Attention). The partial data experimental results of the present invention are as shown in Figure 3 (a) and (b) in

[0097] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for generating video from images based on a non-training strategy and multi-agent attention alignment, characterized in that: The following steps are involved: S1: Get input text and pictures; S2: Separate the input text and segment the input image to obtain the mask of each subject, and perform perceptual processing on the mask; S3: The extracted mask, separated text and corresponding input image are input into the pre-trained image-to-video model based on the diffusion model; S4: Encode the image and then make multiple copies, adding noise to it; S5: Send the separated text to the CLIP text encoder to extract the features of each subject respectively, and copy each subject feature to the same number of copies as the number of target generated video frames, and store these copied subject features in the list set p in turn; S6: Input the above-obtained set p, each subject region mask, and noise into the U-Net network; S7: When calculating cross attention in the U-Net network, a mask is used to obtain the feature q of each subject from multiple subject features q l , main feature q l The text feature k that matches l and v l First calculate q by dot product l With k l , we get an attention weight to represent the correlation between each input and other inputs, and then normalize the attention weights. Finally, we apply these normalized attention weights to the corresponding value vectors to generate attention outputs, where k l and v l Get from the set p, l represents the lth subject; S8: Multiple attention outputs are concatenated and normalized with the sum of all masks to obtain an attention DeAttention that matches the pre-trained image-to-video model based on the diffusion model to replace the attention in the original attention module in the U-Net network, and then a one-step denoising latent variable is obtained; S9: Generate the final latent variables and obtain the latent representation using the classifier-free guidance method through one-step denoising; S10: By repeating denoising T times, the noise is gradually eliminated to obtain the target generated video.

2. The image-generated video generation method based on non-training strategy and multi-subject attention alignment according to claim 1 is characterized in that: In step S2, the specific processing process is as follows: S21: separating the descriptive text containing multiple subjects through input text separation processing, and storing the description of each subject separately in a set; S22: using the separated text description, segment each subject in the image through visual segmentation processing to generate a corresponding subject area mask; S23: Then, perceptual processing is performed on each mask: in each mask, the pixel value of the area corresponding to the subject is set to 1, and the pixel values ​​of the areas of other subjects are set to 0, so that each mask accurately represents the area of ​​its corresponding subject.

3. The image-generated video generation method based on non-training strategy and multi-subject attention alignment according to claim 2 is characterized in that: In step S4, the specific processing process is as follows: S41: After the input image is converted into a latent variable through a pre-trained encoder, it is subjected to noise processing; S42: Duplicate the latent variables multiple times so that their number is consistent with the number of frames of the target generated video; S43: Add a noise of one frame dimension to each latent variable; S44: Concatenate the noisy latent variables along the frame number dimension to obtain a latent representation adapted to video generation.

4. The image-generated video generation method based on non-training strategy and multi-subject attention alignment according to claim 3 is characterized in that: In step S7, the subject feature q l The calculation formula is as follows: q l =q·mask; The calculation formula of attention output is as follows: Among them, atten l is the attention output of the alignment of the text features of the lth subject with the subject features, and d is q l and k l Dimension.

5. The image-generated video generation method based on non-training strategy and multi-agent attention alignment according to claim 4 is characterized in that: In step S9, the formula for obtaining the potential representation is as follows: Among them, ω is the weight, ε is the U-Net network, ε(z t ,p) is a conditional latent variable, ε(z t ) is the unconditional latent variable.

Citation Information

Patent Citations

  • Spatial decoupling personalized multi-subject text video method, device and equipment

    CN118505866A

  • Graphics video model generation method, video generation method and device

    CN118747862A

  • MotionBooth framework-based customized object dynamic video generation method

    CN118869902A

  • Search-based text video method

    CN119011969A

  • Video generation method and system based on large-scale text video model

    CN119031209A

Cited By

  • Cross-modal semantic attention collaborative enhancement video subtitle generation method and system

    CN121009887A