Method for obtaining frame sequence from text and related device
By obtaining frame description text and layout information, combining the embedding of input text and reference images, using the stable diffusion model and space-time self-attention mechanism, the problem of role in the multi-character image frame sequence is solved, and a high-quality, multi-character image frame sequence is generated.
Patent Information
- Application Number
- CN202510758551.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-09
AI Technical Summary
It is difficult for prior art to maintain the consistency of characters between different frames when generating sequences of multi-character image frames, for example, the hair color of the same character in the previous frame may be different from the next frame.
By obtaining frame description text and layout information, combining the embedding of input text and reference images, using a stable diffusion model and space-time self-attention mechanism, the image and text embedding are fused to generate frame sequences, ensuring that the positional relationship of characters in different frames is consistent.
It effectively reduces the possibility that characters affect each other in frames, ensures the consistency between frames, and generates high-quality, multi-character image frame sequences.
Smart Images

Figure CN120279470A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of text-to-image technology, and in particular, to a method for obtaining a frame sequence from text and related devices. Background Art
[0002] Story visualization aims to generate attractive images or videos for text narratives. With the development of text-to-image technology using models, it has become possible to generate images containing multiple characters based on text containing multiple characters.
[0003] However, the current problem is that it is difficult to keep the same character consistent between consecutive frames. For example, the hair of a little boy is black in the previous frame and red in the next frame.
[0004] Therefore, how to generate multiple characters contained in text in one frame image and keep the characters consistent between different image frames is a problem that needs to be solved currently. Summary of the Invention
[0005] In view of the above problems, this application provides a method for obtaining a frame sequence from text and related devices to achieve the purpose of generating multiple characters contained in text in one frame image and keeping the characters consistent between different image frames. The specific solutions are as follows:
[0006] The first aspect of this application provides a method for obtaining a frame sequence from text, including:
[0007] Based on the input text, obtain frame description text and layout information of each frame. The frame includes an image frame for expressing the input text, and the layout information of the target frame indicates the spatial information of each character in the target frame;
[0008] Obtain the text embedding of the input text;
[0009] Based on the reference images of each character, obtain the image embedding of the reference image, where the character is the character expressed in the input text;
[0010] Fuse the image embedding with the layout information to obtain first fusion information, where the first fusion information represents the first positional relationship of each image embedding in each frame;
[0011] Fuse the text embedding with the layout information to obtain second fusion information, where the second fusion information represents the second positional relationship of each text embedding in each frame;
[0012] Based on the spatio-temporal self-attention mechanism, fuse the information to obtain a frame sequence including the image frames, and the information for fusion includes the first fusion information and the second fusion information.
[0013] In one possible implementation, obtaining the text embedding of the input text and the image embedding of the reference image includes:
[0014] Encoding the input text into the conditional space of the Stable Diffusion model to obtain the text embedding of the input text;
[0015] Encoding the reference image into the conditional space of the Stable Diffusion model to obtain the image embedding of the reference image;
[0016] Aligning the text embedding and the image embedding.
[0017] In one possible implementation, before obtaining the image embedding of the reference image, it further includes:
[0018] Removing information irrelevant to the character from the reference image.
[0019] In one possible implementation, fusing the image embedding with the layout information to obtain the first fusion information includes:
[0020] Performing a first operation on the hidden state of the attention mechanism and each image embedding to obtain a first parameter;
[0021] Performing a second operation on the first parameter and a first mask matrix to obtain the hidden state update components of each character, where the first mask matrix is obtained based on the layout information;
[0022] Fusing the hidden state update components of each character to obtain the first fusion information.
[0023] In one possible implementation, fusing the embedding features of the text with the layout information to obtain the second fusion information includes:
[0024] Performing a first operation on the hidden state of the attention mechanism and each text embedding to obtain a second parameter;
[0025] Performing a third operation on the second parameter, an area balance matrix, and a second mask matrix to obtain a third parameter, where the area balance matrix is used to balance the area occupied by each character in the target frame, and the second mask matrix is obtained based on the layout information;
[0026] Performing a fourth operation on the third parameter and each text embedding to obtain the second fusion information.
[0027] In one possible implementation, the information for fusion further includes: information of the character extracted from an image including a single character.
[0028] In a possible implementation, obtaining the frame description text and the layout information of each frame based on the input text includes:
[0029] Outputting the input text as the description text and layout information of multiple image frames based on a language model.
[0030] A second aspect of the present application provides an apparatus for obtaining a frame sequence from text, including:
[0031] A first acquisition module, configured to obtain the frame description text and the layout information of each frame based on the input text, where the frame includes an image frame for expressing the input text, and the layout information of the target frame indicates the spatial information of each role in the target frame;
[0032] A second acquisition module, configured to obtain the text embedding of the input text;
[0033] A third acquisition module, configured to obtain the image embedding of the reference image based on the reference images of each role, where the role is the role described in the input text;
[0034] A first fusion module, configured to fuse the image embedding with the layout information to obtain first fusion information, where the first fusion information represents the first position relationship of each image embedding in each frame;
[0035] A second fusion module, configured to fuse the text embedding with the layout information to obtain second fusion information, where the second fusion information represents the second position relationship of each text embedding in each frame;
[0036] A third fusion module, configured to fuse the information based on a spatio-temporal self-attention mechanism to obtain a frame sequence including the image frames, and the information for fusion includes the first fusion information and the second fusion information.
[0037] A third aspect of the present application provides a computer program product, including computer-readable instructions, which when running on an electronic device, enable the electronic device to implement the method for obtaining a frame sequence from text according to the first aspect or any implementation manner of the first aspect.
[0038] A fourth aspect of the present application provides an electronic device, including at least one processor and a memory connected to the processor, where:
[0039] The memory is used to store a computer program;
[0040] The processor is used to execute the computer program so that the electronic device can implement the method for obtaining a frame sequence from text according to the first aspect or any implementation manner of the first aspect.
[0041] The fifth aspect of the present application provides a computer-readable storage medium carrying one or more computer programs, which can enable an electronic device to implement the method for obtaining a frame sequence from text described in the first aspect or any implementation manner of the first aspect when the one or more computer programs are executed by the electronic device.
[0042] By means of the above technical solution, the method and related device for obtaining a frame sequence from text provided by the present application, based on the input text, obtain the frame description text of the image frames expressing the input text and the layout information of each frame, obtain the text embedding of the input text, and based on the reference images of each role, obtain the image embedding of the reference images, where the role is the role described in the input text. The image embedding is fused with the layout information to obtain the first fusion information, and the text embedding is fused with the layout information to obtain the second fusion information. Based on the spatio-temporal self-attention mechanism, the information is fused to obtain a frame sequence including image frames. Therefore, in the case of multiple roles, multiple roles can be included in one frame image. Also, because the first fusion information represents the first position relationship of each image embedding in each frame, the second fusion information represents the second position relationship of each text embedding in each frame, and the layout information of the target frame indicates the spatial information of each role in the target frame. Therefore, in the information used to obtain the final frame sequence, the spatial information of the roles is fused, so the possibility of mutual influence between different roles in the frame can be reduced, and thus the possibility of one role being affected by other roles can be reduced, which is beneficial to ensuring the consistency of any role between frames in the case of multiple roles. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original elements and elements are not necessarily drawn to scale.
[0044] Figure 1 It is a flowchart of a method for obtaining a frame sequence from text provided by the present application;
[0045] Figure 2 It is a layout example diagram reflecting the layout information provided by the present application;
[0046] Figure 3 It is an architecture diagram for obtaining a frame sequence from text based on a model provided by the present application;
[0047] Figure 4 It is a functional example of obtaining the first fusion information based on the consistency image cross-attention model and obtaining the second fusion information based on the consistency text cross-attention model provided by the present application;
[0048] Figure 5Schematic structural diagram of a device for obtaining a frame sequence from text provided by this application;
[0049] Figure 6 Schematic structural diagram of an electronic device provided by this application. Specific embodiments
[0050] The following describes the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. The terms used in the embodiment part of this application are only used to explain the specific embodiments of this application, and are not intended to limit this application.
[0051] The following describes the embodiments of this application in conjunction with the accompanying drawings. Those of ordinary skill in the art will know that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0052] The terms "first", "second", etc. in the specification, claims and above-mentioned accompanying drawings of this application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing when describing objects with the same attributes in the embodiments of this application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0053] First, the models involved in the embodiments of this application will be described.
[0054] Diffusion model: A generative model that can generate realistic images, videos and other high-dimensional data, and is widely used in multiple fields, achieving remarkable results in high-quality image generation, such as text-to-image models and story visualization generation models. Its main working principle includes a forward diffusion process, gradually adding random noise to clean data until pure noise, mapping the real data distribution to a standard Gaussian distribution; a reverse diffusion process starts from pure noise and gradually removes the noise to obtain a clear image, and gradually approaches the real image by predicting and removing the noise through a neural network. In the training stage, the model learns to predict the noise level and understand how to strip the noise; in the generation stage, it starts from random noise and gradually removes the noise using reverse diffusion to generate an image. The advantages of the diffusion model include high-quality generation, flexibility (supporting conditional generation and combination with other models), and stability (less likely to have problems such as mode collapse compared to generative adversarial networks).
[0055] Stable Diffusion model: It is an efficient text-to-image generation model based on the diffusion model. It can convert text descriptions into images by gradually removing noise, performs excellently in generating high-quality and detail-rich images, and is more efficient in using computing resources.
[0056] The Stable Diffusion model has a close relationship with the diffusion model. First, their basic principles are the same. The Stable Diffusion model continues the noise addition and removal process of the diffusion model, and the steps of generating images are consistent. Second, the Stable Diffusion model has been optimized and improved. By improving the model structure and using techniques such as the latent space, it reduces the computing requirements and improves the generation speed, making it more suitable for practical application scenarios.
[0057] The Stable Diffusion model is designed specifically for text-to-image generation and can generate images that meet the requirements according to user input.
[0058] Figure 1 A method for obtaining a frame sequence from text provided by an embodiment of this application includes the following steps:
[0059] S11. Based on the input text, obtain the frame description text and the layout information of each frame, and obtain the reference images of each character.
[0060] The frame includes an image frame for expressing the input text. For example, generate an image frame that can express the content described in the input text from the input text, and the image frames can form a multimedia file such as a video.
[0061] A character is a character described in the input text. For example, if the text content includes "girl" and "man", then "girl" and "man" are the characters.
[0062] Any one of the image frames is called a target frame, and the layout information of the target frame indicates the quantity and spatial information of each character in the target frame. The spatial information includes size and position. In some examples, the size is represented by the coordinates of the four vertices of a rectangular bounding box, and the position is represented by up, down, left, and right, etc.
[0063] In some implementation manners, the implementation manner of S11 is:
[0064] 110. The user inputs a simple text to a large language model (abbreviated as LLM), such as "Write a short story about two girls, Little A and Little B. Do not use 'he','she', 'it', or 'they' in the story. Generally, do not use 'a person', 'they', 'a girl', 'the trio' to address the subject. When you describe the subject, be sure to use their names!".
[0065] 111. The LLM outputs a story based on simple text to meet the requirements of the simple text. For example, "In a bustling city, there lived two extraordinary girls named Little A and Little B. Little A is a lively and adventurous person with bright blue eyes and curly golden hair, while Little B is a quiet and determined girl with long black hair and a mischievous smile. They are inseparable friends, always exploring new places, seeking adventures, and sharing their dreams..."
[0066] It is understandable that the text output by 111 is an expansion of the simple text.
[0067] 112. The above story text is the input text. The LLM divides the input text into multiple paragraphs based on this input text.
[0068] 113. Each paragraph corresponds to an image frame, and the LLM outputs the description information of the picture of each image frame.
[0069] The picture description information includes description text and layout information.
[0070] The description text of the target frame is used to describe the content expressed by the target frame. Continuing with the above example, such as "They are inseparable friends, always exploring new places, seeking adventures". An example of the layout information is "1. Little A, [a1, b1, c1, d1], 2. Little B, [a2, b2, c2, d2]". 1 and 2 represent the numbers of the characters, Little A and Little B are the names of the characters, and [a1, b1, c1, d1] and [a2, b2, c2, d2] represent the sizes of the areas occupied by the characters in the target frame. The layout reflected by this layout information is as Figure 2 shown, Figure 2 in which the first square from the left is the area represented by [a1, b1, c1, d1], and the second square is the area represented by [a2, b2, c2, d2].
[0071] The reference image is input by the user. It is understandable that the user can design the images of each character and input them in the form of a reference image to obtain the character images that meet the user's requirements.
[0072] In some implementation manners, the description text includes a global description and a detailed description of the character. The global description is the outline of the story of the target frame, and the detailed description is the detailed description of the character obtained by the LLM based on the understanding of the reference image of the character.
[0073] S12. Obtain the text embedding of the input text and the image embedding of the reference image.
[0074] Embedding refers to embedding features.
[0075] The IP-Adapter mainly uses a pre-trained Contrastive Language-Image Pre-training (CLIP) image encoder to extract image features, and adjusts the feature dimensions through normalization by a linear layer and a mapping layer. The key lies in decoupling the cross-attention module, integrating the image features into the generation process of the diffusion model, enabling the model to generate images according to image prompts, and having fewer parameters, which can efficiently enable the pre-trained text-to-image diffusion model to have the ability to utilize image prompts.
[0076] In some implementation manners, the input text is encoded through a pre-trained CLIP image encoder to obtain an encoding result mapped to the conditional space of the Stable Diffusion model, that is, a text embedding is obtained. That is to say, the input text is encoded into the conditional space of the Stable Diffusion model to obtain the text embedding of the input text.
[0077] In some implementation manners, an image is encoded into the conditional space of the Stable Diffusion model through a pre-trained CLIP image encoder to obtain an image embedding of a reference image.
[0078] In order to improve the accuracy of the characters in the finally generated frame sequence, in some implementation manners, before obtaining the image embedding of the reference image, information irrelevant to the characters in the reference image is removed through the mapping layer of the pre-trained IP-Adapter. The information irrelevant to the characters may be information such as background, composition, and style.
[0079] In some implementation manners, after obtaining the text embedding and the image embedding, in order to facilitate subsequent processing, the text embedding and the image embedding are aligned. Exemplarily, the alignment of the text embedding and the image embedding is achieved through the mapping layer of the pre-trained IP-Adapter.
[0080] It can be understood that implementing the above process through the IP-Adapter is only an example and is not limited to the IP-Adapter.
[0081] S13. Integrate the image embedding with the layout information to obtain first fusion information.
[0082] Because the layout information represents the sizes and positions of the respective characters in the target frame, after being integrated with the image embedding, the spatial position relationships of the image embeddings of different characters in the target frame are strengthened. That is, the first fusion information represents the position relationships of the respective image embeddings in each frame. For the sake of distinction, it is called the first position relationship here.
[0083] S14. Integrate the text embedding with the layout information to obtain second fusion information.
[0084] Similarly, the second fusion information represents the positional relationship between each text embedded in each frame, which is referred to as the second positional relationship here.
[0085] It can be understood that the positional relationship between each image embedded in each frame and the positional relationship between each text embedded in each frame can reduce the mutual interference of characters in each frame.
[0086] In this embodiment, image embedding and text embedding are fused with layout information respectively, that is, image embedding and text embedding are fused with layout information respectively, which can enhance the influence of position relationship on embedding, that is, better avoid mutual interference between characters, thereby facilitating obtaining a more accurate frame sequence.
[0087] In some implementations, S13 and S14 are implemented using a cross-attention mechanism, which will be described in the following embodiments.
[0088] S15. Based on the spatiotemporal self-attention mechanism, the information is fused to obtain a frame sequence including image frames.
[0089] The information to be fused includes first fused information and second fused information.
[0090] In some implementations, the information to be fused further includes: information of each character. The information of any one character in the information of each character refers to the information of the character extracted from the image including the single character after the image including the single character is generated.
[0091] Exemplarily, the process of extracting information of each character includes: generating an image including a single character based on the IP-Adapter and a detailed description of the single character, and extracting the character information from each reference image through GroundedSAM.
[0092] GroundedSAM is a visual application that integrates Grounding DINO (a large model for detecting targets) and Segment Anything (SAM) models. It uses the zero-sample detection capability of Grounding DINO to find objects in images through text input, and then uses the fine-grained segmentation capability of SAM, and can also combine Stable Diffusion to generate text and images for segmented areas. The model can achieve tasks such as automatic annotation, visual question answering, and image description generation.
[0093] The spatio-temporal self-attention mechanism is usually used to capture both temporal and spatial information in videos simultaneously. The role of this mechanism is to help the model identify and generate dynamic changes between frames while maintaining the consistency of spatial details in each frame. In spatio-temporal self-attention, the model not only focuses on the pixel relationships within the current frame but also exchanges information between different frames to understand and generate natural motion continuity. This method is crucial for generating high-quality videos as it ensures the temporal consistency of object shapes, positions, and actions.
[0094] Figure 1 In the shown process, both the image embedding and the text embedding are fused with the layout information, which represents the spatial information of each character. Therefore, it can strengthen the influence of spatial information on the generated image, reduce the interference between different characters in each frame, and is conducive to maintaining the consistency of the same character between frames when generating image frames with multiple characters.
[0095] The above process will be described in more detail below in combination with the model structure.
[0096] Take Figure 3 as an example. After the reference image of character 1 and the reference image of character 2 pass through the image encoder, they are encoded into the image embedding of character 1 and the image embedding of character 2. The description text of frame 1 and the description text of frame 2, after passing through the text encoder, are encoded into the text embedding of frame 1 and the text embedding of frame 2.
[0097] The image encoder and the text encoder can be, but are not limited to, the aforementioned IP-Adapter.
[0098] The image embedding of character 1, the image embedding of character 2, as well as the layout information of frame 1 and the layout information of frame 2 are respectively input into the consistency image cross-attention model to obtain the first fusion information, which will be specifically described in combination with Figure 3 Frame 1's text embedding, frame 2's text embedding, as well as frame 1's layout information and frame 2's layout information are respectively input into the consistency text cross-attention model to obtain the second fusion information, which will be specifically described in combination with Figure 3 Frame 1's text embedding, frame 2's text embedding, as well as frame 1's layout information and frame 2's layout information are respectively input into the consistency text cross-attention model to obtain the second fusion information, which will be specifically described in combination with
[0099] Based on the reference images of character 1 and character 2, as well as the detailed descriptions of frame 1 and frame 2, the hidden state guidance model outputs the information of each character. For the specific implementation method, refer to S15 for details.
[0100] The information of each character and the noise data are used as the input of the denoising U-Net network. The denoising U-Net network outputs denoised data, which, after passing through the decoder, generates an image frame containing character 1 and character 2 and can express the description text of frame 1 and the description text of frame 2. The denoising U-Net network can adopt the spatio-temporal self-attention mechanism.
[0101] Among them, Figure 3 the image encoder, text encoder, and denoising U-Net network in
[0102] can be various parts in the Stable Diffusion model. Figure 3 Based on the architecture shown in
[0103] Figure 3 , during the cross-attention calculation in the sampling stage, the image and text control conditions are decoupled. First, the calculation between the text condition and the hidden state is performed to obtain the updated hidden state. Then, the cross-attention between each character reference map condition and the hidden state is carried out, and the hidden state is updated with a fixed weight coefficient.
[0104] To more effectively use the reference maps of multiple characters as control conditions to guide generation, a consistency image cross-attention mechanism is proposed. By using the layout information of the frame (represented by the bounding box of each character), a mask for each character in the cross-attention score matrix is obtained to explicitly control the influence area of different character reference maps on the hidden state, thereby suppressing the interference between different character references. For the case where multiple characters overlap in the picture, averaging the attention score calculation for the overlapping area can effectively prevent the fusion of character features in the overlapping area.
[0105] Fine text (i.e., fine description) is a text description with a long length and rich details. However, long text will cause semantic leakage between different subjects during the cross-attention process (for example, the description belonging to subject A is overly attended to by subject B, interfering with the generation of B). In response, a consistency text cross-attention mechanism calculates the cross-attention mask of each subject in the hidden space according to the position of different subject descriptions, and combines the area size relationship of different subjects obtained from the layout information to adjust the distribution of the text cross-attention scores, so that the subject is not interfered by the descriptions of other subjects and at the same time prevents the interference of the subject size on the attention distribution.
[0106] As Figure 4 shown, Figure 3 the function example of the consistency image cross-attention model in
[0107] h represents the hidden state parameter, also known as the hidden state, which is the hidden state passed into the cross-attention layer (taking the cross-attention layer of the consistency image cross-attention model as an example). Performs a first operation with each image embedding (also known as IP Image Embeddinggs). to obtain a first parameter. An example of the first operation is , where is the first parameter, i.e., the attention score matrix, d represents a preset scaling factor, and softMax represents a normalization operation. Figure 4 In, ip represents the role.
[0108] , is the information of the role (abbreviated as role information) extracted from the image including the single role output by the hidden state guidance model. is the update of the hidden state based on the single role. This operation is an optional step. In the case of performing this step, the first parameter is , is an intermediate parameter for obtaining the first parameter ( Figure 4 is taken as an example here).
[0109] The first mask matrix generated by the layout information (Layout) is , represents the number of subjects (i.e., roles) in the target frame. Taking Figure 3 and Figure 4 as examples, is 2. Among them, is the mask of role ip1, is the mask of role ip2.
[0110] The second operation = ⊙ , where is the mask of any role Performs a mask operation on so that the reference map information of each role only takes effect on the hidden state of its corresponding region, obtaining the hidden state update components of each role. Fusing the hidden state update components of each role, the first fusion information is obtained. The first fusion information contains two parts, Figure 4 In, the first part a represents , and the second part b represents .
[0111] As Figure 4 shown, Figure 3 The functional example of obtaining the second fusion information by the consistency text cross-attention model in is as follows:
[0112] The hidden state passed into the cross-attention layer (taking the cross-attention layer of the consistency text cross-attention model as an example here) (for the sake of distinction, denoted as ) performs a first operation with each text embedding (or the operation result of the text embedding) to obtain a second parameter. For an example of the first operation, refer to the method of obtaining the first parameter, which will not be elaborated here.
[0113] The second parameter is subjected to a third operation with the area scale matrix (Area Scale Matrix) S{S 0, S 1, S2}, the second mask matrix R{r 0, r 1, r2} to obtain a third parameter. Exemplarily, the third operation is to perform a matrix multiplication operation on the second parameter, the area scale matrix, and the second mask matrix (denoted by ), and then perform a normalization operation (softMax) on the result of the matrix multiplication operation.
[0114] Among them, the area scale matrix (Area Scale Matrix, abbreviated as S) is used to balance the area occupied by each role in the target frame. That is to say, the area scale matrix S is used to balance the adjustment degree of the attention scores of large-area and small-area subjects, enhance the attention of small subjects to their corresponding text control, and avoid the loss of small subjects during the generation process.
[0115] The area scale matrix S includes the area balance coefficients of each role in the target frame. The smaller the area occupied by the role, the larger the area balance coefficient. In some implementation manners, the reciprocal of the area balance coefficient of any role is the ratio of the area occupied by the role to the total area occupied by all roles. Exemplarily, the calculation formula for the area balance coefficient of any role is , is the area scale matrix of role i, is the h dimension of the hidden state, is the w dimension of the hidden state, represents the length of R, and Area_i represents the area size occupied by the i-th ip on the compressed picture or the hidden state. Figure 4 In is the area scale matrix of role 0, taking 1.0 as an example, is the area scale matrix of role 1, taking 1.4 as an example, is the area scale matrix of role 2, taking 1.2 as an example.
[0116] The second mask matrix R is obtained based on the layout information. For example, each column in the second mask matrix is the mask of the role corresponding to that column, which is used to indicate the position of the text embedding that needs to be noted in the hidden state. Figure 4In this case, 1 in the second mask matrix indicates that the element value is 1, and -inf indicates negative infinity. It is the second mask matrix for role 0. It is the second mask matrix for role 1. It is the second mask matrix for role 2. The way to obtain the second mask matrix R is as follows: Based on the layout information, obtain the mask matrix for each role (ip). Figure 4 Taking m0 to represent the mask matrix of role 0, m1 to represent the mask matrix of role 1, and m2 to represent the mask matrix of role 2 as examples respectively, flatten the mask matrix of the i-th (such as i being 0, 1, or 2) ip into a long column, and then copy this column ci times (the length of the text description part corresponding to the i-th ip). In this way, the part of the i-th ip in the regional attention control matrix (i.e., the second mask matrix) R is obtained. Repeat the above operation for each ip, and then splice them together to obtain the complete R.
[0117] Perform the fourth operation on the third parameter and each text embedding (or the operation result of the text embedding) to obtain the second fusion information.
[0118] The operation method for fusing the first fusion information and the second fusion information is: , is the fusion result of the first fusion information and the second fusion information, is the second fusion information, is the first fusion information.
[0119] Taking Figure 4 as an example, the consistent image cross-attention model fuses the layout information of each frame with the hidden state fused with the image embedding again to achieve the purpose of influencing the hidden state based on the layout information. Similarly, the consistent text cross-attention model can also achieve the purpose of influencing the hidden state by the layout information.
[0120] Figures 2 - 4 It has the following advantages:
[0121] 1. It is fast and efficient without training or fine-tuning. It does not require a complex and time-consuming training process, a large number of delicate annotations, and data.
[0122] 2. By combining large language models and vision-language models, it can generate high-quality, delicate, rich descriptions, which significantly improves the generation effect.
[0123] 3. It has the ability to generate multiple roles, and the roles highly comply with their own reference images and text descriptions.
[0124] 4. It can automatically generate stories according to user prompts in an interactive form, reducing the user's usage complexity and labor costs.
[0125] The method for obtaining a frame sequence from text provided in the embodiments of the present application is introduced above. The apparatus for executing the method for obtaining a frame sequence from text will be introduced below.
[0126] Figure 5 It is a schematic structural diagram of an apparatus for obtaining a frame sequence from text provided in the embodiments of the present application. As Figure 5 shown, the apparatus includes: a first acquisition module, a second acquisition module, a third acquisition module, a first fusion module, a second fusion module, and a third fusion module.
[0127] The first acquisition module is configured to obtain a frame description text and layout information of each frame based on the input text. The frame includes an image frame for expressing the input text, and the layout information of the target frame indicates the spatial information of each character in the target frame.
[0128] The second acquisition module is configured to obtain a text embedding of the input text.
[0129] The third acquisition module is configured to obtain an image embedding of the reference image based on the reference images of each character, where the character is the character described in the input text.
[0130] The first fusion module is configured to fuse the image embedding with the layout information to obtain first fusion information, and the first fusion information represents the first positional relationship of each image embedding in each frame.
[0131] The second fusion module is configured to fuse the text embedding with the layout information to obtain second fusion information, and the second fusion information represents the second positional relationship of each text embedding in each frame.
[0132] The third fusion module is configured to fuse the information based on a spatio-temporal self-attention mechanism to obtain a frame sequence including the image frames, and the information for fusion includes the first fusion information and the second fusion information.
[0133] In a possible implementation, the manner in which the second acquisition module obtains the text embedding includes: encoding the input text into the conditional space of a stable diffusion model to obtain the text embedding of the input text. The manner in which the third acquisition module obtains the image embedding of the reference image includes: encoding the reference image into the conditional space of a stable diffusion model to obtain the image embedding of the reference image. The third acquisition module is further configured to: align the text embedding and the image embedding.
[0134] In a possible implementation, the apparatus further includes: a removal module configured to remove information irrelevant to the character from the reference image.
[0135] In a possible implementation, the way for the first fusion module to fuse the image embedding with the layout information to obtain the first fusion information includes: performing a first operation on the hidden state of the attention mechanism and each image embedding to obtain a first parameter; performing a second operation on the first parameter and a first mask matrix to obtain the updated hidden state components of each role, where the first mask matrix is obtained based on the layout information; and fusing the updated hidden state components of each role to obtain the first fusion information.
[0136] In a possible implementation, the way for the second fusion module to fuse the text embedding with the layout information to obtain the second fusion information includes: performing a first operation on the hidden state of the attention mechanism and each text embedding to obtain a second parameter; performing a third operation on the second parameter, an area balance matrix, and a second mask matrix to obtain a third parameter, where the area balance matrix is used to balance the area occupied by each role in the target frame, and the second mask matrix is obtained based on the layout information; and performing a fourth operation on the third parameter and each text embedding to obtain the second fusion information.
[0137] In a possible implementation, the device further includes an extraction module for extracting information of a role from an image including a single role.
[0138] In a possible implementation, the way for the first acquisition module to acquire the frame description text and the layout information of each frame based on the input text includes: outputting the input text as the description text and layout information of multiple image frames based on a language model.
[0139] The device can generate multiple roles included in the text in one image frame and can also maintain the consistency of the roles between different image frames.
[0140] An electronic device is further provided in an embodiment of the present application. Refer to Figure 6 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiment of the present application. The electronic device in the embodiment of the present application may include, but is not limited to, fixed terminals such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), desktop computers, and the like. Figure 6 The shown electronic device is only an example and should not bring any limitation to the functions and usage scope of the embodiment of the present application.
[0141] Such as Figure 6As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0142] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or wirelessly to exchange data. Although Figure 6 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0143] In an embodiment of the present application, a computer program product is further provided, including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement any method for obtaining a frame sequence from text provided in an embodiment of the present application.
[0144] In an embodiment of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium carries one or more computer programs, which, when executed by an electronic device, can enable the electronic device to implement any method for obtaining a frame sequence from text provided in an embodiment of the present application.
[0145] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the present application, the connection relationship between the modules indicates that they have a communication connection, which may be specifically implemented as one or more communication buses or signal lines.
[0146] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for this application, in more cases, software program implementation is a better embodiment. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.
[0147] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0148] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)), etc.
Claims
1. A method for obtaining a frame sequence from text, characterized in that, Comprising: Based on the input text, obtain the frame description text and the layout information of each frame. The frame includes an image frame for expressing the input text, and the layout information of the target frame indicates the spatial information of each character in the target frame; Obtain the text embedding of the input text; Based on the reference images of each character, obtain the image embedding of the reference image, where the character is the character described in the input text; Fuse the image embedding with the layout information to obtain first fusion information, which represents the first positional relationship of each image embedding in each frame; Fuse the text embedding with the layout information to obtain second fusion information, which represents the second positional relationship of each text embedding in each frame; Based on the spatio-temporal self-attention mechanism, fuse the information to obtain a frame sequence including the image frames. The information for fusion includes the first fusion information and the second fusion information.
2. The method according to claim 1, characterized in that, The process of obtaining the text embedding and the image embedding includes: Encode the input text into the conditional space of the Stable Diffusion model to obtain the text embedding of the input text; Encode the reference image into the conditional space of the Stable Diffusion model to obtain the image embedding of the reference image; Align the text embedding and the image embedding.
3. The method according to claim 2, wherein Before obtaining the image embedding of the reference image, it further includes: Remove the information unrelated to the character from the reference image.
4. The method according to claim 1, wherein The fusing the image embedding with the layout information to obtain first fusion information includes: Perform a first operation on the hidden state of the attention mechanism and each image embedding to obtain a first parameter; Perform a second operation on the first parameter and a first mask matrix to obtain the updated hidden state components of each character, where the first mask matrix is obtained based on the layout information; Fuse the updated hidden state components of each character to obtain the first fusion information.
5. The method according to claim 1 or 4, characterized in that, The fusing the text embedding with the layout information to obtain second fusion information includes: Perform a first operation on the hidden state of the attention mechanism and each text embedding to obtain a second parameter; Perform a third operation on the second parameter, an area balance matrix, and a second mask matrix to obtain a third parameter. The area balance matrix is used to balance the area occupied by each character in the target frame, and the second mask matrix is obtained based on the layout information; Perform a fourth operation on the third parameter and each text embedding to obtain the second fusion information.
6. The method according to claim 1, characterized in that The information for fusion further includes: the information of the character extracted from the image including a single character.
7. The method according to any one of claims 1-4, characterized in that The obtaining the frame description text and the layout information of each frame based on the input text includes: Based on the language model, output the input text as the description text and layout information of multiple image frames.
8. A computer program product, characterized in that, Including computer-readable instructions, when the computer-readable instructions run on an electronic device, enabling the electronic device to implement the method for obtaining a frame sequence from text as described in any one of claims 1 to 7.
9. An electronic device, characterized in that, Including at least one processor and a memory connected to the processor, where: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the method for obtaining a frame sequence from text as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the method for obtaining a frame sequence from text as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-modal input video condition generation method based on generative adversarial network
CN115345970A
Image synthesis method and device, electronic equipment and storage medium
CN118015138A
Role video generation method and device, electronic equipment and storage medium
CN118015159A
Video generation method and related device
CN118172449A
Spatial decoupling personalized multi-subject text video method, device and equipment
CN118505866A