Image Generation Method, Apparatus, Device, and Storage Medium
By combining the features of the current text and the previous image in the image generation model, the current image is generated, and the image inconsistency caused by text changes is solved, and the continuity and consistency of images in picture book generation is achieved.
Patent Information
- Application Number
- CN202510301466.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-14
AI Technical Summary
In the prior art, when the text content changes, the generated images have poor consistency before and after, and the diffusion model is highly sensitive to text, resulting in inconsistent images.
The previous image corresponding to the current text and the previous text is input into the autoregression module in the image generation model, and the current token sequence is obtained, and the previous image is input into the image feature adaptation module in the image generation model, and the current image is generated by combining the image features and text features through the decoupling cross attention module.
Improve the consistency between the generated images, especially in picture book generation, ensuring the continuity and consistency of story characters and scenes.
Smart Images

Figure CN119810266B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to an image generation method, device, equipment, and storage medium. Background Art
[0002] With the rapid development of artificial intelligence technology, the application of text-to-image is becoming more and more widespread. In fields such as picture book generation, it is usually necessary to depict a complete story through a series of consecutive images, which requires the generated images to have continuity, including the consistency of the story characters' images, the consistency of the scenes, and the consistency of the image styles, etc.
[0003] In the prior art, it is usually based on a diffusion model, and the consistency of the generated images is ensured by editing the content of the text. For example, the description of the character image is kept consistent, so that the characters in the generated images are as similar as possible in appearance, and a stylized description is made in the text, so that the generated images are as close as possible in style. In addition, the random number sampling is kept consistent, so that the content of the front and back images is as close as possible.
[0004] However, due to the continuous change and development of the description of the picture book story content, and the high sensitivity of the diffusion model to the text, when the content of the text changes, the generated images before and after will be inconsistent, resulting in poor consistency of the images. Summary of the Invention
[0005] The present invention provides an image generation method, device, equipment, and storage medium, which are used to solve the defect that the consistency of the images generated before and after is not high when the text content changes in the prior art, and to improve the consistency of the images generated before and after.
[0006] The present invention provides an image generation method, including:
[0007] Input the current text and the previous image corresponding to the previous text into the autoregressive module in the image generation model to obtain the current token sequence output by the autoregressive module;
[0008] Input the previous image into the image feature adaptation module in the image generation model to obtain the image features output by the image feature adaptation module;
[0009] Based on the current token sequence and the image features, determine the current image corresponding to the current text.
[0010] According to the image generation method provided by the present invention, the determining the current image corresponding to the current text based on the current token sequence and the image features includes:
[0011] Input the current token sequence and the image features into the decoupled cross-attention module in the image generation model to obtain the text attention features and image attention features output by the decoupled cross-attention module;
[0012] Input the text attention features and the image attention features into the decoder in the image generation model to obtain the current image output by the decoder.
[0013] According to an image generation method provided by the present invention, the decoder includes a diffusion model.
[0014] According to an image generation method provided by the present invention, the decoder is trained in the following manner:
[0015] Input a sample image into an image feature extraction module to obtain sample image features output by the image feature extraction module;
[0016] Input the sample image features into an initial decoder to obtain a first predicted image;
[0017] Based on the first predicted image and the sample image, iteratively optimize the model parameters of the initial decoder to obtain the decoder.
[0018] According to an image generation method provided by the present invention, the decoupled cross-attention module is trained in the following manner:
[0019] Input the sample image features of the sample image and the sample text token sequence corresponding to the sample text into an initial decoupled cross-attention module to obtain the first sample text attention features and the first sample image attention features output by the initial decoupled cross-attention module;
[0020] Input the first sample text attention features and the first sample image attention features into the decoder to obtain a second predicted image corresponding to the sample text output by the decoder;
[0021] Based on the second predicted image and the sample image, iteratively optimize the model parameters of the initial decoupled cross-attention module to obtain the decoupled cross-attention module.
[0022] According to an image generation method provided by the present invention, the autoregressive module is trained in the following manner:
[0023] Input the sample text and the sample image into an initial autoregressive module to obtain a sample token sequence output by the initial autoregressive module;
[0024] Input the sample token sequence and the sample image features into the decoupled cross-attention module to obtain the second sample text attention features and the second sample image attention features output by the decoupled cross-attention module;
[0025] Input the second sample text attention features and the second sample image attention features into the decoder to obtain the third predicted image corresponding to the sample text output by the decoder;
[0026] Iteratively optimize the model parameters of the initial autoregressive module based on the third predicted image and the sample image to obtain the autoregressive module.
[0027] The present invention also provides an image generation device, including:
[0028] An input module, configured to input the current text and the previous image corresponding to the previous text into the autoregressive module in the image generation model to obtain the current token sequence output by the autoregressive module;
[0029] The input module is further configured to input the previous image into the image feature adaptation module in the image generation model to obtain the image features output by the image feature adaptation module;
[0030] A determination module, configured to determine the current image corresponding to the current text based on the current token sequence and the image features.
[0031] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the image generation method described in any one of the above is implemented.
[0032] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the image generation method described in any one of the above is implemented.
[0033] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the image generation method described in any one of the above is implemented.
[0034] The image generation method, device, equipment, and storage medium provided by the present invention input the current text and the previous image corresponding to the previous text into the autoregressive module in the image generation model to obtain the current token sequence output by the autoregressive module, and input the previous image into the image feature adaptation module in the image generation model. After obtaining the image features output by the image feature adaptation module, based on the current token sequence and the image features, the current image corresponding to the current text is determined. Since when generating the current image, it is not only based on the current text, but is guided by the previous image, and the image features of the previous image are used as a guide to generate the current image, the generated current image can be made more similar to the previous image. Therefore, the consistency between the generated front and back images can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0036] Figure 1 It is a schematic flowchart of the image generation method provided by an embodiment of the present invention.
[0037] Figure 2 It is a schematic structural diagram of the autoregressive network provided by an embodiment of the present invention.
[0038] Figure 3 It is one of the schematic structural diagrams of the image generation model provided by an embodiment of the present invention.
[0039] Figure 4 It is the second schematic structural diagram of the image generation model provided by an embodiment of the present invention.
[0040] Figure 5 It is a schematic structural diagram of the image generation device provided by an embodiment of the present invention.
[0041] Figure 6 It is a schematic physical structure diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0043] With the development of Artificial Intelligence (AI) technology, the application of text-to-image has become increasingly widespread. Among them, picture book generation is a practical application scenario of text-to-image, which is mainly aimed at children's education and has extensive demands in fields such as learning machines or children's book publishing. Picture books concretely present a children's story through cartoon drawings. Picture books often need to depict a complete story through a series of consecutive images. Therefore, it is necessary to ensure the consistency of the image of the story characters, the consistency of the scenes, and the consistency of the image styles, etc. Therefore, how to generate images with feature consistency is very important.
[0044] In the prior art, usually, the text is input into a diffusion model to generate images, and the consistency of the generated images before and after is ensured by editing the content of the text. For example, the description of the character image is kept consistent so that the characters in the generated images are as similar as possible in appearance, and stylized descriptions are made in the text so that the generated image styles are as close as possible. In addition, the random number sampling is kept consistent so that the content of the images before and after is as close as possible. However, although the consistency is maintained in the text description, due to the continuous change and development of the description of the picture book story content, and the diffusion model is highly sensitive to the text, the consistency of some texts is easily diluted as a whole, making the consistency of the generated images before and after uncontrollable, thus resulting in inconsistent images generated before and after.
[0045] The embodiment of the present invention takes into account the above problems and provides an image generation method. In this method, in addition to the current text, the previous image generated according to the previous text can also be input into the autoregressive module in the image generation model to obtain the current token sequence, and the previous image is input into the image feature adaptation module in the image generation model to obtain the image features output by the image feature adaptation module. Thus, the image features corresponding to the previous image are used as a guide, and the current image corresponding to the current text is generated based on the current token sequence. Since the previous image is used as a guide when generating the current image, the image features of the previous image will be used as a standard to generate the current image. Therefore, the consistency between the generated images before and after can be improved.
[0046] The following is combined with Figures 1 to 4The image generation method provided by the embodiments of the present invention will be described. The execution subject of this method can be an electronic device such as a terminal device, a computer, or a server, or a smart device specifically designed for image generation, or an image generation device provided in the electronic device or smart device. The image generation device can be implemented by software, hardware, or a combination of both. This method can be applied to any scenario that requires text-to-image generation, especially in scenarios that require maintaining consistency in image style before and after, such as picture book generation, game design, or film and television production.
[0047] Figure 1 It is a schematic flowchart of the image generation method provided by the embodiments of the present invention. As Figure 1 shown, the method includes:
[0048] Step 101: Input the current text and the previous image corresponding to the previous text into the autoregressive module in the image generation model to obtain the current token sequence output by the autoregressive module.
[0049] In this step, in the text-to-image scenario, usually a piece of text will correspond to at least one image. The current text can be understood as a piece of text for which an image needs to be generated currently, and the previous text is the text based on which the previous image was generated. It should be understood that the moment when the current image needs to be generated can be determined as the current moment, and the moment when the previous image was generated can be determined as the previous moment. Therefore, the previous image corresponding to the previous text can also be understood as the image generated at the previous moment.
[0050] The autoregressive module can adopt Emu2 or other autoregressive networks. Figure 2 It is a schematic structural diagram of the autoregressive network provided by the embodiments of the present invention. As Figure 2 shown, the autoregressive network includes an encoder Encoder, a generative multimodal model, and a decoder Decoder. Among them, the encoder is an image encoding network based on Contrastive Language-Image Pre-training (CLIP). After inputting the current text and the previous image corresponding to the previous text into the encoder of the autoregressive module, the encoder will perform encoding processing on the features of the current text and the previous image, and input the encoded features into the generative multimodal model. Through an autoregressive manner, a current token sequence containing semantic information output by the autoregressive module is obtained. This current token sequence can be used as the input of the decoder Decoder. The decoder Decoder can be an SDXL network or other networks that can decode tokens. The decoder Decoder decodes the tokens based on classification information.
[0051] Step 102: Input the previous image into the image feature adaptation module in the image generation model to obtain the image features output by the image feature adaptation module.
[0052] In this step, the image feature adaptation module can be used to extract the image features in the previous image to guide the generation of the current image subsequently, so that the newly generated current image can maintain feature consistency with the previous image.
[0053] Specifically, after inputting the previous image into the image feature adaptation module, the feature information in the previous image is extracted through the image encoder in the image feature adaptation module. These feature information can be used to represent image content or image style, etc. Then, it is transformed through the linear transformation layer in the image feature adaptation module to combine the feature vectors output by the image encoder, achieving the purpose of adjusting the feature dimension. In addition, the linearly transformed features can be input into the normalization layer for normalization processing, thereby obtaining the image features of the previous image.
[0054] Step 103: Determine the current image corresponding to the current text based on the current token sequence and the image features.
[0055] In this step, after determining the current token sequence and the image features, based on the current token sequence and the image features, with the image features as the guide, the current image corresponding to the current text is generated by decoding the current token sequence. The image features are used as a separate hint. Through the decoupled cross-attention method, a layer specifically for processing the image features is set in each cross-attention layer. In this way, the image features and the text features corresponding to the current token sequence can be processed through their respective cross-attention layers, avoiding information loss that may be caused by direct merging.
[0056] The image generation method provided by the embodiments of the present invention inputs the current text and the previous image corresponding to the previous text into the autoregressive module in the image generation model to obtain the current token sequence output by the autoregressive module, and inputs the previous image into the image feature adaptation module in the image generation model. After obtaining the image features output by the image feature adaptation module, the current image corresponding to the current text is determined based on the current token sequence and the image features. Since when generating the current image, it is not only based on the current text, but also guided by the previous image, and the image features of the previous image are used as a guide to generate the current image, the generated current image can be made more similar to the previous image. Therefore, the consistency between the generated front and back images can be improved.
[0057] In the picture book application scenario, the autoregressive model can endow the picture book generation process with temporal characteristics. In addition to text information, the previous image generated in the previous scenario can also be used as guiding information for the subsequent scenario, making the front and back scenes and stories depicted in the picture book continuous, compensating for the lack of text information in the current scenario, and enriching the details of the current scenario. Moreover, the Image Prompt Adapter (IP-Adapter) is used to enhance the consistency of the story character images before and after. The IP-Adapter extracts the visual characteristics of the character as guiding information and guides the generation of the current image together with the text information, thereby generating the current image containing the guiding information, further ensuring the consistency of the current image and the previous image in terms of characters and content.
[0058] Exemplarily, based on the above embodiments, when determining the current image corresponding to the current text based on the current token sequence and image features, the following method can be adopted:
[0059] Input the current token sequence and image features into the decoupled cross-attention module in the image generation model to obtain the text attention features and image attention features output by the decoupled cross-attention module, and input the text attention features and image attention features into the decoder in the image generation model to obtain the current image output by the decoder.
[0060] Specifically, Figure 3 is one of the structural schematic diagrams of the image generation model provided by the embodiments of the present invention. As Figure 3 shown, the image generation model includes an autoregressive module, an image feature adaptation module, and a decoder. Among them, input the current text "The kitten found a cave" and the previous image corresponding to the previous text into the autoregressive module, and the current token sequence can be obtained. Among them, the text corresponding to the previous image is "A kitten wearing a hat is walking in the forest". Input the previous image into the image feature adaptation module IP-Adapter to obtain image features, and input the current token sequence and image features into the decoder to obtain the current image corresponding to the current text. Among them, Figure 3 t = n in represents the current moment, and t = n - 1 represents the moment when the previous image was generated.
[0061] Figure 4 is the second structural schematic diagram of the image generation model provided by the embodiments of the present invention. As Figure 4 shown, input the current text "A little girl wearing glasses" and the previous image corresponding to the previous text into the autoregressive network, and thus the current token sequence output by the autoregressive network can be obtained.
[0062] In addition, the previous image corresponding to the previous text is input into the image feature adaptation module, and through an image decoder, linear transformation, and normalization processing, the image features of the previous image are obtained. The obtained current token sequence and the image features of the previous image can also be input into the decoupled cross-attention module of the image generation model. Among them, the decoupled cross-attention module contains two cross-attention modules, which process image features and text features respectively, so as to obtain the text attention features and image attention features output by each cross-attention module. Since the image attention features and text attention features are processed through their respective cross-attention modules, the image generation model can independently focus on the important information in the image and text.
[0063] Furthermore, the obtained text attention features and image attention features are input into the decoder in the image generation model to obtain the current image output by the decoder. Among them, the decoder can adopt the U-Net model of the diffusion model, such as SDXL, and input the noise x at time t t , the text attention features and the image attention features into the decoder U-Net model together. Through the way of gradually denoising, in each denoising process, the denoising U-Net model uses the image features and text features to guide the generation of the image, so as to obtain the final clear image. It should be understood that what is obtained after the denoising U-Net model decodes is the latent of the image, and the final current image can also be obtained through the pre-trained Variational Autoencoder (VAE).
[0064] Since the decoder is set as a diffusion model and the final current image is generated by gradually denoising, in each denoising process, the generation process of the image can be controlled more precisely, thereby improving the image quality of the current image.
[0065] Among them, the above-mentioned decoupled cross-attention module can be set in the image feature adaptation module, can also be set in the decoder, or can be set separately.
[0066] It should be noted that at the initial moment, that is, when generating the image for the first time, since there is no previous image corresponding to the previous text, only the text content is input into the autoregressive module. The autoregressive module generates image tokens, and the image tokens are input into the decoder for decoding to generate the corresponding image. At this time, the input of the image feature adaptation module IP-Adapter is empty. At any subsequent moment, the current text and the image tokens corresponding to the previous image can be input into the autoregressive module to obtain the tokens at the current moment, and the previous image is used as the standard input image of the role image feature adaptation module IP-Adapter for feature extraction, so as to guide the decoder to generate the current image with consistent features.
[0067] In addition, the role image can be customized, and any role image can be input into the image feature adaptation module IP-Adapter as the guiding image, so that the role image in the finally generated current image can be controlled.
[0068] In this embodiment, the image features of the previous image are separately extracted as a kind of prompt feature. By setting the decoupled cross-attention module and setting a separate text cross-attention module and image cross-attention module in the decoupled cross-attention module, the text attention features and image attention features are distinguished. Therefore, the image generation model can independently process image features and text features, reduce information loss, and thus improve the correlation between the image and the text description. In the picture book application scenario, the consistency of the picture book content can be maintained in the above manner. In addition, by generating the current token sequence through the autoregressive model and generating the current image through the current token sequence, the consistency of the generated image scene and image style can be ensured. The consistency of the role image can be ensured through the image feature adaptation module, and the role image can be customized. Moreover, when generating images through the image generation method in the present invention, the production efficiency can be improved for the offline picture book generation scenario, and the generation effect and quality can be improved for the online picture book generation scenario.
[0069] Exemplarily, on the basis of the above embodiments, the training process of each module in the image generation model will be introduced in detail below.
[0070] Among them, the decoder is trained in the following manner:
[0071] The sample image is input into the image feature extraction module to obtain the sample image features output by the image feature extraction module. The sample image features are input into the initial decoder to obtain the first predicted image. Based on the first predicted image and the sample image, the model parameters of the initial decoder are iteratively optimized to obtain the decoder.
[0072] Specifically, multiple sample images are collected by a camera device or obtained from an image database, and each sample image is input into an image feature extraction module, such as a Convolutional Neural Network (CNN), to obtain the sample image features output by the image feature extraction module. Then, the extracted sample image features are input into an initial decoder U-Net, and the initial decoder U-Net decodes the sample image features to obtain the first predicted image output by the initial decoder. Based on the difference between the obtained first predicted image and the sample image, loss information is determined, and thus the model parameters of the initial decoder are adjusted according to this loss information. By iteratively executing the above process until the loss is minimized or the obtained model converges, the finally obtained model is determined as the decoder.
[0073] In this embodiment, since the model parameters of the initial decoder can be iteratively optimized based on the first predicted image and the sample image for the training of the decoder, when generating the current image based on the finally obtained decoder, the generated current image can be made to be consistent with the input previous image in style, improving the consistency between the current image and the previously generated images.
[0074] Among them, the decoupled cross-attention module is trained based on the following method:
[0075] The sample image features of the sample image and the sample text token sequence corresponding to the sample text are input into the initial decoupled cross-attention module to obtain the first sample text attention feature and the first sample image attention feature output by the initial decoupled cross-attention module. The first sample text attention feature and the first sample image attention feature are input into the decoder to obtain the second predicted image corresponding to the sample text output by the decoder. Based on the second predicted image and the sample image, the model parameters of the initial decoupled cross-attention module are iteratively optimized to obtain the decoupled cross-attention module.
[0076] Specifically, the sample text is text related to the text used to generate the sample image, such as two adjacent paragraphs in the same story, etc. The sample text is input into the text encoder to obtain the sample text token sequence output by the text encoder. The sample image features extracted from each sample image and the sample text token sequence are input into the initial decoupled cross-attention module. The initial decoupled cross-attention module processes the image features and text features separately, such as processing the image and text through different cross-attention modules respectively, to obtain the first sample text attention feature and the first sample image attention feature output by the initial decoupled cross-attention module. By processing the image and text separately, the interference between different modality features can be reduced, enabling the model to focus more on the key information of each modality.
[0077] Further, the obtained first sample text attention feature and first sample image attention feature can be input into a pre-trained decoder. The decoder decodes the first sample text attention feature and first sample image attention feature to obtain a second predicted image corresponding to the sample text. The loss information is determined based on the difference between the second predicted image and the sample image, and the model parameters of the initial decoupled cross-attention module are iteratively optimized based on the loss information, so as to obtain a finally trained decoupled cross-attention module to minimize the difference between the second predicted image and the sample image. In this way, when using the decoupled cross-attention module to generate the current image subsequently, the difference between the generated current image and the previous image can be reduced.
[0078] In this embodiment, the initial decoupled cross-attention module processes the image feature and text feature separately, and trains the initial decoupled cross-attention module based on the obtained first sample text attention feature and first sample image attention feature. In this way, the trained decoupled cross-attention module can achieve stronger generalization ability when facing new text or images. In addition, by decoupling the feature of the sample text and the feature of the sample image, information loss can be reduced and more key information can be retained, thereby improving the accuracy of the trained decoupled cross-attention module.
[0079] Among them, the autoregressive module is trained based on the following method:
[0080] Input the sample text and the sample image into the initial autoregressive module to obtain a sample token sequence output by the initial autoregressive module, and input the sample token sequence and the sample image feature into the decoupled cross-attention module to obtain the second sample text attention feature and the second sample image attention feature output by the decoupled cross-attention module. After inputting the second sample text attention feature and the second sample image attention feature into the decoder and obtaining the third predicted image corresponding to the sample text output by the decoder, the model parameters of the initial autoregressive module are iteratively optimized based on the third predicted image and the sample image to obtain the autoregressive module.
[0081] Specifically, after the decoder and the decoupled cross-attention module are trained, the autoregressive module can be further trained. The initial autoregressive module is usually a Transformer-based model. After inputting the sample text and the sample image into the initial autoregressive module, a corresponding sample token sequence can be generated, and this sample token sequence is the encoded representation of the sample text.
[0082] The obtained sample token sequence and sample image features are input into the trained decoupled cross-attention module. The decoupled cross-attention module processes the image features and text features separately to generate the second sample text attention features and the second sample image attention features. Processing the image features and text features separately can reduce the interference between features of different modalities, enabling the model to focus more on the key information of each modality.
[0083] The second sample text attention features and the second sample image attention features are input into the trained decoder. The decoder generates the third predicted image corresponding to the sample text based on these features. Based on the difference between the third predicted image and the sample image, loss information is determined, and the model parameters of the initial autoregressive module are iteratively optimized according to this loss information to obtain the autoregressive module. Since the difference between the third predicted image and the sample image is used as the loss information, the difference between the third predicted image and the sample image can be minimized.
[0084] In this embodiment, by inputting the sample text and the sample image into the initial autoregressive module, then inputting the generated sample token sequence and sample image features into the decoupled cross-attention module, and then inputting the second sample text attention features and the second sample image attention features output by the decoupled cross-attention module into the decoder, and finally iteratively optimizing the autoregressive module based on the difference between the obtained third predicted image and the sample image, the finally obtained autoregressive module can generate an image that more accurately matches the text description, improving the quality of image generation.
[0085] The image generation device provided by the present invention will be described below. The image generation device described below can be correspondingly referred to the image generation method described above.
[0086] Figure 5 It is a schematic structural diagram of the image generation device provided by the embodiment of the present invention. Referring to Figure 5 As shown, the image generation device 500 includes:
[0087] An input module 11, configured to input the current text and the previous image corresponding to the previous text into the autoregressive module in the image generation model to obtain the current token sequence output by the autoregressive module;
[0088] The input module 11 is further configured to input the previous image into the image feature adaptation module in the image generation model to obtain the image features output by the image feature adaptation module;
[0089] A determination module 12, configured to determine the current image corresponding to the current text based on the current token sequence and the image features.
[0090] In an exemplary embodiment, the determination module 12 is specifically configured to:
[0091] Input the current token sequence and the image features into the decoupled cross-attention module in the image generation model to obtain the text attention features and image attention features output by the decoupled cross-attention module;
[0092] Input the text attention features and the image attention features into the decoder in the image generation model to obtain the current image output by the decoder.
[0093] In an exemplary embodiment, the decoder includes a diffusion model.
[0094] In an exemplary embodiment, the decoder is trained in the following manner:
[0095] Input a sample image into the image feature extraction module to obtain the sample image features output by the image feature extraction module;
[0096] Input the sample image features into the initial decoder to obtain a first predicted image;
[0097] Iteratively optimize the model parameters of the initial decoder based on the first predicted image and the sample image to obtain the decoder.
[0098] In an exemplary embodiment, the decoupled cross-attention module is trained in the following manner:
[0099] Input the sample image features of the sample image and the sample text token sequence corresponding to the sample text into the initial decoupled cross-attention module to obtain the first sample text attention features and the first sample image attention features output by the initial decoupled cross-attention module;
[0100] Input the first sample text attention features and the first sample image attention features into the decoder to obtain the second predicted image corresponding to the sample text output by the decoder;
[0101] Iteratively optimize the model parameters of the initial decoupled cross-attention module based on the second predicted image and the sample image to obtain the decoupled cross-attention module.
[0102] In an exemplary embodiment, the autoregressive module is trained in the following manner:
[0103] Input the sample text and the sample image into the initial autoregressive module to obtain the sample token sequence output by the initial autoregressive module;
[0104] Input the sample token sequence and the sample image features into the decoupled cross-attention module to obtain the second sample text attention features and the second sample image attention features output by the decoupled cross-attention module;
[0105] Input the second sample text attention features and the second sample image attention features into the decoder to obtain the third predicted image corresponding to the sample text output by the decoder;
[0106] Based on the third predicted image and the sample image, iteratively optimize the model parameters of the initial autoregressive module to obtain the autoregressive module.
[0107] The device in this embodiment can be used to execute the method of any one of the method embodiments on the image generation method side. Its specific implementation process and technical effects are similar to those in the method embodiments on the image generation method side. For details, please refer to the detailed introduction in the method embodiments on the image generation method side, which will not be elaborated here.
[0108] Figure 6 It is a schematic physical structure diagram of an electronic device provided by an embodiment of the present invention. As Figure 6 shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 complete mutual communication through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute an image generation method, which includes: inputting the current text and the previous image corresponding to the previous text into the autoregressive module in the image generation model to obtain the current token sequence output by the autoregressive module; inputting the previous image into the image feature adaptation module in the image generation model to obtain the image features output by the image feature adaptation module; and determining the current image corresponding to the current text based on the current token sequence and the image features.
[0109] In addition, when the logical instructions in the above-mentioned memory 630 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0110] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the image generation method provided by the above-mentioned various methods. The method includes: inputting the current text and the previous image corresponding to the previous text into an autoregressive module in an image generation model to obtain the current token sequence output by the autoregressive module; inputting the previous image into an image feature adaptation module in the image generation model to obtain the image features output by the image feature adaptation module; and determining the current image corresponding to the current text based on the current token sequence and the image features.
[0111] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the image generation method provided by the above-mentioned various methods. The method includes: inputting the current text and the previous image corresponding to the previous text into an autoregressive module in an image generation model to obtain the current token sequence output by the autoregressive module; inputting the previous image into an image feature adaptation module in the image generation model to obtain the image features output by the image feature adaptation module; and determining the current image corresponding to the current text based on the current token sequence and the image features.
[0112] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0113] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image generation method, characterized in that, Including: Input the current text and the previous image corresponding to the previous text into the autoregressive module in the image generation model to obtain the current token sequence containing semantic information output by the autoregressive module; the previous text and the current text are consecutive texts; Input the previous image into the image feature adaptation module in the image generation model to obtain the image features output by the image feature adaptation module; the image features are used to characterize the image style of the previous image; Based on the current token sequence and the image features, with the image style as a guide, decode the current token sequence to generate the current image corresponding to the current text; the current image and the previous image have the same style; The determining of the current image corresponding to the current text based on the current token sequence and the image features includes: Input the current token sequence and the image features into the decoupled cross-attention module in the image generation model to obtain the text attention features and image attention features output by the decoupled cross-attention module; the text attention features and image attention features are processed by different cross-attention modules; Input the text attention features and the image attention features into the decoder in the image generation model to obtain the current image output by the decoder.
2. The image generation method according to claim 1, wherein The decoder includes a diffusion model.
3. The image generation method according to claim 1 or 2, wherein The decoder is trained based on the following method: Input the sample image into the image feature extraction module to obtain the sample image features output by the image feature extraction module; Input the sample image features into the initial decoder to obtain the first predicted image; Based on the first predicted image and the sample image, iteratively optimize the model parameters of the initial decoder to obtain the decoder.
4. The image generation method according to claim 3, wherein The decoupled cross-attention module is trained based on the following method: Input the sample image features of the sample image and the sample text token sequence corresponding to the sample text into the initial decoupled cross-attention module to obtain the first sample text attention features and the first sample image attention features output by the initial decoupled cross-attention module; Input the first sample text attention features and the first sample image attention features into the decoder to obtain the second predicted image corresponding to the sample text output by the decoder; Based on the second predicted image and the sample image, iteratively optimize the model parameters of the initial decoupled cross-attention module to obtain the decoupled cross-attention module.
5. The image generation method according to claim 4, characterized in that The autoregressive module is trained based on the following method: Input the sample text and the sample image into the initial autoregressive module to obtain the sample token sequence output by the initial autoregressive module; Input the sample token sequence and the sample image features into the decoupled cross-attention module to obtain the second sample text attention features and the second sample image attention features output by the decoupled cross-attention module; Input the second sample text attention feature and the second sample image attention feature into the decoder to obtain a third predicted image corresponding to the sample text output by the decoder; Iteratively optimize the model parameters of the initial autoregressive module based on the third predicted image and the sample image to obtain the autoregressive module.
6. An image generation device, characterized in that, Comprising: An input module, configured to input a current text and a previous image corresponding to a previous text into an autoregressive module in an image generation model to obtain a current token sequence containing semantic information output by the autoregressive module; the previous text and the current text are consecutive texts; The input module is further configured to input the previous image into an image feature adaptation module in the image generation model to obtain an image feature output by the image feature adaptation module; the image feature is used to characterize the image style of the previous image; A determination module, configured to decode the current token sequence based on the current token sequence and the image feature, with the image style as a guide, to generate a current image corresponding to the current text; the current image and the previous image have the same style; Specifically, the determination module is configured to: Input the current token sequence and the image feature into a decoupled cross-attention module in the image generation model to obtain a text attention feature and an image attention feature output by the decoupled cross-attention module; the text attention feature and the image attention feature are processed by different cross-attention modules; Input the text attention feature and the image attention feature into a decoder in the image generation model to obtain the current image output by the decoder.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, the image generation method according to any one of claims 1 to 5 is implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the image generation method according to any one of claims 1 to 5 is implemented.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the image generation method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Image generation method and related device
CN118537447A