Text Image Generation Method, Apparatus, Terminal, and Computer-Readable Storage Medium

By integrating high-dimensional text and image representations, the method addresses the issue of inconsistent style and meaning in text-to-image generation, ensuring coherent and consistent image series.

CN115438210BActive Publication Date: 2025-07-15INT DIGITAL ECONOMY ACAD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210912106.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-29
Publication Date
2025-07-15
Estimated Expiration
2042-07-29

AI Technical Summary

Technical Problem

In the prior art, when generating a series of images from long text, there is a lack of style consistency and semantic consistency between images, resulting in a problem of context mismatch.

Method used

By displaying multiple candidate text fragments, receiving user's text fragment determination instructions, extracting the current text representation using the pre-trained text generation model, and inputting the image generation model with historical image representation, generating the current candidate image corresponding to the stitched text, and establishing a multimodal generation model to ensure consistency.

Benefits of technology

It realizes the generation of a series of consistent style and semantic images based on long text, solves the problem of inconsistency in image generation in the prior art, and is suitable for scenarios such as children's picture books and commercial displays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115438210B_ABST
    Figure CN115438210B_ABST
Patent Text Reader

Abstract

The text image generation method, device, terminal, and computer-readable storage medium provided by the present invention, the text image generation method includes: displaying a plurality of candidate text segments, receiving a text segment determination instruction from a user, and determining a current text segment among the plurality of candidate text segments; inputting the current text segment into a pre-trained text generation model, and extracting a current text representation corresponding to the text after splicing the current text segment with all historical text segments; obtaining historical image representations corresponding to all historical images, and inputting the current text representation and the historical image representations into a pre-trained image generation model to generate a current candidate image corresponding to the spliced text. By extracting the current text representation of the spliced text and inputting the current text representation and the historical image representations corresponding to all historical images into the image generation model together, the present invention enables consistency between each image, and further enables a long text to match a series of consistent images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method, device, terminal and computer-readable storage medium for text image generation. Background Art

[0002] With the development of the technology of pre-trained language models (Pretrained Language Model), the model has begun to be able to generate credible and logically consistent long texts. Moreover, with the continuous exploration and innovation of generation technologies, diverse text generation solutions have gradually become possible.

[0003] With the development of pre-trained model technologies for multi-modal (such as text plus image), multiple technologies have emerged that allow users to generate high-definition and credible images according to text descriptions, but are limited to image generation corresponding to a single text. That is to say, image generation models often are limited to single-sentence generation and lack modeling of context images and texts. Therefore, when an image generation model generates a series of images according to a long text, there may be a lack of style consistency and semantic consistency, where the long text includes multiple text segments, and each text segment can correspond to generate an image. For example, in a fairy tale, the style of the picture suddenly switches from an oil painting tone to a realistic tone, resulting in a lack of style consistency in the series of images generated from this long text. Another example is that the story background is in a fairy tale castle, but when describing the dining scene, the picture is of a modern restaurant dining scene, resulting in a lack of semantic consistency in the series of images generated from this long text.

[0004] Therefore, the existing technology has defects and needs to be improved and developed. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method, device, terminal and computer-readable storage medium for text image generation in view of the above-mentioned defects of the existing technology, aiming to solve the problem of the lack of consistency in context semantics between the images generated corresponding to the context when generating a series of images according to a long text in the existing technology.

[0006] The technical solution adopted by the present invention to solve the technical problem is as follows:

[0007] A method for text image generation, comprising:

[0008] Displaying a plurality of candidate text segments, receiving a user's text segment determination instruction, and determining a current text segment among the plurality of candidate text segments;

[0009] Inputting the current text segment into a pre-trained text generation model, and extracting a current text representation corresponding to the text after splicing the current text segment and all historical text segments;

[0010] Obtain the historical image representations corresponding to all historical images, and input the current text representation and the historical image representations into a pre-trained image generation model to generate a current candidate image corresponding to the concatenated text.

[0011] In one implementation, the candidate text segment is generated by the text generation model based on all historical text segments.

[0012] In one implementation, the historical text segments include the initial text segment input by the user;

[0013] Input the initial text segment into the text generation model to generate multiple candidate text segments;

[0014] Receive the user's text segment determination instruction;

[0015] Determine a first text segment from the multiple candidate text segments according to the text segment determination instruction, and extract the first text representation corresponding to the text after concatenating the initial text segment and the first text segment.

[0016] In one implementation, after determining the first text segment from the multiple candidate text segments according to the text segment determination instruction and extracting the first text representation corresponding to the text after concatenating the initial text segment and the first text segment, it further includes:

[0017] Input the first text representation and the blank image representation into the image generation model to generate multiple first candidate images corresponding to the first text segment;

[0018] Receive the user's first image selection instruction, and determine a first image from the multiple first candidate images according to the first image selection instruction.

[0019] In one implementation, the historical text segments include: the initial text segment, and the text segment selected by the user from the multiple candidate texts generated by the text generation model;

[0020] The displaying of the multiple candidate text segments, receiving the user's text segment determination instruction, and determining the current text segment from the multiple candidate text segments includes:

[0021] Generate multiple current candidate text segments in the text generation model based on all historical text segments and display them;

[0022] Receive the user's text segment determination instruction, and determine the current text segment from the multiple candidate text segments.

[0023] In one implementation, the current text segment is input into a pre-trained text generation model, and a current text representation corresponding to the text obtained by concatenating the current text segment with all historical text segments is extracted, including:

[0024] The current text segment is input into a pre-trained text generation model, and the current text representation is jointly extracted according to the sentence vector corresponding to the current text segment and the sentence vectors corresponding to each of the historical text segments;

[0025] Among them, when extracting the sentence vectors corresponding to each text segment, long-range dependencies are modeled, and mapping operations are performed on each sentence vector.

[0026] In one implementation, the historical image includes the image selected by the user from multiple candidate images generated by the image generation model;

[0027] After obtaining the historical image representations corresponding to all historical images, inputting the current text representation and the historical image representations into a pre-trained image generation model to generate a current candidate image corresponding to the concatenated text, further including:

[0028] Receiving a current image determination instruction from the user, and determining a current image among the multiple current candidate images;

[0029] Extracting the current image representations corresponding to the current image and all historical images in the image generation model.

[0030] In one implementation, the training steps of the text generation model and the image generation model include:

[0031] Obtaining a text training set, and training an initial text generation model using the text training set;

[0032] Obtaining a pre-established text-image training set, and training the initial text generation model and the initial image generation model using the text-image training set; the text-image training set includes text segment training data, image training data, and the corresponding relationship between the text segment training data and the image training data;

[0033] After the training is completed, a trained text generation model and an image generation model are obtained.

[0034] In one implementation, the obtaining of the text training set and training the initial text generation model using the text training set includes:

[0035] Obtaining a text training set, and modeling the maximum likelihood probability of generating text segments in the text training set using an autoregressive language model or a non-autoregressive language model.

[0036] In one implementation, when using an autoregressive language model, the calculation formula for the maximum likelihood probability is ;

[0037] wherein, the represents a text segment of length T.

[0038] In one implementation, training the initial text generation model and the initial image generation model using the text image training set includes:

[0039] Obtaining text segment training data in the text image training set;

[0040] Inputting the text segment training data into the initial text generation model to generate a first text training segment;

[0041] Extracting a first training text representation corresponding to the text after splicing the text segment training data and the first text training segment;

[0042] Inputting the first training text representation and a blank picture representation into the initial image generation model to generate a first training image.

[0043] In one implementation, after inputting the first training text representation and a blank picture representation into the initial image generation model to generate a first training image, it further includes:

[0044] Generating a current text training segment according to all historical text training segments, where the historical text training segments include: the text segment training data, and the text training segments generated in the initial text generation model;

[0045] Extracting a current text training representation corresponding to the text after splicing the current text training segment and the historical text training segments;

[0046] Obtaining historical training image representations corresponding to all historical training images, inputting the current text training representation and the historical training image representations into the initial image generation model to generate a current training image; the historical training images include the images generated in the initial image generation model.

[0047] In one implementation, the generating of the current text training segment according to all historical text training segments further includes:

[0048] Comparing the current text training segment with the corresponding target text segment training data in the text image training set, and maximizing the likelihood probability of generating the target text segment training data to calculate the loss of the text generation model.

[0049] In one implementation, obtaining the historical training image representations corresponding to all historical training images, and inputting the current text training representation and the historical training image representations into the initial image generation model to generate a current training image further includes:

[0050] Searching for target image training data corresponding to the target text segment training data, comparing the current training image with the target image training data, and maximizing the likelihood probability of generating the target image training data to calculate the loss of the image generation model.

[0051] The present invention also provides a text-image generation device, including:

[0052] A display module for displaying a plurality of candidate text segments and a plurality of candidate images;

[0053] A determination module for receiving a user's text segment determination instruction to determine a current text segment among the plurality of candidate text segments, and also for receiving a user's image determination instruction to determine a current image among the plurality of candidate images;

[0054] A text feature extraction module for inputting the current text segment into a pre-trained text generation model to generate a plurality of candidate text segments, and extracting a current text representation corresponding to the text obtained by splicing the current text segment and all historical text segments;

[0055] An image generation module for obtaining historical image representations corresponding to all historical images, and inputting the current text representation and the historical image representations into a pre-trained image generation model to generate a current candidate image corresponding to the spliced text.

[0056] The present invention also provides a terminal, including: a memory, a processor, and a text-image generation program stored on the memory and executable on the processor. When the text-image generation program is executed by the processor, the steps of the text-image generation method described above are implemented.

[0057] The present invention also provides a computer-readable storage medium storing a computer program that can be executed to implement the steps of the text-image generation method described above.

[0058] The text image generation method, device, terminal, and computer-readable storage medium provided by the present invention. The text image generation method includes: displaying a plurality of candidate text segments, receiving a text segment determination instruction from a user, and determining a current text segment among the plurality of candidate text segments; inputting the current text segment into a pre-trained text generation model, and extracting a current text representation corresponding to the text obtained by splicing the current text segment and all historical text segments; obtaining historical image representations corresponding to all historical images, and inputting the current text representation and the historical image representations into a pre-trained image generation model to generate a current candidate image corresponding to the spliced text. By extracting the current text representation of the text obtained by splicing the current text segment and all historical text segments, and inputting the current text representation and the historical image representations corresponding to all historical images into the image generation model together, the present invention makes each image have consistency, so that a long text can match a series of consistent images. Description of the Drawings

[0059] Figure 1 It is a flowchart of a preferred embodiment of the text image generation method in the present invention.

[0060] Figure 2 It is a schematic diagram of the text segment input interface in the present invention.

[0061] Figure 3 It is a specific flowchart of step S100 in a preferred embodiment of the text image generation method in the present invention.

[0062] Figure 4 It is a flowchart after step S300 in a preferred embodiment of the text image generation method in the present invention.

[0063] Figure 5 It is a schematic diagram of the text generation model and the image generation model in the present invention.

[0064] Figure 6 It is a specific flowchart of pre-training the text generation model and the image generation model in the present invention.

[0065] Figure 7 It is a specific flowchart of step S20 in a preferred embodiment of the text image generation method in the present invention.

[0066] Figure 8 It is a schematic diagram of the algorithm principle of the multi-modal generation model in the present invention.

[0067] Figure 9 It is a schematic diagram of the principle of the multi-modal generation model in the present invention.

[0068] Figure 10 It is a functional principle block diagram of a preferred embodiment of the text image generation device in the present invention.

[0069] Figure 11 It is a functional principle block diagram of a preferred embodiment of the terminal in the present invention. Specific Embodiments

[0070] To make the objectives, technical solutions and advantages of the present invention clearer and more explicit, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.

[0071] The present invention provides a user with an interaction process with a large-scale pre-trained language model through an intuitive graphical interface, and finally obtains a series of digital images corresponding to the text. In this process, the user can continuously generate text or stories and freely determine the specific writing direction of the text according to the possible options returned by the model or the user's input. As the text direction selected by the user changes, after each interaction update of the model, the model comprehensively captures the key semantic and syntactic and other abstract structural information in the time series based on the long text generated previously, and generates a high-dimensional vector representation of the potential information. This representation, together with the historical images, is input into a multi-modal text-to-image deep learning model, which guides the process of generating images with consistent context style and coherent meaning, and realizes the generation of strongly consistent images according to dynamic text.

[0072] Please refer to Figure 1 , Figure 1 which is a flowchart of the text image generation method in the present invention. As Figure 1 shown, the text image generation method described in the embodiments of the present invention includes the following steps:

[0073] Step S100: Display a plurality of candidate text segments, receive a text segment determination instruction from the user, and determine a current text segment among the plurality of candidate text segments.

[0074] That is to say, the present invention realizes the interaction with the user when generating images from text, allows the user to select the current text segment by himself / herself, and customizes the generation of a series of images with strong consistency according to the current text segment. The candidate text segments are generated by the text generation model according to all historical text segments.

[0075] In one implementation, the historical text segments include the initial text segments input by the user. The text image generation method further includes: inputting the initial text segments into the text generation model to generate a plurality of candidate text segments; receiving a text segment determination instruction from the user; determining a first text segment among the plurality of candidate text segments according to the text segment determination instruction, and extracting a first text representation corresponding to the text after splicing the initial text segments and the first text segment.

[0076] Specifically, the present invention involves a user inputting an initial text segment, which can be a sentence or one or several keywords, as Figure 2 shown. The text generation model can automatically generate subsequent text based on the initial text segment, that is, generate consecutive text segments. When generating the first text segment, the text generation model provides multiple candidate text segments, and the user can select the first text segment by themselves, thus enabling the user to freely determine the specific direction of the text writing. Then, jointly extract the first text representation of the text after splicing the initial text segment and the first text segment. The first text representation is a joint high-dimensional abstract feature, which plays a decisive role in guiding the generation of consistent styles of images.

[0077] In one embodiment, after determining the instruction according to the text segment to determine the first text segment among multiple candidate text segments and extracting the first text representation corresponding to the text after splicing the initial text segment and the first text segment, it further includes: inputting the first text representation and a blank image representation into the image generation model to generate multiple first candidate images corresponding to the first text segment; receiving the user's first image selection instruction, and determining the first image among multiple first candidate images according to the first image selection instruction. That is, in the first round of the loop, without a historical image representation, a blank image representation and the first text representation are used to input into the image generation model together.

[0078] In one embodiment, the historical text segment includes: the initial text segment and the text segments selected by the user among multiple candidates generated by the text generation model. As Figure 3 shown, the step S100 specifically includes:

[0079] Step S110: Generate multiple current candidate text segments according to all historical text segments in the text generation model and display them;

[0080] Step S120: Receive the user's text segment determination instruction and determine the current text segment among multiple candidate text segments.

[0081] Specifically, the historical text segment refers to the text segments selected by the user among the text segments generated since the input of the initial text segment, and also includes the initial text segment. Generating a corresponding image for each text segment is a cycle, and the user can generate a long text through multiple cycles. Therefore, in the current cycle, all text segments generated in the previous cycles and the initial text segment will be considered to improve the consistency of generating images for each text segment. Multiple current candidate text segments are all consistent with the previous text and the user can determine the final text segment by themselves.

[0082] After the step S100, the following steps are: step S200, input the current text segment into a pre-trained text generation model, and extract the current text representation corresponding to the text after splicing the current text segment and all historical text segments. That is to say, in each loop, the joint high-dimensional abstract features are extracted from the previous text.

[0083] In one embodiment, the current text segment is input into a pre-trained text generation model, and the current text representation is jointly extracted according to the sentence vector corresponding to the current text segment and the sentence vectors corresponding to each historical text segment; wherein, when extracting the sentence vectors corresponding to each text segment, long-term dependencies are modeled, and mapping operations are performed on each sentence vector.

[0084] Specifically, common autoregressive language models model the long-term dependencies between words. However, in the present invention, the extraction of sentence vectors between sentences models long-term dependencies, and mapping operations are performed on the sentence vectors, so that what it focuses on is more high-dimensional abstract semantic information. While general language models do not model the connections between sentence representations and do not map them into high-dimensional latent space vectors.

[0085] After the step S200, the following steps are: step S300, obtain the historical image representations corresponding to all historical images, and input the current text representation and the historical image representations into a pre-trained image generation model to generate the current candidate image corresponding to the spliced text.

[0086] The historical image representation is the high-dimensional feature jointly extracted by the image generation model for all historical images.

[0087] In one embodiment, the historical images include the images selected by the user from multiple candidate images generated by the image generation model. As Figure 4 shown, after the step S300, the following steps are further included:

[0088] Step S410, receive the current image determination instruction of the user, and determine the current image from the multiple current candidate images;

[0089] Step S420, extract the current image representations corresponding to the current image and all historical images in the image generation model.

[0090] It can be understood that the current image representation extracted in the current loop is the historical image representation in the next loop.

[0091] As Figure 5As shown, the text generation model generates multiple sentence candidates based on the sentence 0 input by the user. After the user selects sentence 1 from the multiple sentence candidates, the text generation model jointly extracts the high-dimensional text representation 1 based on the two sentence vectors of sentence 0 and 1. Together with the image representation 0 (i.e., the zero vector), it is input into the image generation model. The image generation model generates multiple candidate images. After the user selects image 1, the image generation model maps the image features into a high-dimensional image representation 1, and a single loop ends.

[0092] In a new round of loop, the text generation model generates multiple sentence candidates based on the historical sentences 0 and 1. After the user selects sentence 2, the text generation model extracts the joint high-dimensional abstract features of sentences 0, 1, and 2 to obtain the text representation 2, which is input into the right image generation model together with the image representation 1 of the historical image 1 extracted by the image generation model in the previous loop to generate image 2. That is to say, when generating text fragments and images in each loop, the generated representations all consider the high-dimensional common features extracted from multiple past sentences or images, and this feature plays a decisive role in guiding the consistent style of image generation. This loop will continue to repeat until the user stops selecting the next text fragment.

[0093] In this way, through multiple loops of the present invention, long texts are formed among the text fragments, and a series of digital images corresponding to the long texts are obtained. That is to say, the present invention allows users to interactively generate long sequential texts and corresponding series of images, which has broad application prospects and potential in scenarios such as children's picture books and commercial display production.

[0094] In one implementation, as Figure 6 shown, the training steps of the text generation model and the image generation model include:

[0095] Step S10: Obtain a text training set and train an initial text generation model using the text training set;

[0096] Step S20: Obtain a pre-established text-image training set and train the initial text generation model and the initial image generation model using the text-image training set;

[0097] Step S30: After the training is completed, obtain the trained text generation model and image generation model;

[0098] Among them, the text-image training set includes text fragment training data, image training data, and the corresponding relationship between the text fragment training data and the image training data.

[0099] That is to say, the present invention first trains an initial text generation model with a text training set, where the text training set only contains text. Optional architectures of the initial text generation model include the Transformer architecture or the recurrent neural network (RNN). In addition to being able to normally generate smooth sentences, it needs to have additional architectures such as the long short-term memory model (LSTM) or the attention mechanism in the recurrent neural network (RNN), etc., which can model the extraction of common abstract features of the context and ensure that it applies controlled text generation technology, and can provide the ability to interact with users to generate text.

[0100] Then, train a multi-modal generation model that receives the high-dimensional representation of the text and the high-dimensional abstract representation of the historical image to generate high-definition images. The multi-modal generation model includes an initial text generation model trained with a text-image training set and an untrained initial image generation model, and the two are jointly trained to align the data representations between different modalities. The goal during model training is to maximize the probability of generating the training data.

[0101] In one implementation, step S10 specifically includes: obtaining a text training set and using an autoregressive language model or a non-autoregressive language model to model the maximum likelihood probability of generating text segments in the text training set.

[0102] In one embodiment, a text training set is obtained. When using an autoregressive language model, the calculation formula for the maximum likelihood probability is ; the represents a text segment of length T.

[0103] Specifically, the initial text generation model adopts a text-to-text training method, and both the training data input and output are text. A common autoregressive language model can be used to model the maximum likelihood probability of sentence generation. The meaning of the calculation formula for the maximum likelihood probability is that when the language model predicts the next word, the probability of the next word is conditional on all previous words. Specifically, when the language model generates sentences, if it generates one by one from left to right, it is called autoregressive generation. When generating each word, it is necessary to consider the content generated previously, which is called conditional on previous words. However, if non-autoregressive generation is used, it may be disordered generation (such as insertion) or all the words in a sentence are optimized and generated simultaneously, then the maximum likelihood probability is no longer conditional on previous words.

[0104] In one implementation, as Figure 7 shown, "using the text-image training set to train the initial text generation model and the initial image generation model" in step S20 specifically includes:

[0105] Step S21: Obtain the text segment training data in the text image training set;

[0106] Step S22: Input the text segment training data into the initial text generation model to generate a first text training segment;

[0107] Step S23: Extract the first training text representation corresponding to the text after splicing the text segment training data and the first text training segment;

[0108] Step S24: Input the first training text representation and the blank picture representation into the initial image generation model to generate a first training image.

[0109] In one embodiment, after step S24, the following steps are further included:

[0110] Step S25: Generate the current text training segment according to all historical text training segments;

[0111] The historical text training segments include: the text segment training data, and the text training segments generated in the initial text generation model;

[0112] Step S26: Extract the current text training representation corresponding to the text after splicing the current text training segment and the historical text training segments;

[0113] Step S27: Obtain the historical training image representations corresponding to all historical training images, input the current text training representation and the historical training image representations into the initial image generation model to generate the current training image;

[0114] The historical training images include the images generated in the initial image generation model.

[0115] Specifically, first, crawl a large amount of text-image consistency data. For example, collect and organize text-image data pairs from sources such as picture books, commercial displays, or introductory videos with subtitles for the hearing-impaired, and store the text-image data pairs as a text image training set. When organizing, the supervised training of a neural network generally includes an input and a desired output. The text-image pairs crawled in a specific context have characteristics such as consistent style and semantics. The text image training set can be one or more long text data segments. A long text data segment includes multiple text segment training data, and each text segment training data corresponds to image training data, and the images are consistent in style and semantics. For example, a long text includes many sentences, each sentence can be used as text segment training data, and each sentence corresponds to an image, and the images of the context are consistent in style and semantics.

[0116] For example, sentence 0 can be the first sentence of a long text data segment in the text image training set, or it can be other sentences. The initial text generation model generates sentence 1 based on sentence 0. After jointly extracting the high-dimensional text representation 1 from the two sentence vectors of sentences 0 and 1, it is input into the initial image generation model together with the image representation 0 (i.e., the zero vector). The image generation model generates image 1 and maps the image features of image 1 into a high-dimensional image representation 1. After the initial text generation model generates sentence 2 based on the historical sentences 0 and 1, the initial text generation model extracts the joint high-dimensional abstract features of sentences 0, 1, and 2 to obtain the text representation 2, which is input into the initial image generation model together with the image representation 1 of the historical image 1 extracted by the initial image generation model to generate image 2. The subsequent steps are continuously repeated in a loop until the long text data segment ends.

[0117] In one embodiment, the step S25 further includes: comparing the current text training segment with the corresponding target text segment training data in the text image training set, and maximizing the likelihood probability of generating the target text segment training data to calculate the loss of the text generation model. That is, when generating a text training segment in each loop, the initial text generation model compares with the text segment training data in the text image training set to maximize the likelihood probability of generating the text segment training data.

[0118] In one implementation manner, the step S27 further includes: searching for the target image training data corresponding to the target text segment training data, comparing the current training image with the target image training data, and maximizing the likelihood probability of generating the target image training data to calculate the loss of the image generation model. That is, when generating a training image in each loop, the image generation model compares with the image training data in the text image training set to maximize the likelihood probability of generating the image training data. That is, during training, the model is updated by calculating the likelihood probability deviation of the data. When the user uses it, each time a text segment or an image is generated, multiple candidate text segments or candidate images are generated through sampling control for the user to select.

[0119] There are multiple possible specific implementations of the multimodal generation model. One of them is the diffusion generation model, which models a process of predicting a continuous probability distribution from a disordered Gaussian distribution based on the mechanism of thermodynamics, and continuously guides the probability distribution prediction with the text input information until the distribution fits the semantics of the text generation model. As Figure 8 and Figure 9 shown, the training process is actually randomly sampling the noisy image at the t-th step, inputting the noisy picture and the step number t, and the model predicts the noise , Model training objective: The smaller the error between the predicted noise and the actually added noise, the better. The sampling process (generation process) is as follows: Subtract the noise predicted by the model from the noisy image (the first image is Gaussian distributed noise sampled randomly) (the other parameters in front of the noise can be deduced backward from the above noise addition process), and continuously remove the noise to restore the original image. Among them, the training of the image diffusion model samples an image X0 from the true data distribution q and fits an error estimate in its diffusion process to gradually add noise. When applying the trained model for inference, a noisy image is sampled from Gaussian noise, and noise is continuously removed to obtain a clear image, and this image is required to follow the true distribution q. The models in the prior art only consider single-step semantic modeling and do not take context into account in text representation. In image generation, they also do not consider inputting the high-dimensional image representation of historical images (i.e., non-random Gaussian distributed noise images) into the model. Therefore, the present invention extracts abstract high-dimensional representations based on context text and historical images to guide image generation. Compared with the traditional technical solutions that do not consider previous text and images, it better ensures the consistency between the generated images, and solves the problems that the semantics of the long texts generated by the current text generation models are inconsistent before and after, and the styles of the images generated according to the text mutate suddenly before and after.

[0120] The text generation model and the image generation model pre-trained by the present invention jointly constitute a multi-modal generation model, which can generate an abstract representation of the text based on context modeling, so that not only the context of the current sentence is considered, but also the information of the previous sentences, such as abstract information like style and atmosphere, can be retained and represented; inputting this abstract representation into the multi-modal generation model, during the process of image generation, through the continuous guidance of the abstract representation, the corresponding images generated by each sentence have the same style, coherent meaning, and are not fragmented; the unique language model of the model that generates potential information representations using context sentences and the architecture that generates new images (including multiple candidate images) using historical image and text representations are the guarantees for the long-sequence model to effectively maintain consistency.

[0121] Furthermore, as Figure 10 shown, based on the above text image generation method, the present invention also correspondingly provides a text image generation device, including:

[0122] A display module 100, configured to display a plurality of candidate text segments and a plurality of candidate images;

[0123] A determination module 200, configured to receive a user's text segment determination instruction and determine the current text segment among the plurality of candidate text segments, and is also configured to receive a user's image determination instruction and determine the current image among the plurality of candidate images;

[0124] The text feature extraction module 300 is configured to input the current text segment into a pre-trained text generation model to generate multiple candidate text segments, and extract the current text representation corresponding to the text after splicing the current text segment and all historical text segments.

[0125] The image generation module 400 is configured to obtain the historical image representations corresponding to all historical images, and input the current text representation and the historical image representations into a pre-trained image generation model to generate a current candidate image corresponding to the spliced text.

[0126] Further, as Figure 11 shown, based on the above text image generation method, the present invention also correspondingly provides a terminal, including a processor 10 and a memory 20. Figure 11 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0127] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as the hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk equipped on the terminal, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 20 may also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software installed on the terminal and various types of data, such as program codes installed on the terminal. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a text image generation program 30 is stored on the memory 20, and the text image generation program 30 can be executed by the processor 10 to implement the text image generation method in the present application.

[0128] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chips, and is configured to run the program codes stored in the memory 20 or process data, such as executing the text image generation method, etc.

[0129] In one embodiment, when the processor 10 executes the text image generation program 30 in the memory 20, the following steps are implemented:

[0130] Display multiple candidate text segments, receive a text segment determination instruction from the user, and determine the current text segment among the multiple candidate text segments;

[0131] Input the current text segment into a pre-trained text generation model, and extract the current text representation corresponding to the text after concatenating the current text segment with all historical text segments;

[0132] Obtain the historical image representations corresponding to all historical images, and input the current text representation and the historical image representations into a pre-trained image generation model to generate the current image corresponding to the concatenated text.

[0133] The candidate text segment is generated by the text generation model based on all historical text segments.

[0134] The historical text segments include the initial text segment input by the user;

[0135] Input the initial text segment into the text generation model to generate multiple candidate text segments;

[0136] Receive the user's text segment determination instruction;

[0137] Determine the first text segment among the multiple candidate text segments according to the text segment determination instruction, and extract the first text representation corresponding to the text after concatenating the initial text segment and the first text segment.

[0138] After determining the first text segment among the multiple candidate text segments according to the text segment determination instruction and extracting the first text representation corresponding to the text after concatenating the initial text segment and the first text segment, it further includes:

[0139] Input the first text representation and the blank image representation into the image generation model to generate multiple first candidate images corresponding to the first text segment;

[0140] Receive the user's first image selection instruction, and determine the first image among the multiple first candidate images according to the first image selection instruction.

[0141] The historical text segments include: the initial text segment, and the text segment selected by the user from the multiple candidate texts generated by the text generation model;

[0142] Displaying multiple candidate text segments, receiving the user's text segment determination instruction, and determining the current text segment among the multiple candidate text segments includes:

[0143] Generate multiple current candidate text segments in the text generation model according to all historical text segments and display them;

[0144] Receive the user's text segment determination instruction, and determine the current text segment among the multiple candidate text segments.

[0145] Input the current text segment into a pre-trained text generation model, and extract the current text representation corresponding to the text obtained by concatenating the current text segment with all historical text segments, including:

[0146] Input the current text segment into a pre-trained text generation model, and jointly extract the current text representation based on the sentence vector corresponding to the current text segment and the sentence vectors corresponding to each historical text segment;

[0147] Among them, when extracting the sentence vectors corresponding to each text segment, long-range dependencies are modeled, and mapping operations are performed on each sentence vector.

[0148] The historical image includes the image selected by the user from multiple candidate images generated by the image generation model;

[0149] After obtaining the historical image representations corresponding to all historical images, inputting the current text representation and the historical image representations into a pre-trained image generation model to generate the current candidate image corresponding to the concatenated text, including:

[0150] Receive the user's current image determination instruction, and determine the current image among the multiple current candidate images;

[0151] Extract the current image representations corresponding to the current image and all historical images in the image generation model.

[0152] The training steps of the text generation model and the image generation model include:

[0153] Obtain a text training set, and use the text training set to train an initial text generation model;

[0154] Obtain a pre-established text-image training set, and use the text-image training set to train the initial text generation model and the initial image generation model; the text-image training set includes text segment training data, image training data, and the corresponding relationship between the text segment training data and the image training data;

[0155] After the training is completed, obtain the trained text generation model and image generation model.

[0156] The obtaining of the text training set and using the text training set to train the initial text generation model includes:

[0157] Obtain a text training set, and use an autoregressive language model or a non-autoregressive language model to model the maximum likelihood probability of generating text segments in the text training set.

[0158] When using an autoregressive language model, the calculation formula for the maximum likelihood probability is ;

[0159] Among them, the represents a text segment of length T.

[0160] Training the initial text generation model and the initial image generation model using the text image training set includes:

[0161] Obtaining text segment training data in the text image training set;

[0162] Inputting the text segment training data into the initial text generation model to generate a first text training segment;

[0163] Extracting a first training text representation corresponding to the text after splicing the text segment training data and the first text training segment;

[0164] Inputting the first training text representation and the blank picture representation into the initial image generation model to generate a first training image.

[0165] After inputting the first training text representation and the blank picture representation into the initial image generation model to generate a first training image, it further includes:

[0166] Generating a current text training segment according to all historical text training segments, where the historical text training segments include: the text segment training data and the text training segments generated in the initial text generation model;

[0167] Extracting a current text training representation corresponding to the text after splicing the current text training segment and the historical text training segments;

[0168] Obtaining historical training image representations corresponding to all historical training images, inputting the current text training representation and the historical training image representations into the initial image generation model to generate a current training image; the historical training images include the images generated in the initial image generation model.

[0169] The generating the current text training segment according to all historical text training segments further includes:

[0170] Comparing the current text training segment with the corresponding target text segment training data in the text image training set, and maximizing the likelihood probability of generating the target text segment training data to calculate the loss of the text generation model.

[0171] The obtaining historical training image representations corresponding to all historical training images, inputting the current text training representation and the historical training image representations into the initial image generation model to generate a current training image further includes:

[0172] Search for target image training data corresponding to the training data of the target text segment, compare the current training image with the target image training data, and maximize the likelihood probability of generating the target image training data to calculate the loss of the image generation model.

[0173] The present invention also provides a computer-readable storage medium storing a computer program that can be executed to implement the steps of the text image generation method as described above.

[0174] In summary, for the text image generation method, device, terminal, and computer-readable storage medium disclosed by the present invention, the text image generation method includes: displaying a plurality of candidate text segments, receiving a text segment determination instruction from a user, and determining a current text segment from the plurality of candidate text segments; inputting the current text segment into a pre-trained text generation model, and extracting a current text representation corresponding to the text after splicing the current text segment and all historical text segments; obtaining historical image representations corresponding to all historical images, and inputting the current text representation and the historical image representations into a pre-trained image generation model to generate a current candidate image corresponding to the spliced text. By extracting the current text representation of the text after splicing the current text segment and all historical text segments, and inputting the current text representation and the historical image representations corresponding to all historical images into the image generation model together, the present invention makes each image consistent, and further enables a long text to match a series of consistent images.

[0175] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or changes can be made according to the above description, and all such improvements and changes should fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for generating a text image, characterized in that, including: displaying a plurality of candidate text segments, receiving a text segment determination instruction from a user, and determining a current text segment among the plurality of candidate text segments; inputting the current text segment into a pre-trained text generation model, and extracting a current text representation corresponding to the text obtained by concatenating the current text segment and all historical text segments; obtaining historical image representations corresponding to all historical images, inputting the current text representation and the historical image representations into a pre-trained image generation model, and generating a current candidate image corresponding to the concatenated text; inputting the current text segment into a pre-trained text generation model, and extracting a current text representation corresponding to the text obtained by concatenating the current text segment and all historical text segments, including: inputting the current text segment into a pre-trained text generation model, and jointly extracting the current text representation according to the sentence vector corresponding to the current text segment and the sentence vectors corresponding to each of the historical text segments; wherein, when extracting the sentence vectors corresponding to the respective text segments, long-term order dependencies are modeled, and mapping operations are performed on the respective sentence vectors.

2. The text image generation method according to claim 1, characterized in that, The candidate text segments are generated by the text generation model according to all historical text segments.

3. The text image generation method according to claim 1, wherein The historical text segments include an initial text segment input by the user; inputting the initial text segment into the text generation model to generate a plurality of candidate text segments; receiving a text segment determination instruction from the user; determining a first text segment among the plurality of candidate text segments according to the text segment determination instruction, and extracting a first text representation corresponding to the text obtained by concatenating the initial text segment and the first text segment.

4. The text image generation method according to claim 3, characterized in that After determining the first text segment among the plurality of candidate text segments according to the text segment determination instruction and extracting the first text representation corresponding to the text obtained by concatenating the initial text segment and the first text segment, it further includes: inputting the first text representation and a blank image representation into the image generation model to generate a plurality of first candidate images corresponding to the first text segment; receiving a first image selection instruction from the user, and determining a first image among the plurality of first candidate images according to the first image selection instruction.

5. The text image generation method according to claim 3, wherein The historical text segments include: the initial text segment, and text segments selected by the user from among the plurality of candidates generated by the text generation model; The displaying a plurality of candidate text segments, receiving a text segment determination instruction from a user, and determining a current text segment among the plurality of candidate text segments includes: generating a plurality of current candidate text segments in the text generation model according to all historical text segments and displaying them; receiving a text segment determination instruction from the user, and determining a current text segment among the plurality of candidate text segments.

6. The text image generation method according to claim 1, characterized in that, The historical images include images selected by the user from among the plurality of candidate images generated by the image generation model; After obtaining historical image representations corresponding to all historical images, inputting the current text representation and the historical image representations into a pre-trained image generation model, and generating a current candidate image corresponding to the concatenated text, it further includes: receiving a current image determination instruction from the user, and determining a current image among the plurality of current candidate images; Extract the current image representations corresponding to the current image and all historical images in the image generation model.

7. The text image generation method according to claim 1, characterized in that The training steps of the text generation model and the image generation model include: Obtain a text training set, and use the text training set to train an initial text generation model. Obtain a pre-established text-image training set, and use the text-image training set to train the initial text generation model and the initial image generation model; the text-image training set includes text segment training data, image training data, and the corresponding relationship between the text segment training data and the image training data. After the training is completed, obtain the trained text generation model and image generation model.

8. The text image generation method according to claim 7, characterized in that, The obtaining of the text training set and using the text training set to train the initial text generation model includes: Obtain a text training set, and use an autoregressive language model or a non-autoregressive language model to model the maximum likelihood probability of text segment generation in the text training set.

9. The text image generation method according to claim 8, wherein When an autoregressive language model is adopted, the calculation formula of the maximum likelihood probability is ; Among them, the represents a text segment of length T.

10. The text image generation method according to claim 7, wherein Using the text-image training set to train the initial text generation model and the initial image generation model includes: Obtain the text segment training data in the text-image training set. Input the text segment training data into the initial text generation model to generate a first text training segment. Extract the first training text representation corresponding to the text after splicing the text segment training data and the first text training segment. Input the first training text representation and a blank image representation into the initial image generation model to generate a first training image.

11. The text image generation method according to claim 10, wherein After inputting the first training text representation and a blank image representation into the initial image generation model to generate a first training image, it further includes: Generate a current text training segment according to all historical text training segments, where the historical text training segments include: the text segment training data and the text training segments generated in the initial text generation model. Extract the current text training representation corresponding to the text after splicing the current text training segment and the historical text training segments. Obtain the historical training image representations corresponding to all historical training images, and input the current text training representation and the historical training image representations into the initial image generation model to generate a current training image; the historical training images include the images generated in the initial image generation model.

12. The text image generation method according to claim 11, wherein The generating of the current text training segment according to all historical text training segments further includes: Compare the current text training segment with the corresponding target text segment training data in the text-image training set, and maximize the likelihood probability of generating the target text segment training data to calculate the loss of the text generation model.

13. The text image generation method according to claim 12, wherein The obtaining of the historical training image representations corresponding to all historical training images, and inputting the current text training representation and the historical training image representations into the initial image generation model to generate a current training image further includes: Search for target image training data corresponding to the target text segment training data, compare the current training image with the target image training data, and maximize the likelihood probability of generating the target image training data to calculate the loss of the image generation model.

14. A text image generation device, characterized in that, Including: A display module for displaying a plurality of candidate text segments and a plurality of candidate images; A determination module for receiving a user's text segment determination instruction to determine the current text segment among the plurality of candidate text segments, and also for receiving a user's image determination instruction to determine the current image among the plurality of candidate images; A text feature extraction module for inputting the current text segment into a pre-trained text generation model to generate a plurality of candidate text segments, and extracting the current text representation corresponding to the text after splicing the current text segment and all historical text segments; An image generation module for obtaining historical image representations corresponding to all historical images, inputting the current text representation and the historical image representations into a pre-trained image generation model, and generating a current candidate image corresponding to the spliced text; Inputting the current text segment into a pre-trained text generation model, and extracting the current text representation corresponding to the text after splicing the current text segment and all historical text segments, including: Inputting the current text segment into a pre-trained text generation model, and jointly extracting the current text representation according to the sentence vector corresponding to the current text segment and the sentence vectors corresponding to each of the historical text segments; Wherein, when extracting the sentence vectors corresponding to each text segment, long-term order dependencies are modeled, and mapping operations are performed on each sentence vector.

15. A terminal, characterized in that, Including: A memory, a processor, and a text image generation program stored on the memory and executable on the processor. When the text image generation program is executed by the processor, the steps of the text image generation method according to any one of claims 1 to 13 are implemented.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the text image generation method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Visual storyline generation from text story

    US20200257763A1