Method and apparatus for generating image based on text, electronic device and storage medium
By obtaining initial information, using a text generation model to expand visual effect information and generate target text information, the low-quality problem caused by users' inability to accurately describe images is solved, and high-quality image generation and improved user experience are achieved.
Patent Information
- Application Number
- CN202410282379.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-12
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-03-12
AI Technical Summary
Since users are unable to provide accurate image description vocabulary, the images generated by the text-based graph model are of low quality and cannot meet user needs.
By obtaining initial information, the pre-trained text generation model is used to expand and generate image category information and visual effect information, the target text information is determined, and the target image is generated by inputting the image generation model.
The quality and efficiency of image generation are improved, the application difficulty is reduced, the user experience is enhanced, and the generated images are highly matched with user needs.
Smart Images

Figure CN118037896B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of text expansion, the technical field of image generation, and in particular to a method and device for generating an image based on text, an electronic device, and a storage medium. BACKGROUND
[0002] In recent years, due to the fact that images generated by text-to-image models are comparable to images created by human authors, text-to-image models have made a great contribution in the field of AIGC (Artificial Intelligence Generated Content). Because text-to-image models can directly generate images corresponding to the semantics of a text based on the text, text-to-image models are widely used in various industries. For example, in e-commerce, text-to-image models can be used to generate product images based on text descriptions of the products. In virtual reality, text-to-image models can be used to directly generate vivid and artistic virtual environment images based on text descriptions of the virtual environment.
[0003] In related technologies, because text-to-image models generate corresponding images based on text descriptions, the accuracy and content richness of the images generated by text-to-image models are closely related to the text used to describe the images. However, in actual applications, due to factors such as the user's expertise, users often cannot provide accurate text expressions for describing the images they need to generate, which results in low-quality images that cannot meet the user's needs. SUMMARY
[0004] To solve the above problems, the present disclosure provides a method and device for generating an image based on text, an electronic device, and a storage medium.
[0005] In one aspect of the present disclosure, a method for generating an image based on text is provided, which includes: obtaining initial information of an image to be generated, the initial information including image category information; inputting the initial information into a pre-trained text generation model to obtain at least one text description information, the text description information including the image category information and image effect information; determining target text information based on the at least one text description information; and inputting the target text information into a pre-trained image generation model to obtain at least one target image.
[0006] In another aspect of the embodiments of the present disclosure, an apparatus for generating an image based on text is provided, which comprises: a first obtaining module configured to obtain initial information of an image to be generated, the initial information comprising image category information; a text generating module configured to input the initial information into a pre-trained text generation model to obtain at least one text description information, the text description information comprising the image category information and image effect information; a text determining module configured to determine target text information based on the at least one text description information; and an image generating module configured to input the target text information into a pre-trained image generation model to obtain at least one target image.
[0007] In yet another aspect of the embodiments of the present disclosure, an electronic device is provided, which comprises: a memory configured to store a computer program; and a processor configured to execute the computer program stored in the memory, and when the computer program is executed, a method for generating an image based on text is implemented.
[0008] In still another aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, which stores a computer program, and when the computer program is executed by a processor, a method for generating an image based on text is implemented.
[0009] In the embodiments of the present disclosure, initial information of an image to be generated (including image category information) can be obtained first, then the initial information is expanded by a text generation model to generate at least one text description information, each text description information comprising the image category information and expanded visual effect information, then target text information is determined based on the at least one text description information, and then at least one target image is generated based on the target text information by an image generation model.
[0010] Therefore, only the initial information of the image to be generated needs to be provided by a user, and the text generation model can expand and generate at least one visual effect information of the image category corresponding to the image category information based on the image category information, which has a high professional degree and accurately describes the image category to be generated, so that the image generation model can generate at least one target image of the corresponding category with the corresponding visual effect based on the target text information, and the target image with different visual effects and rich content can be generated through the target text information comprising different visual effect information, so as to allow the user to select the desired image, effectively improve the quality and efficiency of the generated image, and reduce the application difficulty of the text generation technology.
[0011] In addition, since the image category information is included in the target text information, the image generated based on the text description information has a high matching degree with the user demand, which greatly improves the user experience.
[0012] The technical solutions of the present disclosure will be further described in detail below with reference to the drawings and embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0013] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0014] The present disclosure embodiments can be understood more readily by reference to the following detailed description, taken in connection with the accompanying drawings, and wherein:
[0015] Figure 1 is a flowchart of a method for generating an image based on text according to an example embodiment of the present disclosure;
[0016] Figure 2 is a flowchart of step S140 according to an example embodiment of the present disclosure;
[0017] Figure 3 is a structural diagram of a text generation model according to an example embodiment of the present disclosure;
[0018] Figure 4 is a flowchart of step S120 according to an example embodiment of the present disclosure;
[0019] Figure 5 is a flowchart of a method for generating an image based on text according to another example embodiment of the present disclosure;
[0020] Figure 6 is a flowchart of step S260 according to an example embodiment of the present disclosure;
[0021] Figure 7 is a diagram of an image generated by a text generation model according to an example embodiment of the present disclosure;
[0022] Figure 8 is a diagram of an image generated by a text generation model according to an example embodiment of the present disclosure;
[0023] Figure 9 is a structural diagram of an apparatus for generating an image based on text according to an example embodiment of the present disclosure;
[0024] Figure 10 is a structural diagram of an electronic device according to an example embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] Various example embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. Note that the relative arrangement, numerical expressions, and numerical values of components and steps set forth in these embodiments are not limiting to the scope of the present disclosure, unless otherwise specified.
[0026] Those skilled in the art can understand that the terms "first", "second" and the like in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, and do not represent any specific technical meaning, nor represent the inevitable logical sequence between them.
[0027] It should also be understood that in the embodiments of the present disclosure, "multiple" can mean two or more, and "at least one" can mean one, two or more.
[0028] It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, it can be understood as one or more in general, without explicit limitation or in the context of the opposite indication.
[0029] In addition, the term "and / or" in the present disclosure is only a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in the present disclosure generally represents that the front and rear associated objects are in an "or" relationship.
[0030] It should also be understood that the description of the embodiments of the present disclosure emphasizes the differences between the embodiments, and the same or similar parts can be referred to each other, and for the sake of brevity, will not be repeated.
[0031] At the same time, it should be understood that, for the convenience of description, the size of each part shown in the drawings is not drawn according to the actual proportional relationship.
[0032] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application or uses.
[0033] The techniques, methods and devices known to those skilled in the relevant art can not be discussed in detail, but in appropriate cases, the techniques, methods and devices should be considered as part of the specification.
[0034] It should be noted that: similar numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further discussed in the subsequent drawings.
[0035] The disclosed embodiments can be applied to terminal devices, computer systems, servers, and other electronic devices, which can operate with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with terminal devices, computer systems, servers, and other electronic devices include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments that include any of the above systems, and the like.
[0036] Terminal devices, computer systems, servers, and other electronic devices can be described in the general context of computer system-executable instructions, such as program modules, being executed by a computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, and the like, which perform particular tasks or implement particular abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, in which tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules can be located in local or remote computer system storage media, including storage devices.
[0037] In the process of implementing the present disclosure, the inventors found that before generating an image using a text-to-image model, a user needs to first create words for describing the image, and then input the words for describing the image into the text-to-image model, and the text-to-image model generates a corresponding image according to the input words for describing the image. However, due to factors such as user expertise, the user often cannot provide accurate words for describing the image, which leads to the image content generated by the text-to-image model being simple and far from the image required by the user. For example, a user wants to generate a CG (Computer Graphics) image about an airplane, and needs to input words for describing the theme (airplane), style, resolution, brightness, shooting angle, light, and the like of the image into the text-to-image model. However, in actual application, the user usually only inputs “airplane” into the text-to-image model, which leads to the image content generated by the text-to-image model being simple and far from the image required by the user.
[0038] Figure 1 is a flowchart of a method for generating an image based on text provided by an exemplary embodiment of the present disclosure. The present embodiment can be applied to an electronic device, such as a terminal device, a computer system, a server, and the like. Figure 1 As shown in FIG. 1, the method includes the following steps:
[0039] In step S110, initial information of an image to be generated is obtained.
[0040] The initial information includes image category information. The image category information is used to indicate a subject object required to be included in the generated image. For example, the image category information can include at least one of the following: horse, airplane, car, dog, cat, ship, deer, truck, person, etc. For example, if the initial information includes image category information of "car", the generated image includes the object "car".
[0041] In an embodiment, the initial information can further include other information, which can be used to describe or limit the image category information, i.e., the other information can describe the details of the subject corresponding to the image category information. For example, the other information can include information of color, shape, classification, action, etc. of the image category information, and can further include information of image resolution, image style, image tone, etc. For example, the initial information can include: cat (image category information), tiger stripes (other information), yellow and white (other information), running (other information), 4K (other information), science fiction style (other information), etc.
[0042] In an embodiment, the initial information can be obtained in the following way:
[0043] The original information describing the generated image can be obtained, which can be video information, voice information or text information, etc. The original information includes image category information. When the original information is text information, the original information is determined as the initial information. When the original information is video information or text information, the original information is converted into text format, and the original information in text format is determined as the initial information.
[0044] In step S120, the initial information is input into a pre-trained text generation model to obtain at least one text description information.
[0045] The text generation model can include but is not limited to: Auto-Regressive language model, Auto-Encoder language model, Encoder-Decoder language model. For example, the text generation model can be a GPT2 model in the Auto-Regressive language model, or a Seq2Seq model in the Encoder-Decoder language model, etc.
[0046] The text generation model can also be other types of language models, for example, the text generation model can be a ChatGLM model, a Large Language Model Meta AI (LLaMA) model, or a BLOOM model, etc. The text description information includes image category information and visual effect information. The visual effect information is used to describe the visual effect of the image to be generated. That is, the visual effect information is used to describe or limit the image category information, that is, it can describe the relevant details of the subject object (image category information) corresponding to the image category information. For example, the visual effect information can include information such as the resolution, style, size, brightness, and expression environment of the image. For example, assuming that the image category information is an airplane, the visual effect information can include 4K (resolution), science fiction (style), 400 cd / m 2 (brightness), and flying in the Milky Way (image expression environment).
[0047] Step S130, determining target text information based on at least one text description information.
[0048] In one specific implementation, when the text description information is one, the text description information can be determined as the target text information, and when the text description information is multiple, at least one text description information can be selected as the target text description information according to the content of the multiple text description information.
[0049] In another specific implementation, the user can also perform re-editing processing on the text description information to obtain the target text information. For example, one text description information can be selected from the at least one text description information as an initial text description information, and then the initial text description information can be re-edited by artificial or pre-trained text generation model to obtain the target text information.
[0050] Step S140, inputting the target text information into a pre-trained image generation model to obtain at least one target image.
[0051] The image generation model can be a text-to-image model. For example, the image generation model can be implemented by using a Stable Diffusion model, a Diffusion model, a GAN (Generative Adversarial Networks), etc.
[0052] In one specific implementation, the target text information is input into the image generation model, and at least one target image corresponding to the target text information is generated by the image generation model. The number of target images can be one or more. The target image corresponding to the target text information conforms to the semantics of the target text information. The target image corresponds to the to-be-generated image.
[0053] For example, assuming that the target text information includes: an airplane (image category information), 4K (visual effect information), science fiction style (visual effect information), 400 cd / m 2 2 2
[0054] In the embodiments of the present disclosure, the initial information of the to-be-generated image (including image category information) can be obtained first, and then the initial information is expanded by using the text generation model to generate at least one text description information, each text description information including the image category information and the expanded visual effect information, and then the target text information is determined based on the at least one text description information, and the image generation model generates at least one target image based on the target text information.
[0055] Therefore, only the initial information of the to-be-generated image needs to be provided by the user, and the text generation model can expand and generate at least one visual effect information of high professional degree and accurate description of the to-be-generated image category corresponding to the image category information based on the image category information, so that the image generation model can generate at least one target image of the corresponding category and with the corresponding visual effect based on the target text information, and the target image with different visual effects and rich content can be generated through the target text information including different visual effect information, so as to enable the user to select the desired image, effectively improve the quality and efficiency of the generated image, and reduce the application difficulty of the text generation technology.
[0056] In addition, since the image category information is included in the target text information, the image generated by the text description information has high matching degree with the user demand, and the user experience is greatly improved.
[0057] In some optional embodiments, the visual effect information in the embodiments of the present disclosure includes any one or more of the following information: image style, image color, image shooting angle, image brightness effect, visual feeling expressed by the image, mood expressed by the image, environment expressed by the image, image visual attribute, image category posture, and image category quantity. The image style can be a description of the image style, for example, the image style can be a general description of the image style, such as ancient style, science fiction, dream, and post-apocalyptic, or a specific painting style of a painter, such as Picasso, Monet, and Van Gogh, or a style of image texture, such as film, cartoon, oil painting, illustration, impressionism, cubism, abstraction, and pop art. The image color can refer to the background color of the image or the main color tone of the image, for example, the color of the image can be a warm color tone. The visual feeling expressed by the image can refer to the effect presented by light, shadow, and color in the image. The mood expressed by the image can include, but is not limited to, sadness, humor, disappointment, and wit. The environment expressed by the image can include the background environment in the image, for example, the environment expressed by the image can be the universe, grassland, desert, and high-rise building. The environment expressed by the image can also include information about the characters or animals cooperating with the image category information, for example, the image category information is a horse, and the environment expressed by the image is a person wearing a spacesuit, then the target image generated can present a person wearing a spacesuit riding a horse; for another example, the image category information is a girl, and the environment expressed by the image is a young Chinese couple, then the target image generated can present the girl and her lover. The image visual attribute can be used to describe the related attributes of the image. For example, the image visual attribute can include the material, definition, color saturation, contrast, and noise of the image. The image category posture is used to describe or limit the posture of the image category information, for example, the image category information is a girl (character), and the image category posture can be standing, leaning, lying on one side, and lying flat. The image category quantity is used to describe or limit the number of image category information, for example, the image category information is an airplane, and the image category quantity can be 1, 2, 3, etc., for example, when the image category quantity is 1, it means that the number of airplanes in the target image is 2.
[0058] In some optional embodiments, the image visual attribute in the embodiments of the present disclosure includes any one or more of the following information: image shape, image size, image resolution, and image direction.
[0059] In the embodiments of the present disclosure, the image to be generated and the image category information are accurately and meticulously described through the visual effect information, which not only enables the image generation model to generate high-quality images according to the text description information including the visual effect information, but also enables the image generated through the text description information to be suitable for multiple application fields due to the multi-dimensional description information required for generating the image included in the visual effect information.
[0060] In an optional implementation, step S140 in the embodiments of the present disclosure can further include: inputting the target text information into the image generation model; the image generation model enhances the image category information in the target text information, and processes the target text information after the image category information is enhanced based on a cross attention mechanism to obtain at least one target image.
[0061] In a specific implementation, the target text information is input into the image generation model, in which the image category information in the target text information can be incrementally processed. For example, the image category information can be enhanced by increasing the weight of the image category information in the target text information. Then, the image generation model generates at least one target image based on the target text information after the image category information is enhanced by using a cross attention mechanism, and outputs the at least one target image.
[0062] In the embodiments of the present disclosure, by enhancing the image category information in the target text information, the target image generated by the image generation model based on the target text information after the image category information is enhanced is more relevant to the semantics of the image category information. In combination with the powerful learning ability of the image generation model for the target text information, it is realized that the target image with high professional degree, accuracy and meeting user demand can be output.
[0063] Figure 2 FIG. 1 is a flowchart of step S140 according to an example embodiment of the present disclosure. In an optional implementation, as shown in FIG. 1, step S140 can include the following steps: Figure 2
[0064] Step 141, encoding processing is performed on the target text information to obtain text features of the target text information.
[0065] In which the target text information can be encoded by a text encoder to obtain the text features of the target text information.
[0066] For example, the target text information can be encoded by a Text Encoder in CLIP (Contrastive Language-Image Pre-Training) in the image generation model to obtain an embedding vector matrix (text features) of the target text information.
[0067] Step S142, based on a cross attention mechanism, the text features of the target text information are fused with preset noise to obtain fusion features.
[0068] The image generation model can randomly generate a noise matrix of a noisy image as preset noise, and then use a cross attention mechanism to fuse the text features of the target text information with the preset noise, that is, add noise to the text features, to obtain a Cross Attention Map (fusion features).
[0069] At step S143, at least one target image corresponding to the target text information is generated based on the fusion features.
[0070] In one specific implementation, in the latent space of the image generation model, the fusion features are iteratively denoised by a forward diffusion method to obtain an information matrix of the image, and then the information matrix of the image is input into an image decoder, such as an autoencoder decoder, to output at least one target image corresponding to the target text information.
[0071] In the embodiments of the present disclosure, the cross attention mechanism is used to fuse the text features of the target text information with the preset noise to obtain fusion features, and then generate an image corresponding to the target text information based on the fusion features, thereby realizing fast and efficient generation of high-quality images based on the target text information.
[0072] In some optional embodiments, in the process of generating an image corresponding to the target text information by the image generation model, the image category information can be enhanced to make the generated image more consistent with the semantics of the initial information.
[0073] In one optional implementation, the image category information can be enhanced by the following first enhancement method, specifically including:
[0074] After step S141, the part of the text features of the target text information corresponding to the image category information can be enhanced based on preset enhancement parameters to obtain enhanced text features. Correspondingly, in this embodiment, step S142 can include: fusing the enhanced text features with the preset noise to obtain fusion features.
[0075] The preset enhancement parameter can be a hyperparameter greater than 1. For ease of description, in this embodiment, the part of the text features of the target text information corresponding to the image category information is referred to as a first target part. The first target part can be determined in the target text information through a preset format in the target text information. Then, the embedding vector corresponding to the first target part is multiplied by the preset enhancement parameter to achieve enhancement of the first target part, and an enhanced text feature is obtained. Then, the enhanced text feature is fused with the preset noise based on the cross-attention mechanism to obtain a fusion feature.
[0076] In the embodiments of the present disclosure, the part of the text features corresponding to the image category information is first enhanced by using a preset enhancement parameter to obtain an enhanced text feature, and then the enhanced text feature is fused with the preset noise to obtain a fusion feature, so that the part of the fusion feature corresponding to the image category information is enhanced, and the target image generated based on the fusion feature has a higher semantic matching degree with the initial information.
[0077] In another optional implementation, the image category information can be enhanced by the following second enhancement manner, which specifically includes:
[0078] After step S142, the part of the fusion feature corresponding to the image category information is enhanced based on a preset enhancement parameter to obtain an enhanced fusion feature. Correspondingly, in this embodiment, step S143 includes generating at least one target image corresponding to the target text information based on the enhanced fusion feature.
[0079] In this embodiment, for ease of description, the part of the fusion feature corresponding to the image category information is referred to as a second target part. The amplitude of each feature in the fusion feature determines which content the image generation model is more inclined to generate.
[0080] In an optional implementation, the preset enhancement parameter can be set to a hyperparameter greater than 1, for example, the preset enhancement parameter can be set to 1.5. Then, the amplitude of the feature corresponding to the second target part is enhanced by using formula (1) to obtain an enhanced feature corresponding to the image category information in the fusion feature, and the enhanced fusion feature is composed of the enhanced feature corresponding to the image category information in the fusion feature and other features in the fusion feature, the other features being features in the fusion feature other than the feature corresponding to the image category information.
[0081]
[0082] wherein a is the preset enhancement parameter, A is the amplitude of the feature corresponding to the image category information in the fusion feature, is the enhanced feature corresponding to the image category information in the fusion feature.
[0083] In the embodiment of the present disclosure, first, the part of the fusion feature corresponding to the image category information is enhanced based on a preset enhancement parameter to obtain an enhanced fusion feature; and then an image corresponding to the target text information is generated based on the enhanced fusion feature. In this way, by enhancing the part of the enhanced fusion feature corresponding to the image category information, when the image is generated based on the enhanced fusion feature, the image generation model is more inclined to generate content corresponding to the image category information, and thus the generated image is more consistent with the semantics of the initial information.
[0084] In some optional embodiments, the step S120 in the embodiment of the present disclosure can include: inputting the initial information into the text generation model; and processing the initial information based on an uncertainty modeling rule to obtain at least one text description information.
[0085] The image category information and the image effect information of any text description information in the at least one text description information are arranged according to a preset format. The preset format can include a position of the image category information or the visual effect information in the target text information. For example, the preset format can be that the image category information is located at a head position, a tail position or a middle position in the target text information.
[0086] In one specific implementation, the initial information of the image to be generated is input into the text generation model, the text generation model generates a text feature sequence, and then the uncertainty modeling rule is used to perform uncertainty sampling on the text feature sequence to obtain at least one text description information, and the at least one text description information is output.
[0087] In the embodiment of the present disclosure, the powerful text generation capability of the text generation model is combined with the uncertainty sampling, so that a plurality of text description information with diversified content can be generated based on one initial information, and the user experience is improved.
[0088] Figure 3 is a structural schematic diagram of the text generation model provided by an exemplary embodiment of the present disclosure. In one optional embodiment, as shown in Figure 3 the text generation model in the embodiment of the present disclosure includes: a self-recurrent language model, a sampling network and a word mapping network.
[0089] The self-recurrent language model may, for example, include but is not limited to a GPT2 model, etc.; the sampling network can include a mean subnetwork and a variance subnetwork, wherein the mean subnetwork and the variance subnetwork each include a linear layer (Linear Layer) and a layer normalization layer (Layer Normalization, LN); and the word mapping network can be a fully connected layer (Fully Connected Layer).
[0090] Figure 4 is a flowchart of step S120 provided by an example embodiment of the present disclosure, as shown in Figure 4 Step S120 can include the following steps, as shown in
[0091] Step S121, input the initial information into the autoregressive language model, generate a text feature sequence through the autoregressive language model and input it into the sampling network.
[0092] The text feature sequence includes the text feature of the initial information and the text feature of the basic visual effect information corresponding to the initial information. The basic visual effect information corresponding to the initial information is the basic visual effect information predicted by the autoregressive language model according to the initial information.
[0093] The basic visual effect information includes any one or more of the following information: image style, image color, image shooting angle, image brightness effect, visual feeling expressed by the image, emotion expressed by the image, environment expressed by the image, image visual attribute, image category number, and image category number.
[0094] In an optional implementation, the autoregressive language model can generate the basic visual effect information through multiple text predictions based on the initial information; at each text prediction, the autoregressive language model can predict the information after the input initial information and the prediction information predicted based on the initial information. The autoregressive language model extracts features from the initial information and the basic visual effect information to obtain the text feature of the initial information and the text feature of the basic visual effect information.
[0095] Step S122, the sampling network performs uncertainty sampling on the feature sequence to obtain a sampling feature sequence and input it into the word mapping network.
[0096] The sampling feature sequence includes a plurality of sampling features.
[0097] In a specific implementation, the text feature sequence output by the autoregressive language model satisfies a Gaussian distribution (Guassian Distribution), and the sampling network performs uncertainty sampling on each text feature in the text feature sequence based on the Gaussian distribution of the text feature sequence to obtain a sampling feature corresponding to each text feature. The sampling feature sequence is composed of the sampling features corresponding to each text feature.
[0098] The sampling network can also perform uncertainty sampling on each text feature in the text feature sequence based on Bayesian neural networks (Bayesian Neural Networks, BNNs) to obtain a sampling feature corresponding to each text feature.
[0099] In step S123, the word mapping network generates at least one text description information corresponding to the sampling feature sequence.
[0100] In an optional embodiment, the mapping probability between each sampling feature and each word in the preset word table is determined by the word mapping network, and the word mapped by each sampling feature is determined based on the mapping probability between each sampling feature and each word in the preset word table, and the text description information is composed of the word mapped by each sampling feature.
[0101] In the embodiments of the present disclosure, the powerful text generation capability of the autoregressive language model is utilized to efficiently generate the text feature sequence of the text feature including the initial information and the text feature of the basic visual effect information, then the uncertainty sampling is performed on each text feature in the text feature sequence to obtain the sampling feature corresponding to each text feature, and then the word mapping network determines the text description information based on each sampling feature. Since the uncertainty sampling is performed on the text feature, the content homogenization of the text description information is avoided, so that the content diversified text description information can be generated based on the same initial information, and then multiple high-quality images with different styles can be obtained based on the same initial information, which greatly improves the user experience.
[0102] In an optional embodiment, step S122 in the embodiments of the present disclosure can specifically include: for each text feature in the text feature sequence, inputting the text feature into the mean subnetwork and the variance subnetwork in the sampling network respectively to obtain the mean information and the variance information of the text feature; then performing the uncertainty sampling on the text feature based on the mean information and the variance information of the text feature to obtain the sampling feature corresponding to the text feature, and the sampling feature sequence is composed of the sampling features corresponding to each text feature in the text feature sequence.
[0103] The mean information of each text feature includes the mean value of the text feature, and the variance information of each text feature includes the variance value of the text feature.
[0104] Each text feature is subjected to the uncertainty modeling by the mean subnetwork and the variance subnetwork, and then the sampling feature corresponding to each text feature is obtained by sampling in the Gaussian distribution of the text feature sequence based on the uncertainty modeling.
[0105] In a specific implementation, it is assumed that the text feature sequence is X={x1, …, xm, …, xN}, where X is the text feature sequence, x1, xm and xN represent the first text feature, the mth text feature and the Nth text feature in the text feature sequence respectively, and 1<m<N. The mean sequence U (the mean information) of each text feature is determined by the mean subnetwork, U={μ1(x1), …, μm(xm), …, μN(xN)}, where μ1(x1), μm(xm) and μN(xN) represent the mean value of the first text feature, the mth text feature and the Nth text feature respectively. m N m N , where U is the mean sequence (the mean information), μ1(x1), μm(xm) and μN(xN) represent the mean value of the first text feature, the mth text feature and the Nth text feature respectively, and 1<m<N. The variance sequence V (the variance information) of each text feature is determined by the variance subnetwork, V={σ1(x1), …, σm(xm), …, σN(xN)}, where σ1(x1), σm(xm) and σN(xN) represent the variance value of the first text feature, the mth text feature and the Nth text feature respectively.m (x m ),…,μ N (x N )}. m (x m ) and μ N (x N ) represent the mean values of x1, x m and x N , respectively. The variance sequence E(variance information) of each text feature is determined by a variance subnetwork, E = {σ1(x1),…, σ m (x m ),…, σ N (x N )}. m (x m ) and σ N (x N ) represent the variance values of x1, x m and x N , respectively. Based on the variance values and the mean values of each text feature and a preset sampling parameter sequence ε, the Gaussian distribution sampling feature of each text feature is calculated by formula (2), and the Gaussian distribution sampling feature of each text feature is determined as the sampling feature corresponding to each text feature, and the sampling feature sequence
[0106]
[0107] wherein ε = {∈1…∈ m , …∈ N}, ∈1, ∈ m and ∈ N represent sampling parameters, ∈ i ~ M(0, 1), M(0, 1) represents a Gaussian distribution of the text feature constructed by formula (1), with a mean value of 0 and a variance value of 1, and 1≤i≤N. and represent the sampling features of x1, x m and x N , respectively.
[0108] In the embodiments of the present disclosure, the powerful computing power of the mean subnetwork and the variance subnetwork is utilized to quickly determine the mean information and the variance information of each text feature, and then the uncertainty sampling is performed on each text feature based on the mean information and the variance information of each text feature to obtain the sampling feature corresponding to each text feature, thereby realizing efficient and accurate uncertainty sampling of the text feature and providing reliable data support for obtaining the text description information from the sampling feature sequence obtained by the uncertainty sampling.
[0109] In an optional implementation, step S123 in the embodiments of the present disclosure can specifically include: the word mapping network searching for the words corresponding to each sampling feature sequence in the preset word table based on a preset word search strategy, and generating at least one text description information based on the words corresponding to each sampling feature.
[0110] For example, the preset word search strategy can include, but is not limited to, Beam Search, Greedy Search, Top-p search, Top-k search, etc. The preset word search strategy can also include a preset number of texts.
[0111] For example, assuming that the preset word search strategy includes Beam Search and the preset number of texts is 3, the word mapping network uses Beam Search to search for three text description information corresponding to the sampling feature sequence based on the mapping probability between each sampling feature and each word in the preset word table.
[0112] In the embodiments of the present disclosure, the word mapping network can not only efficiently and quickly search for the text description information corresponding to the sampling feature through the preset word search strategy, but also can usually generate multiple text description information including different visual effect information through an initial information in the preset word search strategy.
[0113] Figure 5 is a flowchart of a method for generating an image based on a text provided by another exemplary embodiment of the present disclosure. In an optional implementation, the text generation model can be obtained in the following manner, for example, as shown in Figure 5 The method includes the following steps:
[0114] Step S210, obtaining a training data set.
[0115] The training data set includes a plurality of training samples.
[0116] Each training sample includes image category information and visual effect information.
[0117] In an optional implementation, the text used to describe the image can be obtained from an open-source image-text database as a training sample. Alternatively, the image description text corresponding to the high-quality image generated by the Stable Diffusion model can also be obtained as a training sample.
[0118] Step S220, inputting each training text into a to-be-trained model.
[0119] The to-be-trained model includes a to-be-trained autoregressive language model, a to-be-trained sampling network, and a to-be-trained word mapping network.
[0120] The structures of the to-be-trained autoregressive language model, the to-be-trained sampling network and the to-be-trained word mapping network are the same as those of the autoregressive language model, the sampling network and the word mapping network in the text generation model. For details of the structure of the to-be-trained model, refer to the structure of the text generation model shown in Figure 3 .
[0121] In step S230, the to-be-trained autoregressive language model generates a predicted text feature sequence based on the image type information of the training text for each training sample.
[0122] The predicted feature sequence includes the predicted text feature of the image type information of the training text and the predicted text feature of the predicted basic visual effect information corresponding to the image type information.
[0123] The to-be-trained autoregressive language model generates a predicted text feature sequence based on the image type information of the training text in the same way as the autoregressive language model generates a text feature sequence based on initial information, which is not described here.
[0124] In step S240, the to-be-trained sampling network is used to perform uncertainty sampling on the predicted feature sequence to obtain a predicted sampling feature sequence.
[0125] The predicted sampling feature sequence includes a plurality of predicted sampling features.
[0126] In an optional embodiment, the to-be-trained sampling network in the embodiments of the present disclosure includes a to-be-trained mean sub-network and a to-be-trained variance sub-network.
[0127] Correspondingly, in this embodiment, step S240 can further include: for each predicted text feature in the predicted text feature sequence, inputting the predicted text feature into the to-be-trained mean sub-network and the to-be-trained variance sub-network respectively to obtain the mean information and the variance information of the predicted text feature; and then performing uncertainty sampling on the predicted text feature based on the mean information and the variance information of the predicted text feature to obtain a predicted sampling feature corresponding to the predicted text feature.
[0128] The mean information and the variance information of the predicted text feature obtained by the to-be-trained mean sub-network and the to-be-trained variance sub-network, and the way of performing uncertainty sampling on the predicted text feature based on the mean information and the variance information of the predicted text feature can refer to the corresponding specific implementation modes of step S122, which are not described here.
[0129] In step S250, the to-be-trained word mapping network is used to generate predicted text description information corresponding to the predicted sampling feature sequence.
[0130] The predicted text description information includes image category information and visual effect information in a preset format.
[0131] The manner in which the to-be-trained word mapping network generates the predicted text description information corresponding to the predicted sampling feature sequence is the same as the manner in which the word mapping network generates the text description information corresponding to the sampling feature sequence, and details are not repeated here.
[0132] In step S260, the to-be-trained model is fine-tuned based on each training sample and the predicted sampling feature sequence to obtain the text generation model.
[0133] The preset loss function can be a cross-entropy loss function, a mean square error function, or the like, and can be used as the loss function of the to-be-trained model. The loss function value of the to-be-trained model can be determined by using the preset loss function according to each predicted sampling feature in each training sample and the predicted sampling feature sequence.
[0134] In an optional embodiment, the parameters of the to-be-trained model can be adjusted by using a parameter optimizer. The parameter optimizer can include, but is not limited to, an SGD (Stochastic Gradient Descent), an Adagrad, an Adam (Adaptive Moment Estimation), an AdamW (Adaptive Moment Estimation Weight), an RMSprop (Root Mean Square Prop), an LBFGS algorithm (Limited-memory Broyden–Fletcher–Goldfarb–Shanno), or the like. Specifically, the gradient of each parameter in the to-be-trained model can be calculated by using the parameter optimizer, and each parameter can be fine-tuned in the direction of the gradient. The gradient represents the direction in which the loss function value decreases the most. The above operations of inputting the training text into the to-be-trained autoregressive language model, calculating the loss function value of the to-be-trained model, and fine-tuning the parameters in the to-be-trained model are iteratively performed until the loss function value of the to-be-trained model no longer decreases, and the training of the to-be-trained model is determined to be completed. The text generation model is obtained by using the trained to-be-trained model.
[0135] Exemplarily, the training data set can include 80,000-100,000 training samples, and each training sample includes at least 10 words. The to-be-trained model includes a to-be-trained autoregressive language model, a to-be-trained sampling network, and a to-be-trained word mapping network. The to-be-trained sampling network can include a to-be-trained mean sub-network and a to-be-trained variance sub-network. GPT2 can be selected as the to-be-trained autoregressive language model, and the parameter size of GPT2 is 175M. The to-be-trained mean sub-network and the to-be-trained variance sub-network are both composed of a linear layer and a standardization layer, and the parameter size of the to-be-trained mean sub-network and the to-be-trained variance sub-network is 0.6M. The to-be-trained word mapping network is composed of a full connection layer, and the parameter of the to-be-trained word mapping network is relatively small and can be ignored.
[0136] The training configuration of the to-be-trained model includes a parameter optimizer AdamW, a batch size 20, a learning rate adjustment strategy warm up, a warm up step number 1000, and a learning rate 0.0004.
[0137] The to-be-trained model is trained for 10 epochs on a GPU 3090 based on the training configuration. The epoch refers to the number of times of forward propagation and backward propagation of the entire training data set in the to-be-trained model.
[0138] In the embodiment of the present disclosure, the structure of the to-be-trained model is designed to include the to-be-trained autoregressive language model, the to-be-trained sampling network, and the to-be-trained word mapping network, and the to-be-trained model is trained by using the training text including the image category information and the visual effect information. Thus, the to-be-trained model can quickly learn to generate diversified text description information based on the image category information in the training text, and the text generation model obtained by training can output text description information with more rich and diversified content, thereby improving the quality of the image generated by using the text description information.
[0139] Figure 6 FIG. 4 is a flowchart of step S260 according to an example embodiment of the present disclosure. In an optional implementation, as shown in FIG. 4, step S260 can include the following steps. Figure 6
[0140] Step S261, determining a first loss function value based on each training sample and the predicted sampling feature sequence.
[0141] In an optional implementation, the first preset loss function can be a cross-entropy loss function. In an optional implementation, the cross-entropy loss function value is calculated by using formula (3) based on each training sample and the predicted sampling feature sequence, and the cross-entropy loss function value is determined as the first loss function value.
[0142]
[0143] wherein, is a first loss function value, is a predicted sampling feature, is a corresponding word y in the training text i a weight w in the word mapping network to be trained, c is a corresponding word c in the preset vocabulary, C is the number of words in the preset vocabulary
[0144] Step S262, determining a second loss function value according to the mean information and the variance information of each predicted text feature.
[0145] wherein, the second preset loss function can be a KL (Kullback-Leibler Divergence) divergence loss function. In an optional implementation, based on the mean information and the variance information of each predicted text feature, the KL divergence loss function value is calculated by using formula (4), and the KL divergence loss function value is determined as the second loss function value.
[0146]
[0147] wherein, is a second loss function value, σ' i is a variance value of the predicted text feature, μ' i is a mean value of the predicted text feature.
[0148] Step S263, fine-tuning the parameters of the model to be trained according to the first loss function value and the second loss function value, to obtain a text generation model.
[0149] wherein, the total loss function value of the model to be trained can be calculated by using formula (5) according to the first loss function value and the second loss function value, and then the parameters of the function to be trained are fine-tuned based on the total loss function value to obtain the text generation model.
[0150]
[0151] wherein, is a total loss function value, λ is a hyperparameter, which is used to balance the first loss function value and the second loss function value.
[0152] In the embodiments of the present disclosure, the first loss function value is determined through each training sample and the predicted sampling feature sequence, the second loss function value is determined based on the mean information and the variance information of each predicted text feature, and then the parameters of the to-be-trained model are fine-tuned based on the first loss function value and the second loss function value to obtain the text generation model. Meanwhile, the parameters of the to-be-trained model are adjusted based on the first loss function value and the second loss function value, which can efficiently adjust the parameters in the to-be-trained autoregressive language model and the to-be-trained sampling network, so that the text generation model obtained by training can accurately generate the text feature sequence and the sampling feature sequence.
[0153] In an optional embodiment, the following is an effect verification of the method for generating an image based on a text in the embodiments of the present disclosure.
[0154] In the present embodiment, the text generation model comprises an autoregressive language model, a sampling network and a word mapping network. The sampling network can comprise a mean subnetwork and a variance subnetwork. The autoregressive language model can be a GPT2 model. The mean subnetwork and the variance subnetwork each comprise a linear layer and a layer normalization layer. The word mapping network can be a fully connected layer. The image generation model is a Stable Diffusion model.
[0155] Ten initial information are set, and the ten initial information respectively comprise image category information of airplane, automobile, bird, cat, dog, frog, horse, ship and truck.
[0156] In the present embodiment, four image generation experiment groups are set up, and the images generated by the four image generation experiment groups are evaluated in two dimensions of semantic matching degree and artisticity. The semantic matching degree refers to the matching degree of the image and the semantics of the initial information, which can be determined by the CLIP score between the image and the semantics of the initial text through the CLIP model. The higher the CLIP score, the higher the matching degree between the image and the semantics of the initial text. Artisticity can be scored by the aesthetic model to obtain the aesthetic score of the image. The higher the aesthetic score of the image, the more beautiful the image is. The aesthetic model can train a neural network to obtain the aesthetic score of the image, and the neural network can be a CNN (Convolutional Neural Network). The training images can be obtained by the following method: obtaining a plurality of initial images from an open-source image-text database, and manually scoring and labeling the artisticity of the initial images to obtain a plurality of training images.
[0157] The first image generation experiment group: input the initial information into the Stable Diffusion model, and output the image by the Stable Diffusion model.
[0158] The second image generation experiment group: input the initial information into the GPT2 model, output the image description text by the GPT2 model, input the image description text into the Stable Diffusion model, and output the image by the Stable Diffusion model.
[0159] The third image generation experiment group: input the initial information into the GPT2 model, output the image description text by the GPT2 model, input the image description text into the Stable Diffusion model, enhance the image category information in the image reasoning process, and output the image by the Stable Diffusion model.
[0160] The fourth image generation experiment group: input the initial information into the text generation model, output the text description information by the text generation model, determine the target text information from the output text description information, input the target text information into the Stable Diffusion model, enhance the image category information in the image reasoning process, and output the target image by the Stable Diffusion model.
[0161] Ten initial information are generated into 20 images by the method shown in the four experimental groups, and the average of the CLIP score and the average of the aesthetic score of the images of each group are calculated, and the results are shown in Table 1.
[0162] Table 1
[0163] Average of CLIP score Average of aesthetic score First image generation experiment group 0.2646 25.8303 Second image generation experiment group 0.2512 34.2250 Third image generation experiment group 0.2678 33.7375 Fourth image generation experiment group 0.2644 33.2335
[0164] From Table 1, it can be seen that the average of the CLIP score of the images generated by the method shown in the second image generation experimental group is 0.2512, and the average of the aesthetic score is 34.2250. Although the average of the aesthetic score is improved compared with the images generated by the first image generation experimental group, the average of the CLIP score of the images is lower than that of the images generated by the first image generation experimental group, that is, the matching degree of the voice of the initial information and the image is low.
[0165] The third image generation experiment group enhances the image category information in the inference process of the image, so as to improve the average value of the aesthetic score of the image while ensuring the average value of the CLIP score of the image. However, the diversity of the generated image description text is poor, the generated image style is single, and there is a homogenization situation. For example, the image category information in the initial information is horse. Three image description texts are generated by the method shown in the third image generation experiment group, which are: No. 1. Horse-like creature with long horns, long tongues, and a long nose appearing from the ground, in the style of beeple and Mike Winkelmann, intricate, epic lighting, cinematic composition, hyper realistic, 8k resolution, unreal engine 5. No. 2. Horse-like creature with long horns, long tongues, and a long nose appearing from the ground, in the style of beeple and Mike Winkelmann, intricate, epic lighting, cinematic composition, hyper realistic, 8k resolution, unreal engine 5. No. 3. Horse-like creature with long horns, long tongues, and a long nose appearing from the ground, in the style of beeple and Mike Winkelmann, intricate, epic lighting, cinematic composition, 8k resolution, unreal engine 5. Figure 7 is a schematic diagram of an image generated by an image description text according to an example embodiment of the present disclosure. Wherein Figure 7 A part in is an image generated by the image description text No. 1, Figure 7 B part in is an image generated by the image description text No. 2, Figure 7Part C in is the image generated by the image description text in number 3. Figure 7 From the images shown in , it can be concluded that the images generated by the three image description texts generated using the method shown in the third image generation experimental group have a single style and are homogenized.
[0166] The fourth image generation experiment group added a sampling network to enhance the diversity of text description information. Table 1 shows that the images generated by the methods shown in the fourth image generation experiment group have high average CLIP scores and average aesthetic scores, and the generated text description information is diverse, corresponding to the diverse styles of the generated images. For example, the image category information in the initial information is horse. The three image description texts generated by the methods in the fourth image generation experiment group are: No. 4. Horse in spacesuit, engineer, boots, intricate, elegant, highly detailed, digital painting, artstation, concept art, smooth, sharp focus, illustration, by gregrutkowski and alphonsemucha. No. 5. Horse wearing a jacket and white boots, riding a horse, looking at the camera, standing in a field, wearing glasses, art by gregrutkowski, hyperdetailed, 8k, concept art, trending on artstation. No. 6. Horse with a sword, fantasy, D&D, intricate, rings, smoke, fire, highly detailed, digital painting, artstation, concept art, matte, sharp focus, illustration, hearthstone, Furyblade. The three text description information are respectively determined as target text information. Figure 8 FIG is a schematic diagram of a target image generated by target text information provided by an exemplary embodiment of the present disclosure. Figure 8 Part A in the figure is the target image generated by the target text information numbered 4. Figure 8 Part B in the figure is the target image generated by the target text information numbered 5.Figure 8 Part C in is the target image generated by the target text information numbered 6. Figure 8 The target images shown in the figure show that the target images generated by the three target text information generated by the fourth image generation experimental group method are not only diverse in style and highly artistic, but also highly matched with the image category information "horse" in the initial information.
[0167] Figure 9 This is a schematic diagram of the structure of an embodiment of the apparatus for generating an image based on text disclosed in the present invention. Figure 9 As shown, the device of this embodiment may include:
[0168] A first acquisition module 300 is used to acquire initial information of an image to be generated, wherein the initial information includes image category information;
[0169] A text generation module 310 is configured to input the initial information into a pre-trained text generation model to obtain at least one text description information, wherein the text description information includes the image category information and image effect information;
[0170] The text determination module 320 is used to determine the target text information based on the at least one text description information.
[0171] The image generation module 330 is configured to input the target text information into a pre-trained image generation model to obtain at least one target image.
[0172] In some possible implementations of the present disclosure, the image generation module 330 in the embodiment of the present disclosure is specifically used to: input the target text information into the image generation model; the image generation model enhances the image category information in the target text information, and processes the target text information after the enhanced image category information based on the cross-attention mechanism to obtain the at least one target image.
[0173] In some possible implementations of the present disclosure, the image generation module 330 in the embodiment of the present disclosure includes:
[0174] A text encoding submodule is used to encode each piece of text description information to obtain text features of the text description information;
[0175] A feature fusion submodule is used to fuse the text features with the preset noise based on a cross attention mechanism to obtain a fused feature;
[0176] The image generation submodule is used to generate an image corresponding to the text description information based on the fusion feature.
[0177] In some possible implementation manners of the present disclosure, the image generation module 330 in the present embodiment further includes:
[0178] The first enhancement submodule is configured to perform feature enhancement on the part of the text features corresponding to the image category information based on preset enhancement parameters, to obtain enhanced text features.
[0179] The feature fusion submodule is configured to perform feature fusion on the enhanced text features and the preset noise.
[0180] Alternatively,
[0181] The second enhancement submodule is configured to perform feature enhancement on the part of the fusion features corresponding to the image category information based on preset enhancement parameters, to obtain enhanced fusion features.
[0182] The image generation submodule is configured to generate an image corresponding to the text description information based on the enhanced fusion features.
[0183] In some possible implementation manners of the present disclosure, the text generation module 310 in the present embodiment is specifically configured to input the initial information into the text generation model, and perform processing on the initial information based on an uncertainty modeling rule to obtain the at least one text description information, wherein the image category information and the image effect information of any text description information in the at least one text description information are arranged in a preset format.
[0184] In some possible implementation manners of the present disclosure, the text generation model in the present embodiment includes an autoregressive language model, a sampling network and a word mapping network.
[0185] In some possible implementation manners of the present disclosure, the text generation module 310 in the present embodiment includes:
[0186] The text feature generation submodule is configured to input the initial information into the autoregressive language model, generate a text feature sequence through the autoregressive language model and input the text feature sequence into the sampling network, wherein the text feature sequence includes text features of the initial information and text features of basic visual effect information corresponding to the initial information.
[0187] The sampling feature generation submodule is configured to perform uncertainty sampling on the feature sequence through the sampling network to obtain a sampling feature sequence, and input the sampling feature sequence into the word mapping network, wherein the sampling feature sequence includes a plurality of sampling features.
[0188] The mapping submodule is configured to generate at least one text description information corresponding to the sampling feature sequence through the word mapping network.
[0189] In some possible implementation manners of the present disclosure, the sampling network in the embodiment of the present disclosure comprises: a mean sub-network and a variance sub-network.
[0190] In some possible implementation manners of the present disclosure, the sampling feature generation sub-module in the embodiment of the present disclosure is specifically configured to:
[0191] respectively inputting each text feature in the text feature sequence into the mean sub-network and the variance sub-network to obtain mean information and variance information of the text feature;
[0192] performing uncertainty sampling on the text feature based on the mean information and the variance information of the text feature to obtain a sampling feature corresponding to the text feature.
[0193] In some possible implementation manners of the present disclosure, the mapping sub-module in the embodiment of the present disclosure is specifically configured to:
[0194] The word mapping network searches for words corresponding to each sampling feature sequence in a preset word table based on a preset word search strategy, and generates the at least one text description information based on the words corresponding to each sampling feature.
[0195] In some possible implementation manners of the present disclosure, the visual effect information in the embodiment of the present disclosure comprises any one or more of the following information: image style, image color, image shooting angle, image brightness effect, visual feeling expressed by the image, emotion expressed by the image, environment expressed by the image, image visual attribute, image category number and image category posture.
[0196] In some possible implementation manners of the present disclosure, the image visual attribute in the embodiment of the present disclosure comprises any one or more of the following information: image shape, image size, image resolution and image direction.
[0197] In some possible implementation manners of the present disclosure, the device for generating an image based on a text in the embodiment of the present disclosure further comprises:
[0198] The second acquisition module is configured to acquire a training data set, the training data set comprising a plurality of training samples, each training sample comprising image category information and visual effect information;
[0199] The first training module is configured to input each training text into a to-be-trained model, the to-be-trained model comprising: a to-be-trained autoregressive language model, a to-be-trained sampling network and a to-be-trained word mapping network.
[0200] a second training module configured to, for each training sample, generate, by the to-be-trained autoregressive language model, a predicted text feature sequence based on image type information of the training text, the predicted text feature sequence including predicted text features of the image type information and predicted text features of predicted basic visual effect information corresponding to the image type information;
[0201] a third training module configured to perform uncertainty sampling on the predicted text feature sequence by the to-be-trained sampling network to obtain a predicted sampling feature sequence, the predicted sampling feature sequence including a plurality of predicted sampling features;
[0202] a fourth training module configured to generate, by the to-be-trained word mapping network, predicted text description information corresponding to the predicted sampling feature sequence;
[0203] a fifth training module configured to fine-tune parameters of the to-be-trained model based on each training sample and the predicted sampling feature sequence to obtain the text generation model.
[0204] In some possible implementation manners of the present disclosure, the to-be-trained sampling network in the embodiment of the present disclosure includes a to-be-trained mean subnetwork and a to-be-trained variance subnetwork;
[0205] In some possible implementation manners of the present disclosure, the third training module in the embodiment of the present disclosure is specifically configured to: for each predicted text feature in the predicted text feature sequence, input the predicted text feature into the to-be-trained mean subnetwork and the to-be-trained variance subnetwork respectively to obtain mean information and variance information of the predicted text feature;
[0206] perform uncertainty sampling on the predicted text feature based on the mean information and the variance information of the predicted text feature to obtain a predicted sampling feature corresponding to the predicted text feature.
[0207] In some possible implementation manners of the present disclosure, the fifth training module in the embodiment of the present disclosure includes:
[0208] a first loss value determination sub-module configured to determine a first loss function value based on each training sample and the predicted sampling feature sequence;
[0209] a second loss value determination sub-module configured to determine a second loss function value according to the mean information and the variance information of each predicted text feature;
[0210] a training sub-module configured to fine-tune parameters of the to-be-trained model according to the first loss function value and the second loss function value to obtain the text generation model.
[0211] The apparatus for generating an image based on text and the method for generating an image based on text of the embodiments of the present disclosure correspond to each other in implementation, and can refer to each other for the corresponding content, which will not be described herein again.
[0212] In addition, the embodiments of the present disclosure also provide an electronic device, comprising:
[0213] a memory for storing a computer program;
[0214] a processor for executing the computer program stored in the memory, and when the computer program is executed, the method for generating an image based on text of any of the embodiments of the present disclosure is implemented.
[0215] Figure 10 The structural schematic diagram of an application embodiment of the electronic device of the present disclosure is shown in FIG. 1. Hereinafter, the electronic device according to the embodiments of the present disclosure will be described with reference to FIG. 1. Figure 10 The electronic device can be any one or both of the first device and the second device, or a single device independent of them, which can communicate with the first device and the second device to receive the collected input signals therefrom.
[0216] As shown in FIG. 1, the electronic device comprises one or more processors and a memory. Figure 10
[0217] The processor can be a central processing unit (CPU) or other forms of processing units having data processing and / or instruction execution capabilities, and can control other components in the electronic device to perform desired functions.
[0218] The memory can comprise one or more computer program products, which can comprise various forms of computer readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory can comprise, for example, random access memory (RAM), cache memory and / or the like. The non-volatile memory can comprise, for example, read-only memory (ROM), hard disk, flash memory and / or the like. One or more computer program instructions can be stored on the computer readable storage media, and the processor can run the program instructions to implement the method for generating an image based on text of the embodiments of the present disclosure described above and / or other desired functions.
[0219] In one example, the electronic device can further comprise input devices and output devices, which are interconnected through a bus system and / or other forms of connection mechanism (not shown).
[0220] In addition, the input devices can further comprise, for example, a keyboard, a mouse and / or the like.
[0221] The output device can output various information including the determined distance information, direction information, etc. to the outside. The output device can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, etc.
[0222] Of course, in order to simplify, Figure 10 Only some of the components of the electronic device related to the present disclosure are shown in the middle, and components such as buses, input / output interfaces, etc. are omitted. In addition, the electronic device can further include any other appropriate components according to the specific application.
[0223] In addition to the above-mentioned method and device, an embodiment of the present disclosure can also be a computer program product including computer program instructions, which, when executed by a processor, causes the processor to perform the steps of the method of generating an image based on text according to various embodiments of the present disclosure described in the above part of the specification.
[0224] The computer program product can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, C++, etc., and a conventional procedural programming language such as "C" language or similar programming languages. The program code can be executed entirely on a user computing device, partially on a user device, as an independent software package, partially on a user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0225] In addition, an embodiment of the present disclosure can also be a computer readable storage medium having stored thereon computer program instructions, which, when executed by a processor, cause the processor to perform the steps of the method of generating an image based on text according to various embodiments of the present disclosure described in the above part of the specification.
[0226] The computer readable storage medium can employ any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples (non-exhaustive list) of readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any appropriate combination of the above.
[0227] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by program instruction related hardware, and the foregoing program can be stored in a computer readable storage medium, and the program performs the steps of the above-mentioned method embodiments when executed; and the foregoing storage medium includes ROM, RAM, magnetic disk or optical disk and various storage medium that can store program codes.
[0228] The above describes the basic principles of the present disclosure in combination with specific embodiments, but it should be noted that the advantages, advantages, effects and the like mentioned in the present disclosure are only examples and are not limiting, and these advantages, advantages, effects and the like cannot be considered as the various embodiments of the present disclosure must have. In addition, the above specific details of the disclosure are only for the purpose of example and for the purpose of understanding, and are not limited to the above specific details, and the above details do not limit the present disclosure to be implemented with the above specific details.
[0229] Each embodiment in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same or similar parts between each embodiment can be referred to each other. For system embodiments, since they basically correspond to method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.
[0230] The block diagrams of the devices, apparatuses, equipment, systems involved in the present disclosure are only illustrative examples and are not intended to require or imply the connection, arrangement, configuration shown in the block diagram. As those skilled in the art will recognize, these devices, apparatuses, equipment, systems can be connected, arranged, configured in any way. Words such as "include", "contain", "have" and the like are open-ended words, which mean "including but not limited to", and can be used interchangeably. The words "or" and "and" used herein mean the word "and / or", and can be used interchangeably unless the context clearly indicates otherwise. The word "such as" used herein means the phrase "such as but not limited to", and can be used interchangeably.
[0231] The methods and apparatuses of the present disclosure can be implemented in many ways. For example, the methods and apparatuses of the present disclosure can be implemented by software, hardware, firmware or any combination of software, hardware and firmware. The above order of steps for the method is only for illustration, and the steps of the method of the present disclosure are not limited to the above specifically described order, unless otherwise specifically described. In addition, in some embodiments, the present disclosure can also be implemented as a program recorded in a recording medium, which includes machine readable instructions for implementing the method according to the present disclosure. Therefore, the present disclosure also covers the recording medium storing the program for executing the method according to the present disclosure.
[0232] It is also important to note that the devices, apparatuses and methods described in the disclosure can be embodied in a variety of other forms, including but not limited to oral, written, and / or visual forms. Similarly, the devices, apparatuses and methods described in the disclosure can be embodied as one or more computer-readable storage media having computer-readable code embodied thereon that, when executed by one or more computer processors, cause the computer processors to perform a method as described in the disclosure. The computer-readable storage media can be non-transitory storage media. As used herein, the term "non-transitory" merely means that the computer- readable code does not create a record of signals on the computer-readable storage media, as opposed to being embodied in a transitory signal, such as a modulated data signal. Similarly, the devices, apparatuses and methods described in the disclosure can be embodied as one or more computer programs that, when executed by one or more computer processors, cause the computer processors to perform a method as described in the disclosure.
[0233] The above description of the disclosed aspects is given for illustrative and descriptive purposes. Various modifications to the aspects will be readily apparent to those skilled in the art and the general principles defined herein can be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0234] The above description has been given for the purpose of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present disclosure to forms disclosed herein. Although various example aspects and embodiments have been discussed above, those of ordinary skill in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.
Claims
1. A method for generating an image based on text, characterized in that: include: Acquire initial information of an image to be generated, the initial information including image category information; Inputting the initial information into a pre-trained text generation model to obtain at least one text description information, wherein the text description information includes the image category information and image effect information; determining target text information based on the at least one text description information; Inputting the target text information into a pre-trained image generation model to obtain at least one target image; The text generation model includes: an autoregressive language model, a sampling network and a word mapping network; the inputting the initial information into the pre-trained text generation model to obtain at least one text description information includes: inputting the initial information into the autoregressive language model, generating a text feature sequence through the autoregressive language model and inputting it into the sampling network, the text feature sequence including: text features of the initial information and text features of basic visual effect information corresponding to the initial information; performing uncertainty sampling on the feature sequence through the sampling network to obtain a sampled feature sequence, and inputting it into the word mapping network, the sampled feature sequence including multiple sampled features; generating at least one text description information corresponding to the sampled feature sequence through the word mapping network; The sampling network includes: a mean subnetwork and a variance subnetwork; The performing uncertainty sampling on the feature sequence through the sampling network to obtain a sampled feature sequence includes: For each text feature in the text feature sequence, input the text feature into the mean sub-network and the variance sub-network respectively to obtain mean information and variance information of the text feature; Using a preset sampling parameter sequence, uncertainty sampling is performed in the Gaussian distribution of the text feature sequence based on the mean information and variance information of the text features, the Gaussian distribution sampling features of each text feature are calculated, and the Gaussian distribution sampling features of each text feature are determined as the sampling features corresponding to each text feature to obtain the sampling feature sequence.
2. The method according to claim 1, characterized in that The step of inputting the target text information into a pre-trained image generation model to obtain at least one target image comprises: Inputting the target text information into the image generation model; The image generation model enhances the image category information in the target text information, and processes the target text information after the enhanced image category information based on a cross-attention mechanism to obtain the at least one target image.
3. The method according to claim 1, characterized in that The step of inputting the target text information into a pre-trained image generation model to obtain at least one target image comprises: Performing encoding processing on the target text information to obtain text features of the target text information; Based on the cross attention mechanism, the text features and the preset noise are fused to obtain fused features; At least one target image corresponding to the text description information is generated based on the fused features.
4. The method according to claim 3, characterized in that After obtaining the text features of the target text information, the method further includes: Based on preset enhancement parameters, feature enhancement is performed on a portion of the text feature corresponding to the image category information to obtain an enhanced text feature; Performing feature fusion on the text feature and the preset noise, including: performing feature fusion on the enhanced text feature and the preset noise; or, After obtaining the fusion feature, the method further includes: based on a preset enhancement parameter, performing feature enhancement on a portion of the fusion feature corresponding to the image category information to obtain an enhanced fusion feature; Generating an image corresponding to the text description information based on the fusion feature includes: At least one target image corresponding to the target text information is generated based on the enhanced fusion features.
5. The method according to claim 1, wherein The step of inputting the initial information into a pre-trained text generation model to obtain at least one text description information includes: Inputting the initial information into the text generation model; The text generation model processes the initial information based on uncertain modeling rules to obtain the at least one text description information, and the image category information and image effect information of any text description information in the at least one text description information are arranged in a preset format.
6. The method according to claim 1, characterized in that Generating at least one text description information corresponding to the sampled feature sequence through the word mapping network includes: The word mapping network searches for words corresponding to each sampling feature sequence in a preset word table based on a preset word search strategy, and generates the at least one text description information based on the words corresponding to each sampling feature.
7. The method according to claim 1, characterized in that The image effect information includes any one or more of the following information: image style, image color, image shooting angle, image brightness effect, visual feeling expressed by the image, emotion expressed by the image, environment expressed by the image, image visual attributes, image category posture, and image category quantity.
8. The method according to claim 7, characterized in that The image visual attributes include any one or more of the following information: image shape, image size, image resolution, and image orientation.
9. A device for generating an image based on text, characterized in that: include: A first acquisition module is used to acquire initial information of an image to be generated, wherein the initial information includes image category information; a text generation module, configured to input the initial information into a pre-trained text generation model to obtain at least one text description information, wherein the text description information includes the image category information and image effect information; a text determination module, configured to determine target text information based on the at least one text description information; An image generation module, configured to input the target text information into a pre-trained image generation model to obtain at least one target image; The text generation model includes: an autoregressive language model, a sampling network and a word mapping network; the text generation module includes: a text feature generation submodule, which is used to input the initial information into the autoregressive language model, generate a text feature sequence through the autoregressive language model and input it into the sampling network, wherein the text feature sequence includes: text features of the initial information and text features of basic visual effect information corresponding to the initial information; a sampling feature generation submodule, which is used to perform uncertainty sampling on the feature sequence through the sampling network to obtain a sampling feature sequence and input it into the word mapping network, wherein the sampling feature sequence includes multiple sampling features; a mapping submodule, which is used to generate at least one text description information corresponding to the sampling feature sequence through the word mapping network; The sampling network includes: a mean subnetwork and a variance subnetwork; The sampling feature generation submodule is further used to: For each text feature in the text feature sequence, input the text feature into the mean sub-network and the variance sub-network respectively to obtain mean information and variance information of the text feature; Using a preset sampling parameter sequence, uncertainty sampling is performed in the Gaussian distribution of the text feature sequence based on the mean information and variance information of the text features, the Gaussian distribution sampling features of each text feature are calculated, and the Gaussian distribution sampling features of each text feature are determined as the sampling features corresponding to each text feature to obtain the sampling feature sequence.
10. An electronic device, characterized in that: include: memory for storing computer programs; The processor is configured to execute the computer program stored in the memory, and when the computer program is executed, the method for generating an image based on text according to any one of claims 1 to 8 is implemented.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating an image based on text according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Image generation method and device, electronic equipment and storage medium
CN116797684A
Figure graph optimization method based on human feedback reinforcement learning
CN117593628A