Prompt word optimization method and device for character-generated image, medium and equipment

By extracting the target subject and style types in the initial prompt words entered by the user and supplementing with the rich matching vocabulary, the optimized prompt words can better utilize the creative performance of models such as Stable Diffusion to improve the quality and details of image generation.

CN120181086APending Publication Date: 2025-06-20BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311756157.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the prior art, prompt words entered by users are often insufficiently described and lack modification, resulting in low quality of generated images and the creative performance of models such as Stable Diffusion are not fully utilized.

Method used

By obtaining the initial prompt words entered by the user, extracting the target subject and target style types, determining rich vocabulary that matches the target subject and/or target style types, and supplementing them into the initial prompt words, obtaining the optimized prompt words.

Benefits of technology

The richness, refinement and accuracy of prompt words are improved, making the generated images more fully and refine in terms of subject and artistic style.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181086A_ABST
    Figure CN120181086A_ABST
Patent Text Reader

Abstract

The invention relates to a cue word optimization method and device for a character-generated image, a medium and equipment. The method comprises the following steps: acquiring an initial prompt word input by a user; extracting a target subject and a target style type based on the initial cue word; determining rich vocabularies matched with the target subject and / or the target style type; wherein the rich vocabularies comprise vocabularies used for supplementing the content of the initial cue word; and supplementing the rich vocabulary into the initial cue word to obtain an optimized cue word. According to the technical scheme, the richness and accuracy of the cue words can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular, to a method, apparatus, medium, and device for optimizing prompts for text-to-image generation. Background Art

[0002] In recent years, the ability to generate images based on text has developed rapidly. For example, Stable Diffusion has shown good performance in many application scenarios such as text-to-image. The method of generating images from text is to use the input prompt to generate satisfactory image content. Therefore, in the method of generating images from text, the prompt has a great impact on the quality of the generated image. Most of the prompts input by users have problems such as insufficient description and lack of modifying prompts, and cannot fully utilize the creative performance of models such as Stable Diffusion. Some prompts with more sufficient, accurate, detailed, and modifying descriptions can greatly improve the quality of the generated images. Therefore, how to enrich and optimize the prompts has become a key problem to be solved. Summary of the Invention

[0003] To solve the above technical problems, the present disclosure provides a method, apparatus, medium, and device for optimizing prompts for text-to-image generation to improve the richness and accuracy of the prompts.

[0004] The present disclosure provides a method for optimizing prompts for text-to-image generation, including:

[0005] Obtaining an initial prompt input by a user;

[0006] Extracting a target subject and a target style type based on the initial prompt;

[0007] Determining rich vocabulary matching the target subject and / or the target style type; wherein the rich vocabulary includes: vocabulary for supplementing the content of the initial prompt;

[0008] Supplementing the rich vocabulary into the initial prompt to obtain an optimized prompt.

[0009] In some embodiments, the extracting a target subject and a target style type based on the initial prompt includes:

[0010] Inputting the initial prompt into a pre-trained first text convolutional neural network;

[0011] Extracting the target subject and the target style type matching the initial prompt through the first text convolutional neural network.

[0012] In some embodiments, the first text convolutional neural network includes: a convolutional module with a residual structure, a fully connected module, a first output module, and a second output module;

[0013] Extracting the target subject and the target style type that match the initial prompt word through the first text convolutional neural network includes:

[0014] Extracting the context text features of the initial prompt word through the convolutional module, and inputting the context text features into the fully connected module;

[0015] Sorting the context text features in a preset order through the fully connected module to obtain a text feature vector, and inputting the text feature vector into the first output module and the second output module;

[0016] Performing subject classification on the initial prompt word according to the text feature vector through the first output module to obtain multiple candidate subjects and the weights of each candidate subject;

[0017] Smoothing the weights of each candidate subject, and determining the target subject based on the smoothed weights;

[0018] Performing art style classification on the initial prompt word according to the text feature vector through the second output module to obtain multiple candidate style types and the scores of each candidate style type;

[0019] Smoothing the scores of each candidate style type, and determining the target style type based on the smoothed scores.

[0020] In some embodiments, determining the rich vocabulary that matches the target style type includes:

[0021] Determining at least one candidate art style type that is different from the target style type and matches the target subject;

[0022] Determining the vocabulary representing the candidate art style type as the rich vocabulary that matches the target style type.

[0023] In some embodiments, determining the rich vocabulary that matches the target subject includes:

[0024] Under the specified target style type, determining the modification information for describing the target subject, where the modification information includes: nature, state, feature, and / or attribute;

[0025] Determining the vocabulary representing the modification information as the rich vocabulary that matches the target subject.

[0026] In some embodiments, the enriched vocabulary further includes: vocabulary associated with a preset text-to-image model; determining the enriched vocabulary that matches the target subject and / or the target style type includes:

[0027] Determining a first text-to-image model that matches the target subject and / or a second text-to-image model that matches the target style type;

[0028] Determining the vocabulary associated with the first text-to-image model as the enriched vocabulary that matches the target subject;

[0029] Determining the vocabulary associated with the second text-to-image model as the enriched vocabulary that matches the target style type.

[0030] In some embodiments, after supplementing the enriched vocabulary to the initial prompt to obtain an optimized prompt, the method further includes:

[0031] Reviewing and editing the optimized prompt to obtain a target prompt.

[0032] In some embodiments, the reviewing and editing of the optimized prompt includes:

[0033] Reviewing the optimized prompt according to a pre-established sample list including drawing texts; the sample list at least includes: drawing texts restricted by conditions and not allowed to be used, and some of the drawing texts are labeled;

[0034] Taking any prompt in the optimized prompt as the current prompt;

[0035] When the current prompt matches the drawing text with a preset label in the sample list, adding the label of the drawing text to the current prompt;

[0036] When the current prompt matches the drawing text not allowed to be used in the sample list, deleting the current prompt;

[0037] When the current prompt matches the drawing text restricted by conditions in the sample list, replacing the current prompt.

[0038] In some embodiments, the reviewing and editing of the optimized prompt includes:

[0039] Inputting the optimized prompt into a pre-trained second text convolutional neural network;

[0040] Reviewing and removing incorrect conflicting prompts in the optimized prompt through the second text convolutional neural network.

[0041] In some embodiments, the auditing and editing of the optimized prompt includes:

[0042] Inputting the optimized prompt into a pre-trained Seq2Seq model;

[0043] Extracting the context feature vector of the optimized prompt through the encoder of the Seq2Seq model;

[0044] Using the attention mechanism model by the decoder of the Seq2Seq model to decode the context feature vector, and auditing and removing incorrect conflicting prompts in the optimized prompt according to the decoding result.

[0045] In some embodiments, after extracting the target subject and target style type based on the initial prompt, the method further includes:

[0046] Matching the initial prompt with a preset text mapping table; wherein, the text mapping table is used to record the mapping relationship between designated words and standard words, and the designated words and the standard words are texts with the same semantics but different expressions;

[0047] When the initial prompt matches the target designated word in the text mapping table, mapping the target designated word in the initial prompt to the target standard word according to the mapping relationship to obtain the mapped initial prompt.

[0048] The present disclosure provides a prompt optimization device for generating images from text, including:

[0049] A prompt acquisition module, configured to acquire an initial prompt input by a user;

[0050] An extraction module, configured to extract a target subject and a target style type based on the initial prompt;

[0051] A vocabulary determination module, configured to determine rich vocabulary that matches the target subject and / or the target style type; wherein, the rich vocabulary includes: vocabulary for supplementing the content of the initial prompt;

[0052] A prompt optimization module, configured to supplement the rich vocabulary into the initial prompt to obtain an optimized prompt.

[0053] The present disclosure further provides a computer-readable storage medium, where the computer-readable storage medium stores a program or instructions, and the program or instructions enable a computer to execute the steps of any of the above methods.

[0054] The present disclosure further provides an electronic device, including:

[0055] One or more processors;

[0056] A memory for storing one or more programs or instructions;

[0057] The processor is configured to execute the steps of any of the above methods by calling the programs or instructions stored in the memory.

[0058] The technical solution provided by the embodiments of the present disclosure has the following advantages compared with the prior art:

[0059] The technical solution provided by the embodiments of the present disclosure first obtains an initial prompt word input by a user; secondly extracts a target subject and a target style type based on the initial prompt word; then determines rich vocabulary that matches the target subject and / or the target style type; wherein, the rich vocabulary includes: vocabulary for supplementing the content of the initial prompt word; and supplements the rich vocabulary into the initial prompt word to obtain an optimized prompt word. This technical solution will, for the initial prompt word input by the user, enrich and optimize the prompt word with rich vocabulary in terms of the subject of the image and the style type of the image by extracting the target subject and the target style type, so that the optimized prompt word can be more fully and refined in terms of the expression of the subject and the artistic style. Therefore, this solution can improve the richness, refinement degree and accuracy of the optimized target prompt word. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.

[0061] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0062] Figure 1 It is a flowchart of a method for optimizing a prompt word for generating an image from text provided by an embodiment of the present disclosure;

[0063] Figure 2 It is a schematic diagram showing the artistic style of an image provided by an embodiment of the present disclosure;

[0064] Figure 3 It is a schematic diagram of a process for optimizing a prompt word for generating an image from text provided by an embodiment of the present disclosure;

[0065] Figure 4 It is a structural block diagram of a device for optimizing a prompt word for generating an image from text provided by an embodiment of the present disclosure;

[0066] Figure 5 A structural schematic diagram of the electronic device provided by the embodiment of the present disclosure. Detailed implementation manners

[0067] In order to more clearly understand the above objects, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.

[0068] Many specific details are set forth in the following description in order to fully understand the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments.

[0069] To enrich and optimize the prompt words, the embodiment of the present disclosure provides a method, apparatus, device and medium for optimizing prompt words for text-to-image generation, which can optimize the initial prompt words input by the user based on the target subject and the target style type, and improve the richness and accuracy of the optimized target prompt words. For ease of understanding, the embodiments of the present disclosure are described in detail below.

[0070] Figure 1 A flowchart of a method for optimizing prompt words for text-to-image generation provided by the embodiment of the present disclosure. This method can be executed by a device for optimizing prompt words for text-to-image generation, and the device for optimizing prompt words for text-to-image generation can be implemented in a software and / or hardware manner, specifically, for example, an electronic device or a server. Among them, the electronic device may include a vehicle-mounted processor, a mobile phone, a tablet computer, a desktop computer, a notebook computer and other devices with communication functions. The server may be a cloud server or a server cluster and other devices with storage and computing functions. It should be noted that, in the following embodiments, the electronic device is used as the execution subject for exemplary explanation.

[0071] As Figure 1 shown, the method for optimizing prompt words for text-to-image generation includes the following steps:

[0072] S102. Obtain the initial prompt words input by the user.

[0073] In one embodiment, after the user inputs text to the Stable Diffusion model, the input text of the user is obtained, that is, the initial prompt. The Stable Diffusion model is a model that diffuses in the latent space, including: a text model Text Model, a UNet structure, and an encoder-decoder. The Stable Diffusion model is, for example, the Stable Diffusion model of version 1.5, version 2.0, or Stable Diffusion XL (abbreviated as SDXL), which is not limited herein. The focus of this embodiment is on the optimization of the prompt, so the above Stable Diffusion model will not be described in detail.

[0074] S104. Extract the target subject and the target style type based on the initial prompt.

[0075] In some embodiments, the initial prompt can be input into a pre-trained first text convolutional neural network (Text CNN, Text Convolutional Neural Networks); the target subject and the target style type matching the initial prompt are extracted through the first text convolutional neural network.

[0076] Among them, the first text convolutional neural network can include: a convolutional module with a residual structure, a fully connected module, a first output module, and a second output module. Specifically, the convolutional module can be a six-layer one-dimensional convolution applying a residual structure; the convolutional module is connected to the fully connected module behind, and the fully connected module includes three fully connected layers; the fully connected module is connected to two branches behind: the first output module and the second output module. The implementation process of extracting the target subject and the target style type matching the initial prompt through the first text convolutional neural network can be referred to as follows.

[0077] The context text features of the initial prompt are extracted through the convolutional module, and the context text features are input into the fully connected module; the context text features are used to represent the semantic parsing information of the context of the current initial prompt in the complete text input by the user, the logical relationship between the initial prompt and its context, and other information. In this embodiment, the convolutional module applying a residual structure is used to extract the context text features of the initial prompt. Since the residual structure can better utilize the context information by introducing skip connections, improve the expression ability of the context text features, and thus improve the performance of the task. At the same time, the residual structure directly passes the information of the previous layer to the subsequent layer through the skip connection, which can alleviate the problem of gradient disappearance, so that the gradient is easier to propagate in the network and the network is easier to converge.

[0078] Sort the context text features in a preset order through the fully connected module to obtain a text feature vector, and input the text feature vector into the first output module and the second output module.

[0079] Through the first output module, perform subject classification on the initial prompt word according to the text feature vector to obtain multiple candidate subjects and the weights of each candidate subject; perform smoothing processing on the weights of each candidate subject, and determine the target subject based on the smoothed weights. Specifically, the first output module performs subject classification on the initial prompt word according to the text feature vector to obtain multiple candidate subjects and the weight of each candidate subject; among them, the subject can be multiple preset types of subjects, such as: people, animals, flowers and plants, mountains, waters, etc., and the above weights are between 0 and 1. Perform post-processing such as smoothing on the candidate subjects and their weights to obtain the target subject finally matched by the initial prompt word.

[0080] Through the second output module, perform art style classification on the initial prompt word according to the text feature vector to obtain multiple candidate style types and the scores of each candidate style type; perform smoothing processing on the scores of each candidate style type, and determine the target style type based on the smoothed scores. Specifically, the second output module performs art style classification on the initial prompt word according to the text feature vector to obtain multiple candidate style types and the score of each candidate style type; among them, the art style can be multiple preset types of styles, such as: realistic style, comic style, oil painting style, etc., and the above scores are between 0 and 1. Perform post-processing such as smoothing on the candidate style types and their weights to obtain the target style type finally matched by the initial prompt word.

[0081] The first text convolutional neural network in this embodiment can not only accurately extract the target subject and target style type matched by the initial prompt word, but also, due to the simplicity of the network design, has a fast processing speed, improving the extraction efficiency of the target subject and target style type.

[0082] Considering the user's language habits, the initial prompt word input by the user may contain special words such as slang and common sayings that are difficult for machines to recognize. Based on this, after the above step S104, the method provided in this embodiment may further include:

[0083] Match the initial prompt word with a preset text mapping table; where the text mapping table is used to record the mapping relationship between specified words and standard words, and the specified words and the standard words are texts with the same semantics but different expressions;

[0084] When the initial prompt word matches the target specified word in the text mapping table, map the target specified word in the initial prompt word to the target standard word according to the mapping relationship to obtain the mapped initial prompt word.

[0085] In a specific embodiment, designated words are generally words that are pre-set by users and are not easily understood by machines, such as slang, common sayings, ancient Chinese poems, proverbs, abbreviations, and Internet terms, etc.; this embodiment can determine at least one standard word that has the same semantics as the designated word but different expressions according to the semantics expressed by the designated word. The standard word is a word with a more standard expression and conforming to the cognition of machines and most people. For example, the designated word "round-headed and round-brained" can be mapped to the following standard words: "round face, chubby". Establish a mapping relationship between the designated word and the standard word, and record the designated word and the standard word with the mapping relationship in the text mapping table.

[0086] In the actual application of prompt optimization, the initial prompt is matched with the pre-set text mapping table. When the target designated word in the initial prompt matches the target designated word in the text mapping table, the target designated word in the initial prompt is mapped to the target standard word. Exemplarily, the initial prompt input by the user is "A cute tabby cat with an anime style and a round-headed and round-brained appearance is sunbathing on a lawn full of flowers", and the "round-headed and round-brained" in it is a kind of slang and common saying, which belongs to the designated word recorded in the text mapping table. Therefore, according to the mapping relationship, the target standard word "round face, chubby" is used to replace the target designated word "round-headed and round-brained", and the above initial prompt is converted into "A cute tabby cat with an anime style, a round face, and a chubby appearance is sunbathing on a lawn full of flowers".

[0087] This embodiment uses the text mapping table to map some target designated words with special and difficult-to-understand expressions to target standard words with standard expressions and low understanding difficulty, which can reduce the difficulty of prompt processing, improve the accuracy of the prompt, help the text-to-image model better understand the prompt, avoid understanding deviations, and thus can generate the desired image more accurately.

[0088] S106. Determine rich words that match the target subject and / or target style type; wherein, the rich words include: words used to supplement the content of the initial prompt.

[0089] In some embodiments, a text optimization framework can be obtained based on text-to-image models such as different versions of Stable Diffusion and SDXL. The text optimization framework sorts out the subject and art style types, and provides multiple supplementary, more refined rich words for the prompt representing the subject and the prompt representing the art style type. Under this text optimization framework, determine rich words that match the target subject and / or target style type to optimize and supplement the initial prompt in terms of the subject and / or art style through the rich words.

[0090] S108. Supplement rich vocabulary to the initial prompt to obtain an optimized prompt. In this embodiment, rich vocabulary matching the target subject and / or rich vocabulary matching the target style type is supplemented to the initial prompt, so that the optimized prompt can express more fully and in detail in terms of the subject and artistic style, realizing the optimization process in terms of the subject and artistic style.

[0091] For better understanding, the following embodiments will describe the above steps S106 and S108 in detail.

[0092] In one embodiment, determining rich vocabulary matching the target style type may include: determining at least one candidate artistic style type different from the target style type and matching the target subject; and determining the vocabulary representing the candidate artistic style type as rich vocabulary matching the target style type.

[0093] For example, for Figure 2 the target subject of the tiger shown, in addition to being able to adapt to the extracted target style type, it can also adapt to other artistic style types. Based on this, this embodiment can supplement more rich and detailed vocabulary for the artistic style. Thus, multiple candidate artistic style types such as realistic style, comic style, and watercolor style that are different from the target style type and match the tiger are determined, and the vocabulary representing each of the above candidate artistic style types is determined as rich vocabulary matching the target style type. Accordingly, the above rich vocabulary representing the candidate artistic style type is used as a supplement to the initial prompt and added to the initial prompt to obtain an optimized prompt. Furthermore, when using the optimized prompt in terms of artistic style to generate an image, multiple different artistic styles can be adapted to the same target subject, producing rich and diverse artistic effects.

[0094] In another embodiment, determining rich vocabulary matching the target style type may include: under the specified target style type, determining the modification information for describing the target subject, where the modification information includes: nature, state, feature, and / or attribute; and determining the vocabulary representing the modification information as rich vocabulary matching the target subject.

[0095] Under some specified target style types, more abundant and refined vocabulary can be supplemented for the detailed modification information of the target subject. For example, for a tiger in a realistic style, the more detailed modification information used to describe the tiger can be determined, such as: the attribute of the tiger is a cub or an adult tiger, the state of the tiger is lying or standing, the coat color and stripes in the external features of the tiger, whether the picture containing the tiger is set to high definition, colorful, etc.; the vocabulary representing the above modification information is determined as the rich vocabulary matching the target subject. In another example, for the target subject of a person, rich vocabulary representing modification information such as gender, clothing, hair color, etc. can be determined. Correspondingly, the above rich vocabulary representing modification information is used as a supplement to the initial prompt word and added to the initial prompt word to obtain an optimized prompt word. Furthermore, when using the optimized prompt word on the detailed modification information of the target subject to generate an image, the detailed presentation of the picture can be higher and more accurate. At the same time, in this embodiment, by supplementing rich vocabulary for describing modification information to the target subject under the specified target style type, it can be ensured that the added rich vocabulary can match the current target style type (such as the realistic style), and avoid the appearance of abstract, distorted, and other words that are significantly inconsistent with the realistic style, and words that are significantly inconsistent with the high-resolution and pixel art styles, etc.

[0096] It can be understood that the above-mentioned multiple embodiments can be combined with each other to perform a more comprehensive, detailed, and accurate optimization process on the initial prompt word.

[0097] In another embodiment, the rich vocabulary further includes: vocabulary associated with a preset text-to-image model; the determination of the rich vocabulary matching the target subject and / or the target style type may include: determining a first text-to-image model matching the target subject and / or a second text-to-image model matching the target style type; determining the vocabulary associated with the first text-to-image model as the rich vocabulary matching the target subject; determining the vocabulary associated with the second text-to-image model as the rich vocabulary matching the target style type.

[0098] In this embodiment, for some specific subjects (such as certain vehicle brands) and specific art style types (such as Chinese characteristic art style types), different models are trained respectively to improve the performance of the generated images. For example, for some subjects, a Lora (Low-Rank Adaptation) model is matched as the first text-to-image model, and for some art style types, a Dreambooth model is matched as the second text-to-image generation model. After the target subject and the target style type are extracted, the Lora model matching the target subject is mounted on the target subject, and the Dreambooth model matching the target style type is mounted on the target style type.

[0099] On this basis, a first text-to-image model matching the target subject is determined. For example, the first text-to-image model includes the above-mentioned Lora model. The vocabulary associated with the first text-to-image model is determined as the rich vocabulary matching the target subject, and the rich vocabulary is supplemented into the initial prompt. Thus, the optimized prompt obtained can be adapted to the first text-to-image model.

[0100] A second text-to-image model matching the target style type is determined. For example, the second text-to-image model includes the above-mentioned Dreambooth model. The vocabulary associated with the second text-to-image generation module is determined as the rich vocabulary matching the target style type, and the rich vocabulary is supplemented into the initial prompt. Thus, the optimized prompt obtained can be adapted to the second text-to-image model.

[0101] In practical applications, due to the technical characteristics of some Lora models and Dreambooth models, specific vocabulary is required to activate the models and make them effective. Therefore, the initial prompt needs to be optimized and supplemented so that the optimized prompt can be adapted to the models. For example, the initial prompt is "Aa", which represents the model number of a certain car. However, the Lora model cannot directly use "Aa" to generate images of relevant vehicles. In this case, according to the preset vocabulary "car" of the Lora model, the vocabulary "car" is supplemented to "Aa" to obtain an effective prompt that can be adapted to the Lora model: "Aa, car". This prompt is adapted to the Lora model and can make the Lora model effective, thereby generating images of vehicles with the model number Aa.

[0102] Moreover, for common artistic styles such as realism, oil painting, and comics, the Dreambooth model can generally be directly used; however, for specific and rare artistic styles such as Dunhuang mural art style, the Dreambooth model generally needs to add some additional vocabulary to be used. Based on this, in this embodiment, some vocabulary that can activate the above-mentioned specific artistic style types is pre-configured for the Dreambooth model, which is determined as the rich vocabulary matching the target style type, and the initial prompt is optimized and supplemented with the rich vocabulary accordingly, so that the optimized prompt can be adapted to the Dreambooth model. Thus, the Dreambooth model can use the optimized prompt to generate images in the Dunhuang mural art style.

[0103] According to the above embodiments, the prompts representing specific subjects and specific artistic style types can be optimized to expand the scope of application of the prompts.

[0104] According to the above embodiment, after the above step S108, i.e. obtaining the optimized prompt word, the method provided by the embodiment of the present disclosure may further include: reviewing and editing the optimized prompt word to obtain the target prompt word. The editing includes but is not limited to: adding, replacing, deleting and / or correcting errors.

[0105] In one embodiment, editing includes: adding tags, deleting and / or replacing; accordingly, reviewing and editing the optimized prompt words may include: reviewing the optimized prompt words according to a pre-established sample list including drawing texts, and editing the optimized prompt words according to the review results. Among them, the hot updated sample list is used to record drawing texts of different levels such as recommended use, conditional use and not allowed use in the image generation process; drawing texts that are conditionally restricted and not allowed to be used can be called negative sample texts, such as: taboo texts involving illegal and irregular texts, and texts that are prohibited due to copyright management and other conditions. Recommended drawing texts can be called positive sample texts, such as: label texts used to supplement detail modifiers such as attributes. Some of the above drawing texts can be pre-set with labels.

[0106] Based on the sample list including: drawing texts that are subject to conditional use and those that are not allowed to be used, edit the optimized prompt words according to the review results. Please refer to the following examples for details.

[0107] Any prompt word among the optimized prompt words is used as the current prompt word.

[0108] When the current prompt word matches the drawing text with a preset label in the sample list, the label of the drawing text is added to the current prompt word. For example, the current prompt word is "dragon", and the user hopes that the Wensheng graph model will generate a Chinese dragon instead of a Western dragon by default, so the drawing text describing the dragon in the sample list is set with a label with the attribute of Chinese dragon. When the current prompt word matches the drawing text describing the dragon in the sample list, a label will be added to the current prompt word, that is, the label of "Chinese dragon" is added to the current prompt word "dragon".

[0109] When the current prompt word matches the drawing text that is not allowed in the sample list, the current prompt word will be deleted. When the current prompt word matches the sensitive text such as nudity, pornography, violence, politics, etc. in the sample list, the current prompt word will be deleted.

[0110] When the current prompt word matches the drawing text in the sample list that is subject to conditional use, the current prompt word is replaced. Regarding the understanding of conditional use, for example, for works that have applied for copyright registration, in order to avoid infringement, copyright-managed text can be included in the sample list. Copyright-managed text can include: text indicating that unauthorized use of the work is prohibited, and can also include text of the work that is authorized for use. It can be understood that the above copyright registration is only an example of conditional use, and there may be other situations in actual applications.

[0111] In a specific embodiment, the current prompt word is Mickey Mouse, and the drawing text that is used under conditional restrictions in the sample list is matched to: Mickey Mouse is an IP class subject object that is prohibited from use without authorization. In this case, the current prompt word is replaced, and Mickey Mouse is replaced with a cartoon mouse without copyright restrictions. Alternatively, the current prompt word is Mickey Mouse, and the drawing text that is used under conditional restrictions in the sample list is matched to: Mickey Mouse is an IP class subject object that is prohibited from use without authorization, and Jerry Mouse is an IP class subject object that can be used with authorization. In this case, the current prompt word is replaced, and Mickey Mouse is replaced with an authorized Jerry Mouse.

[0112] This embodiment can optimize the prompt words more accurately and more conducive to model generation in real time and in a distributed manner through the sample list, and can also avoid the generation of illegal and irregular content to a certain extent.

[0113] In one embodiment, editing includes error correction; accordingly, reviewing and editing the optimized prompt words may include: inputting the optimized prompt words into a pre-trained second text convolutional neural network; and reviewing and removing erroneous conflicting prompt words in the optimized prompt words through the second text convolutional neural network.

[0114] For example, generating a pixel-style image requires a prompt word of "low resolution", while the globally optimized prompt words include prompt words such as "high resolution" and "high detail". Obviously, the "low resolution" required for the pixel style is a conflicting prompt word with "high resolution" and "high detail". Based on this, for conflicting prompt words, this embodiment outputs a value for each prompt word based on the second text convolutional neural network to review whether it is a conflicting word that needs to be deleted, and performs text vocabulary error correction based on the output result, eliminating the wrong conflicting prompt words in the optimized prompt words. Eliminating the conflicting prompt words can effectively avoid the problem of confusion caused by the text-generated graph model.

[0115] In another embodiment, for the sentence editing of error correction, the review and editing of the optimized prompt words can also be implemented in the following manner:

[0116] Input the optimized prompt into a pre-trained Seq2Seq model; extract the context feature vector of the optimized prompt through the encoder of the Seq2Seq model; use the attention mechanism model in the decoder of the Seq2Seq model to decode the context feature vector, and review and eliminate the incorrect conflicting prompts in the optimized prompt according to the decoding result.

[0117] In this embodiment, the Seq2Seq model is used to learn to eliminate incorrect conflicting prompts from the optimized prompt. The Seq2Seq model is divided into two parts: an encoder and a decoder. The encoder part is used to obtain the high-level features of the optimized prompt, that is, the context feature vector. The decoder part uses the attention mechanism model, specifically an LSTM recurrent neural network model with an attention mechanism; during decoding, corresponding representations can be generated according to the context feature vector, so as to more accurately review and eliminate incorrect conflicting prompts and obtain a higher accuracy rate.

[0118] According to the above embodiments, this embodiment provides a method for optimizing prompts for text-to-image generation as Figure 3 shown, specifically including the following steps:

[0119] (1) Obtain the initial prompt input by the user;

[0120] (2) Extract the target subject and target style type matching the initial prompt through the first text convolutional neural network;

[0121] (3) Determine at least one candidate art style type different from the target style type and matching the target subject, and determine the vocabulary representing the candidate art style type as the rich vocabulary matching the target style type;

[0122] (4) Under the specified target style type, determine the modification information for describing the target subject; determine the vocabulary representing the modification information as the rich vocabulary matching the target subject;

[0123] (5) Determine the vocabulary associated with the Lora model as the rich vocabulary matching the target subject;

[0124] (6) Determine the vocabulary associated with the Dreambooth model as the rich vocabulary matching the target style type;

[0125] (7) Supplement the above rich vocabulary into the initial prompt to obtain the optimized prompt;

[0126] (8) Review the optimized prompt according to the pre-established sample list and perform editing such as adding tags, deleting, and replacing;

[0127] (9) Using the second text convolutional neural network, the incorrect conflicting prompt words in the optimized prompt words are reviewed and eliminated.

[0128] The final target prompt word is obtained through the above steps.

[0129] In summary, the prompt word optimization method for text-generated images provided in the embodiment of the present disclosure includes: first, obtaining the initial prompt word input by the user; second, extracting the target subject and the target style type based on the initial prompt word; then determining the rich vocabulary that matches the target subject and / or the target style type; wherein the rich vocabulary includes: vocabulary used to supplement the content of the initial prompt word; and, supplementing the rich vocabulary to the initial prompt word to obtain the optimized prompt word. This technical solution will optimize the prompt word by enriching the vocabulary on the subject and style type of the image for the initial prompt word input by the user, by extracting the target subject and the target style type, so that the optimized prompt word can be more fully and refined in terms of the subject and artistic style; then, the accuracy of the prompt word can be improved by editing such as adding tags, deleting and replacing. Therefore, this solution can improve the richness, refinement and accuracy of the optimized target prompt word.

[0130] Furthermore, the present disclosure can provide an engineering solution for optimizing prompt words for different versions of the Stable Diffusion model and SDXL. The initial prompt words entered by the user are optimized in terms of subject and artistic style classification, and a certain degree of interception is performed on illegal and prohibited entries through review and editing such as adding tags, deletion and replacement. The initial prompt words are optimized and supplemented according to the vocabulary associated with the text graph model, and the prompt words can be specially optimized for the generation of Chinese artistic styles and specific themes (such as a certain brand of car).

[0131] Corresponding to the prompt word optimization method for generating images from text provided in the embodiment of the present disclosure, the embodiment of the present disclosure also provides a prompt word optimization device for generating images from text. Figure 4 The structural block diagram of the prompt word optimization device for generating images from text provided in the embodiment of the present disclosure is as follows: Figure 4 As shown, the prompt word optimization device for generating an image from text includes:

[0132] The prompt word acquisition module 210 is used to acquire the initial prompt word input by the user;

[0133] An extraction module 220, for extracting a target subject and a target style type based on the initial prompt word;

[0134] A vocabulary determination module 230, configured to determine rich vocabulary that matches the target subject and / or the target style type; wherein, the rich vocabulary includes: vocabulary for supplementing the content of the initial prompt;

[0135] A prompt optimization module 240, configured to supplement the rich vocabulary into the initial prompt to obtain an optimized prompt.

[0136] In some embodiments, the extraction module 220 is further configured to:

[0137] Input the initial prompt into a pre-trained first text convolutional neural network;

[0138] Extract a target subject and a target style type that match the initial prompt through the first text convolutional neural network.

[0139] In some embodiments, the first text convolutional neural network includes: a convolutional module with a residual structure, a fully connected module, a first output module, and a second output module;

[0140] The extraction module 220 is further configured to:

[0141] Extract context text features of the initial prompt through the convolutional module, and input the context text features into the fully connected module;

[0142] Sort the context text features in a preset order through the fully connected module to obtain a text feature vector, and input the text feature vector into the first output module and the second output module;

[0143] Classify the initial prompt for the subject through the first output module according to the text feature vector to obtain multiple candidate subjects and the weights of each candidate subject;

[0144] Smooth the weights of each candidate subject, and determine the target subject based on the smoothed weights;

[0145] Classify the initial prompt for the artistic style through the second output module according to the text feature vector to obtain multiple candidate style types and the scores of each candidate style type;

[0146] Smooth the scores of each candidate style type, and determine the target style type based on the smoothed scores.

[0147] In some embodiments, the vocabulary determination module 230 is further configured to:

[0148] Determine at least one candidate artistic style type that is different from the target style type and matches the target subject;

[0149] Determine the vocabulary representing the candidate art style type as the rich vocabulary that matches the target style type.

[0150] In some embodiments, the vocabulary determination module 230 is further configured to: under the specified target style type, determine the modification information for describing the target subject, where the modification information includes: nature, state, feature, and / or attribute;

[0151] Determine the vocabulary representing the modification information as the rich vocabulary that matches the target subject.

[0152] In some embodiments, the rich vocabulary further includes: vocabulary associated with a preset text-to-image model; the vocabulary determination module 230 is further configured to:

[0153] Determine a first text-to-image model that matches the target subject and / or a second text-to-image model that matches the target style type;

[0154] Determine the vocabulary associated with the first text-to-image model as the rich vocabulary that matches the target subject;

[0155] Determine the vocabulary associated with the second text-to-image model as the rich vocabulary that matches the target style type.

[0156] In some embodiments, the device further includes a review and editing module, which is configured to: review and edit the optimized prompt to obtain a target prompt.

[0157] In some embodiments, the review and editing module is further configured to:

[0158] Review the optimized prompt according to a pre-established sample list including drawing texts; the sample list at least includes: drawing texts restricted by conditions and not allowed to be used, and some of the drawing texts are labeled;

[0159] Take any prompt in the optimized prompt as the current prompt;

[0160] When the current prompt matches the drawing text with a preset label in the sample list, add the label of the drawing text to the current prompt;

[0161] When the current prompt matches the drawing text not allowed to be used in the sample list, delete the current prompt;

[0162] When the current prompt matches the drawing text restricted by conditions in the sample list, replace the current prompt.

[0163] In some embodiments, the review and editing module is further configured to:

[0164] Input the optimized prompt into a pre-trained second text convolutional neural network;

[0165] Review and eliminate incorrect conflicting prompts in the optimized prompt through the second text convolutional neural network.

[0166] In some embodiments, the review and editing module is further configured to:

[0167] Input the optimized prompt into a pre-trained Seq2Seq model;

[0168] Extract the context feature vector of the optimized prompt through the encoder of the Seq2Seq model;

[0169] Use the attention mechanism model in the decoder of the Seq2Seq model to decode the context feature vector, and review and eliminate incorrect conflicting prompts in the optimized prompt according to the decoding result.

[0170] In some embodiments, the apparatus further includes a mapping module, which is configured to:

[0171] Match the initial prompt with a preset text mapping table; wherein, the text mapping table is used to record the mapping relationship between designated words and standard words, and the designated words and the standard words are texts with the same semantics but different expressions;

[0172] When the initial prompt matches the target designated word in the text mapping table, map the target designated word in the initial prompt to the target standard word according to the mapping relationship to obtain the mapped initial prompt.

[0173] The prompt optimization device for generating images from text disclosed in the above embodiments can execute the prompt optimization method for generating images from text disclosed in the above embodiments, and has the same or corresponding beneficial effects. To avoid repetition, it will not be elaborated here.

[0174] Embodiments of the present disclosure further provide a computer-readable storage medium, which stores programs or instructions, and the programs or instructions cause a computer to execute the steps of any of the above methods.

[0175] Obtain an initial prompt input by a user;

[0176] Extract a target subject and a target style type based on the initial prompt;

[0177] Determine rich vocabulary that matches the target subject and / or the target style type; wherein, the rich vocabulary includes: vocabulary for supplementing the content of the initial prompt;

[0178] Supplement the rich vocabulary into the initial prompt to obtain an optimized prompt.

[0179] Optionally, when the computer-executable instructions are executed by a computer processor, they can also be used to execute the technical solutions of the above-mentioned method for optimizing the prompt for generating an image from text provided by the embodiments of the present disclosure, and achieve the corresponding beneficial effects.

[0180] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments of the present disclosure can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solutions of the embodiments of the present disclosure, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a floppy disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (FLASH), a hard disk, or an optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present disclosure.

[0181] The embodiments of the present disclosure further provide an electronic device, including: one or more processors; a memory for storing one or more programs or instructions; the processor is configured to execute the steps of any of the above methods by calling the programs or instructions stored in the memory, and achieve the corresponding beneficial effects.

[0182] Figure 5 It is a schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present disclosure. As Figure 5 shown, the electronic device includes one or more processors 301 and a memory 302.

[0183] The processor 301 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to execute desired functions.

[0184] The memory 302 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 301 may run the program instructions to implement the method for optimizing the prompt for generating an image from text in the embodiments of the present disclosure described above, and / or other desired functions. Various contents such as input signals, signal components, noise components, etc. may also be stored in the computer-readable storage media.

[0185] In one example, the electronic device may further include: an input device 303 and an output device 304, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0186] In addition, the input device 303 may further include, for example, a keyboard, a mouse, and the like.

[0187] The output device 304 may output various information to the outside, including the determined distance information, direction information, etc. The output device 304 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0188] Of course, for simplicity, Figure 5 only some of the components related to the present disclosure in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application scenarios, the electronic device may further include any other appropriate components.

[0189] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0190] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but rather will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for optimizing prompts for text-to-image generation, characterized in that, Including: Obtain an initial prompt entered by the user; Extract a target subject and a target style type based on the initial prompt; Determine rich vocabulary that matches the target subject and / or the target style type; wherein, the rich vocabulary includes: vocabulary for supplementing the content of the initial prompt; Supplement the rich vocabulary into the initial prompt to obtain an optimized prompt.

2. The method according to claim 1, characterized in that, The extracting the target subject and the target style type based on the initial prompt includes: Input the initial prompt into a pre-trained first text convolutional neural network; Extract the target subject and the target style type that match the initial prompt through the first text convolutional neural network.

3. The method according to claim 2, characterized in that, The first text convolutional neural network includes: a convolutional module with a residual structure, a fully connected module, a first output module, and a second output module; The extracting the target subject and the target style type that match the initial prompt through the first text convolutional neural network includes: Extract the context text features of the initial prompt through the convolutional module, and input the context text features into the fully connected module; Sort the context text features in a preset order through the fully connected module to obtain a text feature vector, and input the text feature vector into the first output module and the second output module; Perform subject classification on the initial prompt according to the text feature vector through the first output module to obtain multiple candidate subjects and the weights of each candidate subject; Smooth the weights of each candidate subject, and determine the target subject based on the smoothed weights; Perform art style classification on the initial prompt according to the text feature vector through the second output module to obtain multiple candidate style types and the scores of each candidate style type; Smooth the scores of each candidate style type, and determine the target style type based on the smoothed scores.

4. The method according to claim 1, characterized in that, The determining the rich vocabulary that matches the target style type includes: Determine at least one candidate art style type that is different from the target style type and matches the target subject; Determine the vocabulary representing the candidate art style type as the rich vocabulary that matches the target style type.

5. The method according to claim 1, characterized in that, The determining the rich vocabulary that matches the target subject includes: Under the specified target style type, determine the modification information for describing the target subject, and the modification information includes: nature, state, feature, and / or attribute; Determine the vocabulary representing the modification information as the rich vocabulary that matches the target subject.

6. The method according to claim 1, characterized in that, The rich vocabulary further includes: vocabulary associated with a preset text-to-image model; the determining the rich vocabulary that matches the target subject and / or the target style type includes: Determine a first text-to-image model that matches the target subject and / or a second text-to-image model that matches the target style type; Determine the vocabulary associated with the first text-to-image model as the rich vocabulary that matches the target subject; Determine the vocabulary associated with the second text-to-image model as the rich vocabulary that matches the target style type.

7. The method according to claim 1, characterized in that, After supplementing the rich vocabulary into the initial prompt to obtain an optimized prompt, the method further includes: Reviewing and editing the optimized prompt to obtain a target prompt.

8. The method according to claim 7, characterized in that, The reviewing and editing of the optimized prompt includes: Reviewing the optimized prompt according to a pre-established sample list including drawing texts; the sample list at least includes: drawing texts restricted by conditions and not allowed to be used, and some of the drawing texts are labeled; Taking any prompt in the optimized prompt as the current prompt; When the current prompt matches the drawing text with a preset label in the sample list, adding the label of the drawing text to the current prompt; When the current prompt matches the drawing text not allowed to be used in the sample list, deleting the current prompt; When the current prompt matches the drawing text restricted by conditions in the sample list, replacing the current prompt.

9. The method according to claim 7, wherein The reviewing and editing of the optimized prompt includes: Inputting the optimized prompt into a pre-trained second text convolutional neural network; Reviewing and removing incorrect conflicting prompts in the optimized prompt through the second text convolutional neural network.

10. The method according to claim 7, wherein The reviewing and editing of the optimized prompt includes: Inputting the optimized prompt into a pre-trained Seq2Seq model; Extracting the context feature vector of the optimized prompt through the encoder of the Seq2Seq model; Using the attention mechanism model by the decoder of the Seq2Seq model to decode the context feature vector, and reviewing and removing incorrect conflicting prompts in the optimized prompt according to the decoding result.

11. The method according to claim 1, wherein After extracting the target subject and target style type based on the initial prompt, the method further includes: Matching the initial prompt with a preset text mapping table; wherein, the text mapping table is used to record the mapping relationship between specified words and standard words, and the specified words and the standard words are texts with the same semantics but different expressions; When the initial prompt matches the target specified word in the text mapping table, mapping the target specified word in the initial prompt to the target standard word according to the mapping relationship to obtain the mapped initial prompt.

12. A prompt optimization device for generating images from text, wherein Including: A prompt acquisition module, configured to acquire an initial prompt input by a user; An extraction module, configured to extract a target subject and a target style type based on the initial prompt; A vocabulary determination module, configured to determine rich vocabulary that matches the target subject and / or the target style type; wherein, the rich vocabulary includes: vocabulary for supplementing the content of the initial prompt; A prompt optimization module, configured to supplement the rich vocabulary into the initial prompt to obtain an optimized prompt.

13. A computer-readable storage medium, wherein The computer-readable storage medium stores a program or instructions, and the program or instructions cause a computer to execute the steps of the method according to any one of claims 1 to 11.

14. An electronic device, wherein Including: One or more processors; A memory for storing one or more programs or instructions; The processor is configured to execute the steps of the method according to any one of claims 1 to 11 by calling the programs or instructions stored in the memory.

Citation Information

Cited By

  • Image processing method and device based on large model, medium, electronics and product

    CN121392544A