Method and device for prompting word expansion, computer device and storage medium

By acquiring initial prompt words and expanding them using a prompt word generation model, the problem of prompt words in text-to-image generation models failing to meet user needs and achieve batch automated production is solved. This provides a more diverse and flexible selection of prompt words and enables the automatic acquisition of high-quality prompt words.

CN117764068BActive Publication Date: 2026-01-13DOUYIN VISION CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311789659.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2026-01-13
Estimated Expiration
2043-12-22

AI Technical Summary

Technical Problem

Existing text-to-image models cannot meet user needs and the requirements for automated batch production. They cannot effectively match users' customized needs and cannot provide high-quality prompts for large batches.

Method used

By acquiring initial prompt words, selecting a prompt word expansion mode, and obtaining reference prompt words related to the initial prompt words according to the expansion strategy, the prompt word generation model is used for expansion, providing more diverse and free choices to meet user needs.

Benefits of technology

It achieves close alignment with user suggestion requirements, lowers the threshold for suggestion writing, automatically acquires a large number of high-quality suggestions, and meets the needs of batch automated production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117764068B_ABST
    Figure CN117764068B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of computer, in particular to a method and device for prompt word expansion, computer equipment and storage medium, the method comprising: obtaining an initial prompt word, wherein the initial prompt word is related to a target image to be obtained by a target user; selecting a prompt word expansion mode, and determining a target initial prompt word from the initial prompt word according to the number of prompt words corresponding to the prompt word expansion mode; and obtaining a reference prompt word related to the target initial prompt word according to an expansion strategy corresponding to the target initial prompt word and the prompt word expansion mode, wherein the reference prompt word is an expanded prompt word of the target initial prompt word. The present disclosure can closely align with the prompt word demand of the target user, provide more diverse and free choices for the target user, greatly reduce the threshold of prompt word writing, and automatically obtain a large number of high-quality prompt words.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and specifically to methods, apparatus, computer devices, and storage media for expanding prompt words. Background Technology

[0002] With the development of large language models, models for generating images from text have emerged. The principle behind text-to-image generation is to encode the input text into vector representations, and then use a deep learning model to convert these vector representations into corresponding images.

[0003] In common text-to-image generation scenarios, prompt word engineering always plays a crucial role. However, many users initially face a high learning curve and difficulty in quickly getting started. While many prompt word plugins exist, they often fail to effectively match users' customized needs or meet the demands of large-scale production of high-quality prompt words. Furthermore, existing text-to-image generation models also cannot satisfy user requirements and the need for automated batch production.

[0004] Therefore, the prompts in current text-to-image generation models have technical problems that fail to meet user needs and the requirements for batch automated production. Summary of the Invention

[0005] In view of this, the present disclosure provides a method, apparatus, computer device and storage medium for expanding prompt words, in order to solve the technical problem that the prompt words of the current text-to-image model cannot meet the needs of users and the needs of batch automated production.

[0006] Firstly, this disclosure provides a method for expanding prompt words, the method comprising:

[0007] Obtain initial prompt words, where the initial prompt words are related to the target image to be acquired by the target user;

[0008] Select the prompt word expansion mode, and determine the target initial prompt word from the initial prompt words based on the number of prompt words corresponding to the prompt word expansion mode;

[0009] Based on the target initial prompt word and the expansion strategy corresponding to the prompt word expansion mode, obtain reference prompt words related to the target initial prompt word, where the reference prompt words are the expanded prompt words of the target initial prompt word.

[0010] In the embodiments of the present disclosure, by obtaining the initial prompt word and the selected prompt word expansion model, the target initial prompt word is determined from the obtained initial prompt word according to the number of prompt words corresponding to the selected prompt word expansion model, and then the reference prompt word obtained by expanding the target initial prompt word is obtained according to the target initial prompt word and the expansion strategy corresponding to the prompt word expansion model. In this way, the embodiments of the present disclosure can closely align the prompt word needs of the target user, provide more diverse and free choices for the target user, greatly reduce the prompt word writing threshold, and automatically obtain a large number of high-quality prompt words, thereby solving the technical problems in the related art that the prompt words cannot meet the user needs and batch automatic production needs.

[0011] In an optional implementation, obtaining the initial prompt word comprises:

[0012] Obtaining the subject keyword input by the target user for describing the target image;

[0013] Analyzing the subject keyword to obtain an expansion keyword;

[0014] Obtaining the initial prompt word according to the subject keyword and the expansion keyword.

[0015] In the embodiments of the present disclosure, the subject keyword input by the target user is reasonably supplemented, so that the target image is described more richly, and the target user can quickly and conveniently accumulate the initial prompt word.

[0016] In an optional implementation, obtaining the initial prompt word comprises:

[0017] Obtaining a reference image input by the target user;

[0018] Analyzing the reference image to obtain the initial prompt word describing the reference image.

[0019] In the embodiments of the present disclosure, the reference image input by the target user is analyzed, so that the generated initial prompt word can represent a more accurate target image, and the target user can quickly and conveniently accumulate the initial prompt word.

[0020] In an optional implementation, in a case where the prompt word expansion mode is a first mode, obtaining the reference prompt word related to the target initial prompt word according to the target initial prompt word and the expansion strategy corresponding to the prompt word expansion mode comprises:

[0021] The target initial prompt word is multiple, and the reference prompt word is obtained by the prompt word generation model according to the multiple target initial prompt words.

[0022] A preset number of reference prompt words are obtained, wherein the correlation degree between the reference prompt words is less than a correlation degree threshold.

[0023] In the embodiments of the present disclosure, the number of target initial prompt words is determined according to the prompt word expansion mode selected by the target user, and then the reference prompt words meeting the expansion strategy are obtained through the prompt word generation model, so as to closely align the prompt word demand of the target user.

[0024] In an optional implementation, in the case where the prompt word expansion mode is the second mode, the reference prompt words related to the target initial prompt word are obtained according to the target initial prompt word and the expansion strategy corresponding to the prompt word expansion mode, including:

[0025] The target initial prompt word is one, the target initial prompt word is split to determine a plurality of attributes of the target initial prompt word and a target sub-prompt word corresponding to each attribute;

[0026] In response to the selection of the target attribute, the other sub-prompt words are retained, and the target sub-prompt word corresponding to the target attribute is replaced through the prompt word generation model;

[0027] Based on the replaced target sub-prompt word and the other sub-prompt words, the reference prompt words related to the target initial prompt word are obtained.

[0028] In the embodiments of the present disclosure, the number of target initial prompt words is determined according to the prompt word expansion mode selected by the target user, then a plurality of attributes are obtained by splitting the target initial prompt word, the target sub-prompt word corresponding to the target attribute is replaced through the prompt word generation model in response to the selection of the target attribute by the target user, the other sub-prompt words are retained, and then the reference prompt words meeting the expansion strategy are obtained, so as to closely align the prompt word demand of the target user.

[0029] In an optional implementation, in the case where the prompt word expansion mode is the third mode, the reference prompt words related to the target initial prompt word are obtained according to the target initial prompt word and the expansion strategy corresponding to the prompt word expansion mode, including:

[0030] When the target initial prompt word is one and contains a plurality of sub-prompt words, in response to the selection of the target sub-prompt word, the target sub-prompt word is retained, and the other sub-prompt words are replaced through the prompt word generation model;

[0031] Based on the target sub-prompt word and the replaced other sub-prompt words, the reference prompt words related to the target initial prompt word are obtained.

[0032] In the embodiments of the present disclosure, the number of target initial prompt words is determined according to the prompt word expansion mode selected by the target user, then the other sub-prompt words are replaced through the prompt word generation model in response to the selection of the target sub-prompt word needed to be retained by the target user, and then the reference prompt words meeting the expansion strategy are obtained, so as to closely align the prompt word demand of the target user.

[0033] In an optional implementation, the initial prompt word is obtained, including:

[0034] The initial prompt word in the prompt word generation model is obtained, wherein the prompt word generation model is a model obtained by initializing the initial prompt word generation model based on a randomness parameter, a diversity parameter and an association parameter, the randomness parameter and the diversity parameter are parameters to be adjusted by the target user when setting the initial prompt word generation model, and the association parameter is a parameter required when the reference prompt word is loaded in the background.

[0035] In the embodiments of the present disclosure, the initial prompt word generation model is initialized by obtaining the settings of the front-end parameters by the target user and the settings of the association parameters by the background, so that more diverse and free choices are provided for the target user when the initial prompt word is obtained.

[0036] In an optional implementation, the method further includes:

[0037] The first quantity of reference prompt words is obtained, wherein the first quantity of reference prompt words is determined according to a batch value in the association parameter;

[0038] The feedback information of the first quantity of reference prompt words is sent;

[0039] The second quantity of reference prompt words is obtained, wherein the second quantity of reference prompt words is determined according to a quantity of reference prompt words generated per request and a request collection frequency per time in the association parameter.

[0040] In the embodiments of the present disclosure, by reasonably setting the obtained association parameter, a large quantity of high-quality prompt words are obtained.

[0041] In a second aspect, the present disclosure provides a prompt word expansion device, which includes:

[0042] The first obtaining module is configured to obtain an initial prompt word, wherein the initial prompt word is related to a target image to be obtained by a target user;

[0043] The determining module is configured to select a prompt word expansion mode, and determine a target initial prompt word from the initial prompt word according to a prompt word quantity corresponding to the prompt word expansion mode;

[0044] The second obtaining module is configured to obtain a reference prompt word related to the target initial prompt word according to an expansion strategy corresponding to the prompt word expansion mode and the target initial prompt word, wherein the reference prompt word is an expansion prompt word of the target initial prompt word.

[0045] In a third aspect, the present disclosure provides a computer device, comprising: a memory and a processor, which are connected with each other in communication, and the memory stores computer instructions, and the processor executes the computer instructions to perform the method of the prompt word expansion of the first aspect or any of the corresponding embodiments.

[0046] In a fourth aspect, the present disclosure provides a computer readable storage medium, which stores computer instructions for making a computer execute the method of the prompt word expansion of the first aspect or any of the corresponding embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0047] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the drawings needed in the specific embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present disclosure, and other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0048] Figure 1 is a flowchart of the method of the prompt word expansion according to some embodiments of the present disclosure;

[0049] Figure 2 is a flowchart of the prompt word interactive learning mode according to some embodiments of the present disclosure;

[0050] Figure 3 is a flowchart of the prompt word interactive learning mode according to some other embodiments of the present disclosure;

[0051] Figure 4 is a flowchart of the prompt word interactive learning mode according to some other embodiments of the present disclosure;

[0052] Figure 5 is a flowchart of the method of the prompt word expansion according to some embodiments of the present disclosure;

[0053] Figure 6 is a flowchart of the method of the prompt word expansion according to some embodiments of the present disclosure;

[0054] Figure 7 is a structural block diagram of the device of the prompt word expansion according to some embodiments of the present disclosure;

[0055] Figure 8 is a hardware structure diagram of the computer device of the present disclosure. DETAILED DESCRIPTION

[0056] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0057] In common text-to-image generation scenarios, prompt word engineering always plays a crucial role. However, existing prompt word plugins cannot effectively match users' customized needs in practical use, nor can they meet the large-scale production tasks of high-quality prompt words. Furthermore, some existing text-to-image generation models also fail to meet user needs and the requirements for automated batch production. To address these issues, according to an embodiment of this disclosure, a method for prompt word expansion is provided. It should be noted that the steps shown in the flowcharts can be executed in a computer system such as a set of computer-executable instructions. Also, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0058] This embodiment provides a method for expanding prompt words. Figure 1 This is a flowchart of a method for expanding prompt words according to an embodiment of this disclosure, such as... Figure 1 As shown, this method can be applied to the front-end client, and the method flow includes the following steps:

[0059] Step S101: Obtain initial prompt words, wherein the initial prompt words are related to the target image to be obtained by the target user.

[0060] Optionally, with the development of models, models for text-to-text and text-to-image generation have emerged.

[0061] Specifically, in the process of generating an image from text, it is necessary to first rely on prompts input by the user, and then output the corresponding image based on these prompts. In this embodiment of the disclosure, the client can obtain the initial prompts input by the target user to describe the target image to be obtained, or it can obtain the initial prompts from a text-to-text model, such as a prompt generation model. It should be noted that the number of initial prompts can be one or multiple. In addition, the target user can be a specific user or a user cluster.

[0062] Step S102: Select the prompt word expansion mode, and determine the target initial prompt word from the initial prompt words according to the number of prompt words corresponding to the prompt word expansion mode.

[0063] Optionally, the client acquires the prompt word expansion mode selected by the target user on the terminal screen, acquires the number of prompt words corresponding to the selected prompt word expansion mode, and then obtains target initial prompt words meeting the number of prompt words from the initial prompt words. For example, the prompt word expansion mode is the first mode (such as the multi-example sentence mode), and the number of prompt words corresponding to the first mode is multiple, such as 3-100. In this case, the number of target initial prompt words obtained from the initial prompt words is 3-100.

[0064] In step S103, a reference prompt word related to the target initial prompt word is acquired according to the target initial prompt word and an expansion strategy corresponding to the prompt word expansion mode, where the reference prompt word is an expanded prompt word of the target initial prompt word.

[0065] Optionally, the client acquires the reference prompt word related to the target initial prompt word according to the number of target initial prompt words and the expansion strategy corresponding to the currently selected prompt word expansion mode, where the reference prompt word is an expanded prompt word of the target initial prompt word. In addition, the expansion strategy can generally be the personalized demand of the target user for the number of prompt words and the homogeneity of results, and the personalized demand is combined with the number of target initial prompt words to acquire the reference prompt word related to the initial prompt word.

[0066] In the embodiments of the present disclosure, the initial prompt word and the selected prompt word expansion model are acquired, the target initial prompt word is determined from the acquired initial prompt word according to the number of prompt words corresponding to the selected prompt word expansion model, and then the reference prompt word expanded from the target initial prompt word is obtained according to the target initial prompt word and the expansion strategy corresponding to the prompt word expansion mode. In this way, the embodiments of the present disclosure can closely align with the prompt word demand of the target user, provide more diverse and free choices for the target user, greatly reduce the threshold of prompt word writing, automatically acquire a large number of high-quality prompt words, and solve the technical problems in the related art that the prompt words cannot meet the user demand and batch automatic production demand.

[0067] In some optional embodiments, the initial prompt word is acquired, including:

[0068] The subject keyword input by the target user for describing the target image is acquired.

[0069] The subject keyword is analyzed to acquire an expansion keyword.

[0070] The initial prompt word is obtained according to the subject keyword and the expansion keyword.

[0071] Optionally, when the initial prompt word is selected, the initial prompt word can be obtained by interacting with an external device, such as a client and a database of a cloud server, or can be generated based on some auxiliary methods: based on a prompt word plug-in in the client, an extended writing method can be provided to help the target user quickly and conveniently accumulate the initial prompt word: the subject keyword input by the target user to describe the target image is obtained, for example, the target image is a portrait of a man with a beard, and the subject keyword can be a male painter with a beard.

[0072] After the subject keyword is obtained, the prompt word plug-in reasonably fills the subject keyword to obtain some extended keywords that depict the details of the target image to supplement the subject keyword, so that the subject keyword is more richly described and is more conducive to generating a picture that conforms to the internal logic and is more in line with the actual demand.

[0073] The vocabulary composed of the subject keyword and the extended keyword is referred to as the initial prompt word. The set including the subject keyword and the extended keyword can also be referred to as the initial prompt word.

[0074] In the embodiments of the present disclosure, the subject keyword input by the target user is reasonably supplemented, so that the target image is more richly described, and the target user is helped to quickly and conveniently accumulate the initial prompt word.

[0075] In some optional embodiments, obtaining the initial prompt word includes:

[0076] Obtaining a reference image input by the target user;

[0077] Analyzing the reference image to obtain an initial prompt word describing the reference image.

[0078] Optionally, the embodiments of the present disclosure propose an auxiliary method of generating an initial prompt word: the target user inputs some reference images that the target user considers to have a similarity greater than a first preset threshold (such as 0.9) to the target image to the client. It can be understood that the number of the reference images can be one or more, and the reference images are usually excellent images that the target user considers to be close to the target image.

[0079] At this time, the prompt word plug-in of the client analyzes the reference image to generate a description vocabulary of the reference image, and the description vocabulary is used as the initial prompt word.

[0080] In addition, for experienced users, the prompt word written or accumulated by the user can also be used as the initial prompt word.

[0081] In this embodiment of the disclosure, by analyzing the reference image input by the target user, the generated initial prompt words can represent a more accurate target image, helping the target user to quickly and easily accumulate initial prompt words.

[0082] In some optional implementations, when the prompt word expansion mode is in the first mode, reference prompt words related to the target initial prompt word are obtained based on the target initial prompt word and the expansion strategy corresponding to the prompt word expansion mode, including:

[0083] There are multiple initial target prompts. Based on these initial target prompts, a preset number of reference prompts are obtained through a prompt generation model. The correlation between the reference prompts is less than a correlation threshold.

[0084] Optionally, after the client obtains the initial target prompt words, it can perform interactive learning on the initial target prompt words based on the prompt word generation model to obtain expanded reference prompt words. Specifically, when the prompt word expansion mode selected by the target user is the first mode (such as the multi-example whole sentence mode), since there are multiple initial target prompt words corresponding to the first mode, based on the expansion strategy, such as the expansion strategy of the first mode is to avoid homogenization of the obtained reference prompt words, that is, to want the correlation between the various reference prompt words to be less than the correlation threshold (such as 0.7), the client will guide the prompt word generation model to learn the initial target prompt words and generate reference prompt words that conform to the above expansion strategy (i.e., non-homogeneous results, large-scale acquisition).

[0085] It should be noted that homogenization refers to a small range between reference prompts, resulting in results belonging to a similar small scope.

[0086] This disclosure addresses scenarios where the target user does not require precise prompts. For example, the initial prompts provided by the target user may not be limited to "beard" or "painter," and the obtained reference prompts may not necessarily contain a specific subject or specific vocabulary description; their diversity is rich.

[0087] like Figure 2 As shown, the process involves inputting initial target prompts; determining whether all initial target prompts have been input; if not, inputting more initial target prompts; if so, using a prompt generation model to learn these initial target prompts and obtain reference prompts. Additionally, after inputting initial prompts, positive examples can be selected as the learning objects for the prompt generation model. These positive examples are initial target prompts that more accurately and completely represent the target image.

[0088] In this embodiment of the disclosure, the number of initial target prompts is determined according to the prompt expansion mode selected by the target user, and then reference prompts that meet the expansion strategy are obtained through the prompt generation model, so as to achieve the effect of closely aligning with the prompt needs of the target user.

[0089] In some optional implementations, when the prompt word expansion mode is the second mode, reference prompt words related to the target initial prompt word are obtained based on the target initial prompt word and the expansion strategy corresponding to the prompt word expansion mode, including:

[0090] The initial target prompt is a single word. The initial target prompt is then broken down to determine multiple attributes of the initial target prompt and the sub-prompts corresponding to each attribute.

[0091] In response to the selection of the target attribute, other sub-cue words are retained, and the target sub-cue words corresponding to the target attribute are replaced by the cue word generation model;

[0092] Based on the replaced target sub-hint and other sub-hints, obtain reference hints related to the initial target hint.

[0093] Optionally, when the target user selects the second mode (which could be the intelligent splitting mode) for the prompt word expansion mode on the client, since the second mode corresponds to only one initial target prompt word, if the target user has a need to customize a target sub-prompt word corresponding to a specific target attribute among the multiple attributes contained in the initial target prompt word, the client will respond to the replacement operation performed by the target user on the client. Specifically, the client first splits the initial target prompt word, determines the multiple attributes of the initial target prompt word and the sub-prompt words corresponding to each attribute, and replaces the target sub-prompt words corresponding to the target attribute using the prompt word generation model based on the target user's selection of the target attribute. Then, it retains the other sub-prompt words that the target user did not perform a replacement operation on. The attributes contained in the initial target prompt word can be the gender attribute, clothing attribute, accessory attribute, etc., of the people contained in the target image. For example, if the target user wants to customize the target sub-prompt word corresponding to the target attribute (such as the clothing attribute), which was originally Hanfu, the target user customizes the target sub-prompt word to Hanfu, and then retains the other sub-prompt words in the initial target prompt word.

[0094] The replaced target sub-prompt and other sub-prompts are used as the updated target initial prompt, and the client then obtains relevant reference prompts based on the updated target initial prompt.

[0095] like Figure 3The process involves inputting an initial target prompt; intelligently splitting the initial target prompt into different attributes; customizing target sub-prompts corresponding to some or all of the attributes to be replaced; replacing the custom target sub-prompts corresponding to some or all of the attributes; and obtaining reference prompts based on the replaced and retained parts. Additionally, after inputting the initial target prompt, positive examples can be selected and used as the intelligent splitting objects. These positive examples are initial target prompts that more accurately and completely represent the target image.

[0096] In this embodiment of the disclosure, the number of initial target prompts is determined according to the prompt expansion mode selected by the target user. Then, the initial target prompts are split into multiple attributes. Based on the target user's selection of the target attributes, the target sub-prompts corresponding to the target attributes are replaced by the prompt generation model, while other sub-prompts are retained. In this way, reference prompts that meet the expansion strategy are obtained, thereby achieving the effect of closely aligning with the prompt requirements of the target user.

[0097] In some optional implementations, when the prompt word expansion mode is the third mode, reference prompt words related to the target initial prompt word are obtained based on the target initial prompt word and the expansion strategy corresponding to the prompt word expansion mode, including:

[0098] When there is only one target initial prompt word and it contains multiple sub-prompt words, in response to the selection of the target sub-prompt word, the target sub-prompt word is retained, and the other sub-prompt words are replaced by the prompt word generation model;

[0099] Based on the target sub-hint and the other sub-hints after replacement, obtain reference hints related to the target initial hint.

[0100] Optionally, when the target user selects the third mode (which can be a custom splitting mode) as the prompt word expansion mode on the client, since the number of target initial prompt words corresponding to the third mode is 1, and the target initial prompt word contains multiple sub-prompt words, if the target user has a specific target sub-prompt word that they want to keep, the client only needs to respond to the target sub-prompt word that the target user selects to keep on the client, and then use the prompt word generation model to replace the other sub-prompt words. After that, the replaced other sub-prompt words and the target sub-prompt word are used as the updated target initial prompt word, and the client can then obtain the relevant reference prompt words based on the updated target initial prompt word.

[0101] like Figure 4The process involves inputting the initial target prompt; obtaining the desired sub-prompts; retaining the target sub-prompts and replacing the remaining sub-prompts; and obtaining reference prompts based on the replaced and retained parts. Additionally, after inputting the initial target prompts, positive examples can be selected from them. These positive examples serve as the basis for the target user to determine whether to retain certain sub-prompts. The positive examples are initial target prompts that more accurately and completely represent the target image.

[0102] In this embodiment of the disclosure, the number of initial target prompt words is determined according to the prompt word expansion mode selected by the target user. Then, according to the target sub-prompt words that the target user needs to retain, other sub-prompt words are replaced by a prompt word generation model to obtain reference prompt words that meet the expansion strategy, thereby achieving the effect of closely aligning with the prompt word requirements of the target user.

[0103] In some alternative implementations, obtaining the initial prompt word includes:

[0104] Obtain the initial prompt words from the prompt word generation model. The prompt word generation model is the model obtained after initializing the initial prompt word generation model based on randomness parameters, diversity parameters, and association parameters. The randomness parameters and diversity parameters are the parameters to be adjusted by the target user when setting the initial prompt word generation model, and the association parameters are the parameters required when loading reference prompt words in the background.

[0105] Optionally, when obtaining the initial prompt word, the client in this embodiment of the disclosure can also obtain the initial prompt word from the prompt word generation model. Before obtaining the prompt word generation model, the initial prompt word generation model needs to be initialized. The initialization process mainly includes setting and adjusting parameters, loading parameters, and initializing the model, which is completed through collaboration between the front-end client and the back-end.

[0106] The client loads user-defined parameters for the target user (such as custom randomness and diversity parameters when generating reference prompts).

[0107] It should be noted that if the target user desires a greater diversity and richer variety of reference prompts, the values ​​of the randomness and diversity parameters can be increased. Conversely, if the target user desires a more stable variety of reference prompts, the values ​​of the randomness and diversity parameters can be decreased. Alternatively, some recommended values ​​can be provided to the target user for reference when setting these values. For example, the randomness parameter for model X could be set to 1.4, and the diversity parameter to 0.9, while the randomness parameter for model Y could be set to 0.7, and the diversity parameter to 0.9.

[0108] The backend loads reference prompts and requires associated parameters (including the initial prompt generation model type, batch size, number of reference prompts generated per request, number of requests per collection, etc.), loads / calls the selected initial prompt generation model, and guides the initialization of the initial prompt generation model.

[0109] In this embodiment of the disclosure, the initial prompt word generation model is initialized by obtaining the target user's settings for front-end parameters and the settings of existing related parameters in the back-end. This provides the target user with more diverse and free choices when obtaining the initial prompt words.

[0110] In some alternative implementations, the method further includes:

[0111] Obtain a first number of reference suggestion words, wherein the first number of reference suggestion words is determined based on the batch value in the associated parameters;

[0112] Send feedback information for the first number of reference prompts;

[0113] Obtain a second number of reference prompt words, wherein the second number of reference prompt words is determined based on the number of reference prompt words generated per request in the associated parameters and the number of requests collected in a single collection.

[0114] Optionally, after the client-side prompt interaction is completed and reference prompts are obtained, this embodiment of the disclosure can further align the prompt requirements of the target user through a generation-feedback-regeneration instruction fine-tuning mechanism. The client will first obtain a first number of reference prompts, which can be a small batch quantity, such as 10. It should be noted that the number of reference prompts generated in a single batch during this process is based on the batch size parameter of the associated parameter. The larger this parameter is set, the more feedback-regeneration rounds are usually required to meet the requirements in a single batch. However, once the requirements are aligned, the robustness of batch collection will be better. Therefore, the batch size value is preferably set in the range of 5-10.

[0115] After the target user obtains the first number of reference prompts through the client, if they approve the quality of the single batch generation, they will send feedback information, such as confirmation messages, to the backend. The backend can then begin batch collection, sending a second number of reference prompts to the client. This second number is the batch collection quantity, for example, 50. In this step, the backend in this embodiment performs batch collection based on two parameters: `minibatch` (the number of prompts generated per request) and `requests` (the number of requests per collection). The quantity collected per collection (i.e., the second number of reference prompts) = `minibatch * requests`. When the required collection quantity is fixed, to improve utilization and reduce costs, it is generally recommended to use a larger `minibatch`, provided the context length of the prompt generation model allows. This increases the number of prompts obtained per request, thereby reducing resource waste caused by repeated acquisition of contextual information in multiple requests.

[0116] In this embodiment of the disclosure, by reasonably setting the acquired related parameters, a large-scale acquisition of high-quality prompt words can be achieved.

[0117] In some alternative implementations, such as Figure 5 As shown, the complete main process of this embodiment, from start to finish, mainly includes the following steps: initialization of the initial prompt word generation model, selection of prompt word expansion mode, acquisition of target initial prompt words, interactive learning of prompt words, generation and evaluation of small batch results (if all results meet the requirements, proceed directly to the next step; if not all results meet the requirements, mark the non-compliant results, provide their serial numbers and a summary of the reasons, and generate a new batch of results until they meet the requirements before proceeding to the next step), batch collection, and return of plaintext results and ciphertext templates. Among these, as... Figure 6 As shown, after obtaining the encrypted template, the target user can directly enter a batch generation state that is completely identical to when the template was generated. Furthermore, before starting batch collection, the target user can choose whether to make further fine-tuning to accommodate any slight changes in their needs. After the target user decides to start batching, a new batch of plaintext results and a new encrypted template containing all the execution processes will be generated.

[0118] Based on the disclosed content of the above embodiments, this embodiment of the disclosure selects operation examples of Mode 1 (corresponding to the scenario of acquiring a large number of prompt words), Mode 2 (corresponding to the scenario of intelligent splitting to realize the replacement of target sub-prompt words corresponding to custom attributes), and Mode 3 (corresponding to the scenario of replacing other unretained sub-prompt words with target sub-prompt words that the target user has selected to retain) in prompt word interaction learning for further explanation:

[0119] Regarding Mode 1:

[0120] 1. Select initial prompt words to initialize the model;

[0121] 2. Randomness parameter settings (higher values ​​result in more randomness, default 1.5), diversity parameter settings (higher values ​​result in more diversity, default 0.8);

[0122] 3. Select Mode 1 for prompt word expansion mode;

[0123] 4. Select the initial prompt word for your input target;

[0124] 5. Enter multiple positive examples of the initial target prompt words in the text box, and repeat this step until all positive examples have been entered;

[0125] 6. Select "Start Generation" to obtain suggested keywords;

[0126] 7. Select Regenerate, evaluate the batch results, select the unsatisfactory result number, enter it in the text box, select Regenerate, and repeat this step until the batch results reach a satisfactory state;

[0127] 8. Select "Collect", enter the name of the results you want to save in the text box, and collect in batches.

[0128] Regarding Mode Two:

[0129] 1. Select initial prompt words to initialize the model;

[0130] 2. Randomness parameter settings (higher values ​​result in more randomness, default 1.5), diversity parameter settings (higher values ​​result in more diversity, default 0.8);

[0131] 3. Select Mode 2 for prompt word expansion mode;

[0132] 4. Select the initial prompt word for your input target;

[0133] 5. Enter a positive example of a target initial prompt word in the text box;

[0134] 6. Select Start Generation. Based on the intelligent splitting results, enter the target sub-suggestion word corresponding to the attribute to be replaced in the text box, retain other sub-suggestion words, and obtain the reference suggestion words;

[0135] 7. Select Regenerate, evaluate the batch results, select the unsatisfactory result number, enter it in the text box, select Regenerate, and repeat this step until the batch results reach a satisfactory state;

[0136] 8. Select "Collect", enter the name of the results you want to save in the text box, and collect in batches.

[0137] Regarding Mode 3:

[0138] 1. Select initial prompt words to initialize the model;

[0139] 2. Randomness parameter settings (higher values ​​result in more randomness, default 1.5), diversity parameter settings (higher values ​​result in more diversity, default 0.8);

[0140] 3. Select Mode 3 for the prompt word expansion mode;

[0141] 4. Select the initial prompt word for your input target;

[0142] 5. Enter a positive example of a target initial prompt word in the text box;

[0143] 6. Select "Start Generation", enter the target sub-keyword to be retained in the text box, replace the other sub-keywords that are not retained, and obtain the reference keyword;

[0144] 7. Select Regenerate, evaluate the batch results, select the unsatisfactory result number, enter it in the text box, select Regenerate, and repeat this step until the batch results reach a satisfactory state;

[0145] 8. Select "Collect", enter the name of the results you want to save in the text box, and collect in batches.

[0146] This embodiment also provides a device for expanding prompt words, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0147] This embodiment provides a device for expanding prompt words, such as... Figure 7 As shown, it includes:

[0148] The first acquisition module 701 is used to acquire initial prompt words, wherein the initial prompt words are related to the target image to be acquired by the target user;

[0149] The determination module 702 is used to select the prompt word expansion mode and determine the target initial prompt word from the initial prompt words according to the number of prompt words corresponding to the prompt word expansion mode;

[0150] The second acquisition module 703 is used to acquire reference prompts related to the target initial prompt based on the target initial prompt and the expansion strategy corresponding to the prompt expansion mode, wherein the reference prompts are the expanded prompts of the target initial prompt.

[0151] In some optional implementations, the first acquisition module 701 includes:

[0152] The first acquisition unit is used to acquire the main keywords input by the target user to describe the target image;

[0153] The second acquisition unit is used to analyze the main keywords and acquire expanded keywords;

[0154] The first unit is used to obtain initial suggestion words based on the main keywords and expanded keywords.

[0155] In some optional implementations, the first acquisition module 701 includes:

[0156] The third acquisition unit is used to acquire the reference image input by the target user;

[0157] The fourth acquisition unit is used to analyze the reference image and obtain initial prompt words describing the reference image.

[0158] In some optional implementations, when the prompt word expansion mode is the first mode, the second acquisition module 703 includes:

[0159] The fifth acquisition unit is used when there are multiple initial target prompt words. Based on these multiple initial target prompt words, a preset number of reference prompt words are obtained through a prompt word generation model. The correlation between the reference prompt words is less than a correlation threshold.

[0160] In some optional implementations, when the prompt word expansion mode is the second mode, the second acquisition module 703 includes:

[0161] The determination unit is used when there is only one target initial prompt word. The target initial prompt word is split into multiple attributes and the sub-prompt words corresponding to each attribute.

[0162] The first replacement unit is used to respond to the selection of the target attribute, retain other sub-cue words, and replace the target sub-cue words corresponding to the target attribute through the cue word generation model;

[0163] The sixth acquisition unit is used to acquire reference prompts related to the target initial prompt based on the replaced target sub-prompt and other sub-prompts.

[0164] In some optional implementations, when the prompt word expansion mode is the third mode, the second acquisition module 703 includes:

[0165] The second replacement unit is used when the initial target prompt word is one and contains multiple sub-prompt words. In response to the selection of the target sub-prompt word, the target sub-prompt word is retained, and other sub-prompt words are replaced by the prompt word generation model.

[0166] The seventh acquisition unit is used to acquire reference prompts related to the target initial prompt based on the target sub-prompt and the other replaced sub-prompts.

[0167] In some optional implementations, the first acquisition module 701 includes:

[0168] The eighth acquisition unit is used to acquire the initial prompt words in the prompt word generation model. The prompt word generation model is the model obtained after initializing the initial prompt word generation model based on randomness parameters, diversity parameters, and association parameters. The randomness parameters and diversity parameters are the parameters to be adjusted by the target user when setting the initial prompt word generation model, and the association parameters are the parameters required when loading reference prompt words in the background.

[0169] In some alternative embodiments, the device further includes:

[0170] The third acquisition module is used to acquire a first number of reference prompt words, wherein the first number of reference prompt words is determined based on the batch value in the associated parameters;

[0171] The sending module is used to send feedback information to the first number of reference prompt words;

[0172] The fourth acquisition module is used to acquire a second number of reference prompt words, wherein the second number of reference prompt words is determined based on the number of reference prompt words generated in each request in the associated parameters and the number of requests in a single collection.

[0173] In this embodiment, the device for expanding prompts is presented in the form of a functional unit. Here, a unit refers to an ASIC circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above-mentioned functions.

[0174] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0175] This disclosure also provides a computer device having the above-described features. Figure 7 The device shown expands the prompt words.

[0176] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of this disclosure, such as... Figure 8As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 8 Take a processor 10 as an example.

[0177] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0178] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.

[0179] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device as shown by a landing page for an app. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, which can be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0180] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0181] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.

[0182] This disclosure also provides a computer-readable storage medium in which the methods described in this disclosure can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium after being downloaded over a network. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium may be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium may also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code that, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0183] Although embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for expanding prompt words, characterized in that, The method includes: Obtain initial prompt words, wherein the initial prompt words are related to the target image to be acquired by the target user; Select a prompt word expansion mode, and determine the target initial prompt word from the initial prompt words based on the number of prompt words corresponding to the prompt word expansion mode; Based on the target initial prompt word and the expansion strategy corresponding to the prompt word expansion mode, reference prompt words related to the target initial prompt word are obtained, wherein the reference prompt words are expanded prompt words of the target initial prompt word; when the prompt word expansion mode is the second mode, obtaining reference prompt words related to the target initial prompt word based on the target initial prompt word and the expansion strategy corresponding to the prompt word expansion mode includes: The target initial prompt word is one, and the target initial prompt word is split to determine multiple attributes of the target initial prompt word and sub-prompt words corresponding to each attribute; In response to the selection of the target attribute, other sub-cue words are retained, and the target sub-cue words corresponding to the target attribute are replaced by the cue word generation model; Based on the replaced target sub-prompt word and the other sub-prompt words, obtain reference prompt words related to the target initial prompt word.

2. The method according to claim 1, characterized in that, The process of obtaining the initial prompt word includes: Obtain the main keywords input by the target user to describe the target image; Analyze the main keywords to obtain expanded keywords; The initial prompt words are obtained based on the main keywords and the expanded keywords.

3. The method according to claim 1, characterized in that, The process of obtaining the initial prompt word includes: Obtain the reference image input by the target user; The reference image is analyzed to obtain the initial prompt words describing the reference image.

4. The method according to claim 1, characterized in that, When the prompt word expansion mode is the first mode, the step of obtaining reference prompt words related to the target initial prompt word based on the target initial prompt word and the expansion strategy corresponding to the prompt word expansion mode includes: The target initial prompt words are multiple. Based on the multiple target initial prompt words, a preset number of reference prompt words are obtained through a prompt word generation model, wherein the correlation between the reference prompt words is less than a correlation threshold.

5. The method according to claim 1, characterized in that, When the prompt word expansion mode is the third mode, the step of obtaining reference prompt words related to the target initial prompt word based on the target initial prompt word and the expansion strategy corresponding to the prompt word expansion mode includes: When the initial target prompt word is one and contains multiple sub-prompt words, in response to the selection of the target sub-prompt word, the target sub-prompt word is retained, and other sub-prompt words are replaced by the prompt word generation model; Based on the target sub-prompt word and the other sub-prompt words after replacement, obtain reference prompt words related to the target initial prompt word.

6. The method according to claim 1, characterized in that, The process of obtaining the initial prompt word includes: Obtain the initial prompt word from the prompt word generation model, wherein the prompt word generation model is a model obtained after initializing the initial prompt word generation model based on randomness parameters, diversity parameters, and association parameters. The randomness parameters and the diversity parameters are parameters to be adjusted by the target user when setting the initial prompt word generation model, and the association parameters are parameters required when loading the reference prompt words in the background.

7. The method according to claim 6, characterized in that, The method further includes: Obtain a first number of reference prompt words, wherein the first number of reference prompt words is determined based on the batch value in the association parameter; Send feedback information for the first number of the reference prompt words; Obtain a second number of reference prompt words, wherein the second number of reference prompt words is determined based on the number of reference prompt words generated per request in the associated parameters and the number of requests collected in a single collection.

8. A device for expanding prompt words, characterized in that, The device includes: The first acquisition module is used to acquire an initial prompt word, wherein the initial prompt word is related to the target image to be acquired by the target user; The determining module is used to select a prompt word expansion mode and determine a target initial prompt word from the initial prompt words based on the number of prompt words corresponding to the prompt word expansion mode; The second acquisition module is configured to acquire reference prompts related to the target initial prompt based on the target initial prompt and the expansion strategy corresponding to the prompt expansion mode, wherein the reference prompts are expanded prompts of the target initial prompt. When the prompt expansion mode is the second mode, the second acquisition module includes: The determination unit is used when there is only one target initial prompt word. The target initial prompt word is split into multiple attributes and the sub-prompt words corresponding to each attribute. The first replacement unit is used to respond to the selection of the target attribute, retain other sub-cue words, and replace the target sub-cue words corresponding to the target attribute through the cue word generation model; The sixth acquisition unit is used to acquire reference prompts related to the target initial prompt based on the replaced target sub-prompt and other sub-prompts.

9. A computer device, characterized in that, include: A memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, the processor executing the computer instructions to perform the prompt word expansion method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the prompt word expansion method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Intelligent cue word optimization method and system for generating images through characters

    CN116012492A

  • Image generation method and device, electronic equipment and computer readable storage medium

    CN116580127A

  • Image generation method and device, electronic equipment and storage medium

    CN117170559A