Prompt word expansion method and apparatus, and computer device and storage medium
By obtaining the initial prompt words and obtaining reference prompt words according to the expansion strategy, the problem that prompt words in the prior art cannot meet user needs and batch automated production is solved, and more efficient and flexible prompt word expansion is achieved.
Patent Information
- Application Number
- PCT/CN2024/122530
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-09-29
- Publication Date
- 2025-06-26
AI Technical Summary
The prompt words of the existing text-generated image model cannot meet the user's customized needs and cannot achieve batch automation production needs.
By obtaining the initial prompt word, selecting the prompt word expansion mode, and obtaining reference prompt words according to the expansion strategy to meet user needs and achieve batch automation production.
It achieves the prompt word requirements of closely aligned with the target users, provides users with more diverse and free choices, lowers the threshold for prompt word writing, and automatically obtains a large number of high-quality prompt words.
Smart Images

Figure CN2024122530_26062025_PF_FP_ABST
Abstract
Description
Method, device, computer equipment and storage medium for expanding prompt words
[0001] This application claims priority to the Chinese patent application filed on December 22, 2023, with application number 202311789659.5 and invention name “Method, device, computer equipment and storage medium for prompt word expansion”. The entire contents of that application are incorporated by reference into this application. Technical Field
[0002] The present disclosure relates to the field of computer technology, and in particular to a method, device, computer equipment, and storage medium for expanding prompt words. Background Art
[0003] With the development of large language models, models for generating images from text have emerged. The principle of generating images from text is to encode the input text into a vector representation and then use a deep learning model to convert the vector representation into the corresponding image.
[0004] In common text-to-image generation scenarios, prompt word engineering always plays a key role. However, many users often face high barriers to entry and difficulty quickly getting started. While many prompt word plugins exist, they fail to effectively match user customization needs in practice and are unable to meet the needs of mass-producing high-quality prompt words. Furthermore, some existing text-to-image generation models fail to meet both user needs and the demands of mass automated production.
[0005] Therefore, the prompt words of the current text-to-image model have technical problems that cannot meet user needs and batch automated production requirements.
[0006] Summary of the Invention
[0007] In view of this, the present disclosure provides a method, apparatus, computer device and storage medium for expanding prompt words to solve the technical problem that the prompt words of the current text-to-image model cannot meet user needs and batch automated production requirements.
[0008] In a first aspect, the present disclosure provides a method for expanding prompt words, the method comprising: obtaining an initial prompt word, wherein the initial prompt word is related to a target image to be obtained by a target user; selecting a prompt word expansion mode, and determining a target initial prompt word from the initial prompt words according to the number of prompt words corresponding to the prompt word expansion mode; obtaining a reference prompt word related to the target initial prompt word according to an expansion strategy corresponding to the target initial prompt word and the prompt word expansion mode, wherein the reference prompt word is an expanded prompt word of the target initial prompt word.
[0009] In the disclosed embodiment, an initial prompt word and a selected prompt word expansion model are obtained, and a target initial prompt word is determined from the obtained initial prompt words based on the number of prompt words corresponding to the selected prompt word expansion model. Then, a reference prompt word is obtained after expanding the target initial prompt word according to the expansion strategy corresponding to the target initial prompt word and the prompt word expansion model. In this way, the disclosed embodiment can closely align with the prompt word needs of target users, providing them with more diverse and free choices. While significantly lowering the threshold for writing prompt words, it can also automatically obtain a large number of high-quality prompt words, resolving the technical problem in related technologies where prompt words cannot meet user needs and the requirements of mass automated production.
[0010] In an optional embodiment, obtaining the initial prompt word includes: obtaining main keywords input by the target user for describing the target image; analyzing the main keywords to obtain expanded keywords; and obtaining the initial prompt word based on the main keywords and the expanded keywords.
[0011] In the disclosed embodiment, by reasonably supplementing the subject keywords input by the target user, the target image is more richly portrayed, helping the target user to quickly and easily complete the accumulation of initial prompt words.
[0012] In an optional implementation, obtaining the initial prompt word includes: obtaining a reference image input by the target user; and analyzing the reference image to obtain the initial prompt word describing the reference image.
[0013] In the disclosed embodiment, by analyzing the reference image input by the target user, the generated initial prompt words can represent the target image more accurately, helping the target user to quickly and easily complete the accumulation of initial prompt words.
[0014] In an optional embodiment, when the prompt word expansion mode is the first mode, reference prompt words related to the target initial prompt word are obtained according to the expansion strategy corresponding to the target initial prompt word and the prompt word expansion mode, including: there are multiple target initial prompt words, and based on the multiple target initial prompt words, a preset number of reference prompt words are obtained through a prompt word generation model, wherein the correlation between the reference prompt words is less than the correlation threshold.
[0015] In the embodiment of the present disclosure, the number of target initial prompt words is determined according to the prompt word expansion mode selected by the target user, and then the reference prompt words that meet the expansion strategy are obtained through the prompt word generation model, so as to achieve the effect of closely aligning the prompt word requirements of the target user.
[0016] In an optional embodiment, when the prompt word expansion mode is the second mode, reference prompt words related to the target initial prompt word are obtained according to the expansion strategy corresponding to the target initial prompt word and the prompt word expansion mode, including: the target initial prompt word is one, the target initial prompt word is split, and multiple attributes of the target initial prompt word and sub-prompt words corresponding to each attribute are determined; in response to the selection of the target attribute, other sub-prompt words are retained, and the target sub-prompt words corresponding to the target attribute are replaced by the prompt word generation model; based on the replaced target sub-prompt words and other sub-prompt words, reference prompt words related to the target initial prompt word are obtained.
[0017] In the disclosed embodiment, the number of target initial prompt words is determined according to the prompt word expansion mode selected by the target user, and then the target initial prompt words are split to obtain multiple attributes. According to the target user's selection of the target attribute, the target sub-prompt words corresponding to the target attribute are replaced by the prompt word generation model, and other sub-prompt words are retained, thereby obtaining reference prompt words that meet the expansion strategy, thereby achieving the effect of closely aligning the prompt word requirements of the target user.
[0018] In an optional embodiment, when the prompt word expansion mode is the third mode, a reference prompt word related to the target initial prompt word is obtained according to the expansion strategy corresponding to the target initial prompt word and the prompt word expansion mode, including: when there is one target initial prompt word and it contains multiple sub-prompt words, in response to the selection of the target sub-prompt word, the target sub-prompt word is retained, and other sub-prompt words are replaced by the prompt word generation model; based on the target sub-prompt word and the other replaced sub-prompt words, a reference prompt word related to the target initial prompt word is obtained.
[0019] In the embodiment of the present disclosure, the number of target initial prompt words is determined according to the prompt word expansion mode selected by the target user. Then, based on the target sub-prompt words that the target user needs to retain, other sub-prompt words are replaced through the prompt word generation model to obtain reference prompt words that meet the expansion strategy, thereby achieving the effect of closely aligning the prompt word requirements of the target user.
[0020] In an optional embodiment, obtaining the initial prompt word includes: obtaining the initial prompt word in the prompt word generation model, wherein the prompt word generation model is a model obtained after initializing the initial prompt word generation model based on randomness parameters, diversity parameters and association parameters, the randomness parameters and diversity parameters are parameters to be adjusted by the target user when setting the initial prompt word generation model, and the association parameters are parameters required when loading the reference prompt word in the background.
[0021] In the embodiment of the present disclosure, the initial prompt word generation model is initialized by obtaining the target user's settings for the front-end parameters and the settings of the existing associated parameters in the background. In this way, when obtaining the initial prompt words, the target user is provided with more diverse and free choices.
[0022] In an optional embodiment, the method also includes: obtaining a first number of reference prompt words, wherein the first number of reference prompt words is determined based on the batch value in the associated parameter; sending feedback information on the first number of reference prompt words; obtaining a second number of reference prompt words, wherein the second number of reference prompt words is determined based on the number of reference prompt words generated by each request in the associated parameter and the number of requests collected in a single time.
[0023] In the embodiment of the present disclosure, by reasonably setting the acquired associated parameters, a large number of high-quality prompt words can be acquired.
[0024] In a second aspect, the present disclosure provides a device for expanding prompt words, which includes: a first acquisition module, used to obtain an initial prompt word, wherein the initial prompt word is related to a target image to be obtained by a target user; a determination module, used to select a prompt word expansion mode, and determine a target initial prompt word from the initial prompt words according to the number of prompt words corresponding to the prompt word expansion mode; a second acquisition module, used to obtain a reference prompt word related to the target initial prompt word according to the target initial prompt word and the expansion strategy corresponding to the prompt word expansion mode, wherein the reference prompt word is an expanded prompt word of the target initial prompt word.
[0025] In a third aspect, the present disclosure provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to execute the method for expanding prompt words of the first aspect or any corresponding embodiment thereof.
[0026] In a fourth aspect, the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to cause a computer to execute the method for expanding prompt words according to the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0028] FIG1 is a flow chart of a method for expanding prompt words according to some embodiments of the present disclosure;
[0029] FIG2 is a flow chart of a prompt word interactive learning mode according to some embodiments of the present disclosure;
[0030] FIG3 is a flow chart of a prompt word interactive learning mode according to other embodiments of the present disclosure;
[0031] FIG4 is a flow chart of a prompt word interactive learning mode according to some further embodiments of the present disclosure;
[0032] FIG5 is a schematic diagram of the overall process of a method for expanding prompt words according to some embodiments of the present disclosure;
[0033] FIG6 is a schematic diagram of a process of loading a ciphertext template according to some embodiments of the present disclosure;
[0034] FIG7 is a structural block diagram of an apparatus for expanding prompt words according to some embodiments of the present disclosure;
[0035] FIG8 is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0036] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present disclosure.
[0037] In common text-to-image generation scenarios, prompt word engineering always plays a key role. However, existing prompt word plug-ins cannot well match the user's customized needs in actual use, nor can they meet the task of mass production of high-quality prompt words. At the same time, some existing text-to-image generation models are also unable to meet user needs and mass automated production needs. In order to solve the above problems, according to an embodiment of the present disclosure, a method embodiment of prompt word expansion is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.
[0038] In this embodiment, a method for expanding prompt words is provided. FIG1 is a flow chart of the method for expanding prompt words according to an embodiment of the present disclosure. As shown in FIG1 , the method can be applied to a front-end client. The method flow includes the following steps:
[0039] Step S101: obtaining an initial prompt word, wherein the initial prompt word is related to a target image to be obtained by a target user.
[0040] Optionally, with the development of models, models that generate text from text and models that generate images from text come into being.
[0041] Specifically, in the process of generating an image from text, it is necessary to first rely on prompt words input by the user, and then output the corresponding image based on these prompt words. In this case, in the disclosed embodiment, the client can obtain the initial prompt words input by the target user to describe the target image to be obtained, or it can obtain the initial prompt words from a text-to-text model, such as a prompt word generation model. It should be noted that the number of initial prompt words can be one or more. In addition, the target user can be a specific user or a user cluster.
[0042] Step S102: selecting a prompt word expansion mode, and determining a target initial prompt word from the initial prompt words according to the number of prompt words corresponding to the prompt word expansion mode.
[0043] Optionally, the client obtains the prompt word expansion mode selected by the target user on the terminal screen, obtains the number of prompt words corresponding to the selected prompt word expansion mode, and then obtains the target initial prompt words that meet the prompt word quantity from the initial prompt words. For example, if the prompt word expansion mode is the first mode (such as the multiple example whole sentence mode), the corresponding number of prompt words is multiple, such as 3-100, then the number of target initial prompt words obtained from the initial prompt words is 3-100.
[0044] Step S103 : obtaining a reference prompt word related to the target initial prompt word according to the target initial prompt word and the expansion strategy corresponding to the prompt word expansion mode, wherein the reference prompt word is an expanded prompt word of the target initial prompt word.
[0045] Optionally, the client retrieves reference prompts related to the target initial prompt based on the number of target initial prompts and the expansion strategy corresponding to the currently selected prompt expansion mode. Reference prompts are expansion prompts of the target initial prompt. Furthermore, the expansion strategy can typically be based on the target user's personalized needs for the number of prompts or homogeneity of results. By combining these personalized needs with the number of target initial prompts, reference prompts related to the initial prompt can be retrieved.
[0046] In the disclosed embodiment, an initial prompt word and a selected prompt word expansion model are obtained, and a target initial prompt word is determined from the obtained initial prompt words based on the number of prompt words corresponding to the selected prompt word expansion model. Then, a reference prompt word is obtained after expanding the target initial prompt word according to the expansion strategy corresponding to the target initial prompt word and the prompt word expansion model. In this way, the disclosed embodiment can closely align with the prompt word needs of target users, providing them with more diverse and free choices. While significantly lowering the threshold for writing prompt words, it can also automatically obtain a large number of high-quality prompt words, resolving the technical problem in related technologies where prompt words cannot meet user needs and the requirements of mass automated production.
[0047] In some optional implementations, obtaining the initial prompt word includes: obtaining main keywords input by the target user to describe the target image; analyzing the main keywords to obtain expanded keywords; and obtaining the initial prompt word based on the main keywords and the expanded keywords.
[0048] Optionally, when selecting the initial prompt word, the initial prompt word can be obtained externally, such as through interaction between the client and the cloud or other server databases, or it can be generated based on some auxiliary methods: a prompt word plug-in within the client can provide an expansion method to help the target user quickly and easily complete the accumulation of initial prompt words: obtain the main keywords input by the target user to describe the target image, for example, if the target image is a portrait, the main keyword can be a male painter with a beard.
[0049] After obtaining the main keyword, the prompt word plug-in will reasonably fill in the main keyword and obtain some expanded keywords that describe the details of the target image to supplement the main keyword, making the main keyword more vivid and more conducive to generating pictures that conform to the internal logic and are more in line with actual needs.
[0050] The vocabulary consisting of the main keywords and the expanded keywords is then referred to as the initial prompt word. The set containing the main keywords and the expanded keywords can also be referred to as the initial prompt word.
[0051] In the disclosed embodiment, by reasonably supplementing the subject keywords input by the target user, the target image is more richly portrayed, helping the target user to quickly and easily complete the accumulation of initial prompt words.
[0052] In some optional implementations, obtaining the initial prompt word includes: obtaining a reference image input by the target user; and analyzing the reference image to obtain the initial prompt word describing the reference image.
[0053] Optionally, an embodiment of the present disclosure proposes a method of generating initial prompt words in a described auxiliary manner: the target user inputs some reference images to the client that the target user believes have a similarity with the target image greater than a first preset threshold (for example, 0.9). It can be understood that the number of these reference images can be one or more, and the reference images are usually excellent images that the target user believes are relatively close to the target image.
[0054] At this time, the prompt word plug-in of the client completes the analysis of the reference image, generates descriptive words for the reference image, and uses these descriptive words as initial prompt words.
[0055] In addition, experienced users can also use prompt words written or accumulated by themselves as initial prompt words.
[0056] In the disclosed embodiment, by analyzing the reference image input by the target user, the generated initial prompt words can represent the target image more accurately, helping the target user to quickly and easily complete the accumulation of initial prompt words.
[0057] In some optional embodiments, when the prompt word expansion mode is the first mode, reference prompt words related to the target initial prompt word are obtained according to the expansion strategy corresponding to the target initial prompt word and the prompt word expansion mode, including: there are multiple target initial prompt words, and based on the multiple target initial prompt words, a preset number of reference prompt words are obtained through a prompt word generation model, wherein the correlation between the reference prompt words is less than the correlation threshold.
[0058] Optionally, after the client obtains the target initial prompt word, it can perform interactive learning on the target initial prompt word based on the prompt word generation model to obtain expanded reference prompt words. Specifically, when the prompt word expansion mode selected by the target user is the first mode (such as the diverse example whole sentence mode), since there are multiple target initial prompt words corresponding to the first mode, based on the expansion strategy, for example, the expansion strategy of the first mode is to avoid homogeneity of the results of the reference prompt words, that is, the correlation between the desired reference prompt words is less than the correlation threshold (such as 0.7), the client will guide the prompt word generation model to learn the target initial prompt word and generate reference prompt words that meet the above expansion strategy (i.e., heterogeneous results and large-scale acquisition).
[0059] It should be noted that homogeneity means that the span between reference prompt words is small and the results belong to a smaller range.
[0060] The embodiments of the present disclosure are aimed at scenarios where the target user does not need precise prompt words. For example, the target initial prompt words provided by the target user may not be limited to "beard" or "painter", and the reference prompt words obtained may not necessarily contain a specific subject or a specific vocabulary description, and their diversity is rich.
[0061] As shown in Figure 2, the target initial prompt word is input; a determination is made as to whether all target initial prompt words have been entered. If not, the target initial prompt word is then input again. If completed, the prompt word generation model learns these target initial prompt words to obtain reference prompt words. Furthermore, after inputting the initial prompt words, positive examples can be selected from them and used as learning objects for the prompt word generation model. Positive examples are target initial prompt words that more accurately and completely represent the target image.
[0062] In the embodiment of the present disclosure, the number of target initial prompt words is determined according to the prompt word expansion mode selected by the target user, and then the reference prompt words that meet the expansion strategy are obtained through the prompt word generation model, so as to achieve the effect of closely aligning the prompt word requirements of the target user.
[0063] In some optional embodiments, when the prompt word expansion mode is the second mode, reference prompt words related to the target initial prompt word are obtained according to the expansion strategy corresponding to the target initial prompt word and the prompt word expansion mode, including: the target initial prompt word is one, the target initial prompt word is split, and multiple attributes of the target initial prompt word and sub-prompt words corresponding to each attribute are determined; in response to the selection of the target attribute, other sub-prompt words are retained, and the target sub-prompt words corresponding to the target attribute are replaced by the prompt word generation model; based on the replaced target sub-prompt words and other sub-prompt words, reference prompt words related to the target initial prompt word are obtained.
[0064] Optionally, when the target user selects the second mode (which may be the smart split mode) for the prompt word expansion mode on the client, since the number of target initial prompt words corresponding to the second mode is one, if the target user desires to customize the target sub-prompt word corresponding to a target attribute among the multiple attributes included in the target initial prompt word, the client will respond to the target user's replacement operation performed on the client. Specifically, the client first splits the target initial prompt word, determines the multiple attributes of the target initial prompt word and the sub-prompt words corresponding to each attribute, and then uses the prompt word generation model to replace the target sub-prompt words corresponding to the target attribute based on the target user's selection of the target attribute, while retaining the other sub-prompt words that the target user did not replace. The attributes included in the aforementioned target initial prompt word may be the gender, clothing, and accessory attributes of the person depicted in the target image. For example, if the target sub-prompt word corresponding to the target attribute (such as clothing) that the target user wants to customize is originally Hanfu, the target user customizes the target sub-prompt word to Hanfu, and retains the other sub-prompt words in the target initial prompt word.
[0065] The replaced target sub-prompt word and other sub-prompt words are used as updated target initial prompt words, and the client then obtains related reference prompt words based on the updated target initial prompt words.
[0066] As shown in Figure 3, the system inputs an initial target prompt word; intelligently splits the initial target prompt word into different attributes; customizes the target sub-prompt words corresponding to some or all of the attributes you want to replace; replaces the target sub-prompt words corresponding to the customized attributes; and generates reference prompt words based on the replaced and retained parts. Furthermore, after inputting the initial target prompt word, you can select positive examples from it and use them as the targets for intelligent splitting. Positive examples are target initial prompt words that more accurately and completely represent the target image.
[0067] In the disclosed embodiment, the number of target initial prompt words is determined according to the prompt word expansion mode selected by the target user, and then the target initial prompt words are split to obtain multiple attributes. According to the target user's selection of the target attribute, the target sub-prompt words corresponding to the target attribute are replaced by the prompt word generation model, and other sub-prompt words are retained, thereby obtaining reference prompt words that meet the expansion strategy, thereby achieving the effect of closely aligning the prompt word requirements of the target user.
[0068] In some optional embodiments, when the prompt word expansion mode is the third mode, reference prompt words related to the target initial prompt word are obtained according to the expansion strategy corresponding to the target initial prompt word and the prompt word expansion mode, including: when there is one target initial prompt word and it contains multiple sub-prompt words, in response to the selection of the target sub-prompt word, the target sub-prompt word is retained, and other sub-prompt words are replaced by the prompt word generation model; based on the target sub-prompt word and the other replaced sub-prompt words, a reference prompt word related to the target initial prompt word is obtained.
[0069] Optionally, when the prompt word expansion mode selected by the target user on the client is the third mode (which can be a custom split mode), since the number of target initial prompt words corresponding to the third mode is 1, and the target initial prompt word contains multiple sub-prompt words, if the target user himself has a clear target sub-prompt word that he wants to retain, the client only needs to respond to the target sub-prompt word that the target user selects on the client to retain, and then use the prompt word generation model to replace other sub-prompt words, and then use the replaced other sub-prompt words and the target sub-prompt word as the updated target initial prompt word. The client then obtains relevant reference prompt words based on the updated target initial prompt word.
[0070] As shown in Figure 4, the target initial prompt word is input; the desired sub-prompt word portion is obtained; the target sub-prompt word is retained and the remaining sub-prompt words are replaced; and a reference prompt word is obtained based on the replaced and retained portions. Furthermore, after inputting the target initial prompt word, positive examples can be selected from the initial prompt word. These positive examples serve as the basis for the target user to determine whether to retain some sub-prompt words. Positive examples are those that more accurately and completely represent the target image.
[0071] In the embodiment of the present disclosure, the number of target initial prompt words is determined according to the prompt word expansion mode selected by the target user. Then, based on the target sub-prompt words that the target user needs to retain, other sub-prompt words are replaced through the prompt word generation model to obtain reference prompt words that meet the expansion strategy, thereby achieving the effect of closely aligning the prompt word requirements of the target user.
[0072] In some optional embodiments, obtaining the initial prompt word includes: obtaining the initial prompt word in the prompt word generation model, wherein the prompt word generation model is a model obtained after initializing the initial prompt word generation model based on randomness parameters, diversity parameters and association parameters, the randomness parameters and diversity parameters are parameters to be adjusted by the target user when setting the initial prompt word generation model, and the association parameters are parameters required when loading the reference prompt word in the background.
[0073] Optionally, when obtaining the initial prompt word, the client of the embodiment of the present disclosure can also obtain the prompt word generation model and obtain the initial prompt word from the prompt word generation model. Before obtaining the prompt word generation model, the initial prompt word generation model needs to be initialized. The initialization process mainly includes parameter setting and adjustment, parameter loading, initialization model guidance, etc., and is completed through the collaboration of the front-end client and the back-end:
[0074] The client loads the target user's customized parameters (such as the corresponding randomness parameters and diversity parameters when generating reference prompt words).
[0075] It should be noted that if the target user wants the reference prompt words to be more diverse and richer, the values of the randomness parameter and the diversity parameter can be increased. If the target user wants the reference prompt words to be more stable, the values of the randomness parameter and the diversity parameter can be lowered. It is also possible to provide the target user with some recommended values so that the target user can refer to the recommended values for setting. For example, the randomness parameter of the X model is set to 1.4, the diversity parameter is set to 0.9, the randomness parameter of the Y model is set to 0.7, and the diversity parameter is set to 0.9, and so on.
[0076] The background loads the associated parameters required for reference prompt words (including the initial prompt word generation model type, batch size, the number of reference prompt words generated per request, the number of requests collected in a single time, etc.), loads / calls the selected initial prompt word generation model, and guides the initialization of the initial prompt word generation model.
[0077] In the embodiment of the present disclosure, the initial prompt word generation model is initialized by obtaining the target user's settings for the front-end parameters and the settings of the existing associated parameters in the background. In this way, when obtaining the initial prompt words, the target user can be provided with more diverse and free choices.
[0078] In some optional embodiments, the method further includes: obtaining a first number of reference prompt words, wherein the first number of reference prompt words is determined based on the batch value in the associated parameter; sending feedback information on the first number of reference prompt words; and obtaining a second number of reference prompt words, wherein the second number of reference prompt words is determined based on the number of reference prompt words generated for each request in the associated parameter and the number of requests collected in a single time.
[0079] Optionally, after the client prompt word interaction is completed and the reference prompt words are obtained, the disclosed embodiment can further align the prompt word requirements of the target user through a generate-feedback-regenerate instruction fine-tuning mechanism. The client will first obtain a first number of reference prompt words, which can be a small batch number, such as 10. It should be noted that the number of single batches generated in the small batch generation of reference prompt words in this process is based on the batch size parameter of the associated parameter. The larger this parameter is set, the more feedback-regeneration rounds are generally required to fully meet the requirements in a single batch. However, once the requirements are aligned, the robustness of the batch collection will also be improved. Therefore, the batch size value is preferably set in the range of 5-10.
[0080] After the target user obtains the first number of reference prompt words through the client, if he approves the quality of the single batch generation, he will send feedback information, such as sending some confirmation instructions to the background, after which the background can start batch collection and send the second number of reference prompt words to the client. The second number at this time is the number of batch collections, such as 50. In this step, the background of the embodiment of the present disclosure will perform batch collection based on the two parameters of the associated parameters minibatch (the number of prompt words produced per request) and requests (the number of requests for a single collection). The number of single collections (i.e., the second number of reference prompt words) = minibatch*requests. When the number to be collected is certain, in order to improve utilization and reduce costs, it is usually recommended to use a larger minibatch under the premise that the context length of the prompt word generation model allows. This can increase the number of prompt words obtained in a single request, thereby reducing the waste of resources caused by repeated acquisition of context background when providing multiple requests.
[0081] In the embodiment of the present disclosure, by reasonably setting the acquired associated parameters, a large number of high-quality prompt words can be acquired.
[0082] In some optional implementations, as shown in FIG5 , the complete main process of the disclosed embodiment, from start to finish, primarily includes the following steps: initialization of the initial prompt word generation model, selection of the prompt word expansion mode, acquisition of the target initial prompt word, interactive learning of prompt words, generation and evaluation of small batches of results (if all results meet the requirements, the next step can be directly advanced; if not, the non-compliant results are marked, their sequence numbers and reasons are provided, and a new batch of results is generated until they meet the requirements), batch collection, and return of plaintext results and ciphertext templates. As shown in FIG6 , after obtaining the encrypted template, the target user can directly enter the state of waiting for batch generation, which is exactly the same as when the template was generated. Furthermore, before starting batch collection, the target user can also choose whether to make further fine-tuning to accommodate any slight changes in the target user's requirements. After the target user decides to start the batch collection, a new batch of plaintext results and a new ciphertext template containing all the execution processes for this time are generated.
[0083] Based on the disclosure of the above embodiments, the present embodiment selects operation examples of Mode 1 (corresponding to the scenario of obtaining a large number of prompt words), Mode 2 (corresponding to intelligent splitting, realizing the scenario of replacing the target sub-prompt words corresponding to custom attributes), and Mode 3 (corresponding to the scenario where the target user has selected the target sub-prompt words to be retained, replacing other sub-prompt words that have not been retained) in the interactive learning of prompt words to be explained:
[0084] About Mode 1:
[0085] 1. Select the initial prompt word to generate the model initialization;
[0086] 2. Randomness parameter settings (the higher the more random, the default is 1.5), diversity parameter settings (the higher the more diverse, the default is 0.8);
[0087] 3. Select Mode 1 for the prompt word expansion mode;
[0088] 4. Select and enter the target initial prompt word;
[0089] 5. Enter multiple positive examples of the target initial prompt word in the text box, and repeat this step until all positive examples have been entered;
[0090] 6. Select Start Generation to get reference words;
[0091] 7. Select Generate Again, evaluate the batch results, select the unsatisfactory result number, enter it in the text box, select Generate Again, and repeat this step until the batch results are satisfactory;
[0092] 8. Select Collect, enter the name of the result you want to save in the text box, and collect in batches.
[0093] About Mode 2:
[0094] 1. Select the initial prompt word to generate the model initialization;
[0095] 2. Randomness parameter settings (the higher the more random, the default is 1.5), diversity parameter settings (the higher the more diverse, the default is 0.8);
[0096] 3. Select Mode 2 for the prompt word expansion mode;
[0097] 4. Select and enter the target initial prompt word;
[0098] 5. Enter a positive example of the target initial prompt word in the text box;
[0099] 6. Select Start Generation. Based on the output smart split results, enter the target sub-prompt word portion corresponding to the attribute to be replaced in the text box, retain the other sub-prompt words, and obtain the reference prompt word.
[0100] 7. Select Generate Again, evaluate the batch results, select the unsatisfactory result number, enter it in the text box, select Generate Again, and repeat this step until the batch results are satisfactory;
[0101] 8. Select Collect, enter the name of the result you want to save in the text box, and collect in batches.
[0102] About Mode 3:
[0103] 1. Select the initial prompt word to generate the model initialization;
[0104] 2. Randomness parameter settings (the higher the more random, the default is 1.5), diversity parameter settings (the higher the more diverse, the default is 0.8);
[0105] 3. Select Mode 3 for the prompt word expansion mode;
[0106] 4. Select and enter the target initial prompt word;
[0107] 5. Enter a positive example of the target initial prompt word in the text box;
[0108] 6. Select Start Generation and enter the target sub-prompt word to be retained in the text box to replace the other sub-prompt words that are not retained to obtain the reference prompt word;
[0109] 7. Select Generate Again, evaluate the batch results, select the unsatisfactory result number, enter it in the text box, select Generate Again, and repeat this step until the batch results are satisfactory;
[0110] 8. Select Collect, enter the name of the result you want to save in the text box, and collect in batches.
[0111] This embodiment also provides a device for expanding prompt words, which is used to implement the above-mentioned embodiments and preferred embodiments. Details already described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented using software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0112] This embodiment provides a device for expanding prompt words, as shown in Figure 7, including: a first acquisition module 701, used to obtain an initial prompt word, wherein the initial prompt word is related to a target image to be obtained by a target user; a determination module 702, used to select a prompt word expansion mode and determine a target initial prompt word from the initial prompt words based on the number of prompt words corresponding to the prompt word expansion mode; a second acquisition module 703, used to obtain a reference prompt word related to the target initial prompt word based on the target initial prompt word and the expansion strategy corresponding to the prompt word expansion mode, wherein the reference prompt word is an expanded prompt word of the target initial prompt word.
[0113] In some optional embodiments, the first acquisition module 701 includes: a first acquisition unit, used to obtain the main keywords input by the target user to describe the target image; a second acquisition unit, used to analyze the main keywords and obtain expanded keywords; a first obtaining unit, used to obtain initial prompt words based on the main keywords and expanded keywords.
[0114] In some optional implementations, the first acquisition module 701 includes: a third acquisition unit, configured to acquire a reference image input by a target user; and a fourth acquisition unit, configured to analyze the reference image and acquire an initial prompt word describing the reference image.
[0115] In some optional embodiments, when the prompt word expansion mode is the first mode, the second acquisition module 703 includes: a fifth acquisition unit, which is used to obtain a preset number of reference prompt words through a prompt word generation model based on multiple target initial prompt words when there are multiple target initial prompt words, wherein the correlation between the reference prompt words is less than the correlation threshold.
[0116] In some optional embodiments, when the prompt word expansion mode is the second mode, the second acquisition module 703 includes: a determination unit, configured to split the target initial prompt word when there is one target initial prompt word, determine multiple attributes of the target initial prompt word and sub-prompt words corresponding to each attribute; a first replacement unit, configured to retain other sub-prompt words in response to selection of a target attribute, and replace the target sub-prompt words corresponding to the target attribute using the prompt word generation model;
[0117] The sixth acquisition unit is configured to acquire a reference prompt word related to the target initial prompt word based on the replaced target sub-prompt word and other sub-prompt words.
[0118] In some optional embodiments, when the prompt word expansion mode is the third mode, the second acquisition module 703 includes: a second replacement unit, which is used to retain the target sub-prompt word in response to the selection of the target sub-prompt word when the target initial prompt word is one and contains multiple sub-prompt words, and replace other sub-prompt words through the prompt word generation model; a seventh acquisition unit, which is used to obtain reference prompt words related to the target initial prompt word based on the target sub-prompt word and the replaced other sub-prompt words.
[0119] In some optional embodiments, the first acquisition module 701 includes: an eighth acquisition unit, used to obtain the initial prompt word in the prompt word generation model, wherein the prompt word generation model is a model obtained after initializing the initial prompt word generation model based on randomness parameters, diversity parameters and association parameters, the randomness parameters and diversity parameters are parameters to be adjusted by the target user when setting the initial prompt word generation model, and the association parameters are parameters required when loading the reference prompt word in the background.
[0120] In some optional embodiments, the device also includes: a third acquisition module, used to obtain a first number of reference prompt words, wherein the first number of reference prompt words is determined based on the batch value in the associated parameter; a sending module, used to send feedback information on the first number of reference prompt words; a fourth acquisition module, used to obtain a second number of reference prompt words, wherein the second number of reference prompt words is determined based on the number of reference prompt words generated by each request in the associated parameter and the number of requests collected in a single time.
[0121] The prompt word expansion device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0122] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0123] The embodiment of the present disclosure further provides a computer device having the apparatus for expanding prompt words as shown in FIG. 7 .
[0124] Please refer to Figure 8, which is a structural diagram of a computer device provided by an optional embodiment of the present disclosure. As shown in Figure 8, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 8 takes a processor 10 as an example.
[0125] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.
[0126] The memory 20 stores instructions that can be executed by at least one processor 10, so as to enable at least one processor 10 to execute the method shown in the above embodiment.
[0127] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created based on the use of a computer device for displaying a small program landing page, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0128] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0129] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0130] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.
[0131] Although the embodiments of the present disclosure have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A method for expanding prompt words, comprising: Acquire an initial prompt word, wherein the initial prompt word is related to a target image to be acquired by a target user; Selecting a prompt word expansion mode, and determining a target initial prompt word from the initial prompt words according to the number of prompt words corresponding to the prompt word expansion mode; According to the target initial prompt word and the expansion strategy corresponding to the prompt word expansion mode, a reference prompt word related to the target initial prompt word is acquired, wherein the reference prompt word is an expanded prompt word of the target initial prompt word.
2. The method according to claim 1, wherein the obtaining of the initial prompt word comprises: Acquire subject keywords input by the target user for describing the target image; Analyze the main keywords to obtain expanded keywords; The initial prompt word is obtained according to the main keyword and the expanded keyword.
3. The method according to claim 1, wherein the obtaining of the initial prompt word comprises: Acquire a reference image input by the target user; The reference image is analyzed to obtain the initial prompt words describing the reference image.
4. The method according to claim 1, wherein when the prompt word expansion mode is the first mode, acquiring a reference prompt word related to the target initial prompt word according to the target initial prompt word and the expansion strategy corresponding to the prompt word expansion mode comprises: There are multiple target initial prompt words, and a preset number of reference prompt words are acquired through a prompt word generation model based on the multiple target initial prompt words, wherein the correlation between the reference prompt words is less than a correlation threshold.
5. The method according to claim 1, wherein when the prompt word expansion mode is the second mode, acquiring a reference prompt word related to the target initial prompt word according to the expansion strategy corresponding to the target initial prompt word and the prompt word expansion mode comprises: The target initial prompt word is one, the target initial prompt word is split, and multiple attributes of the target initial prompt word and sub-prompt words corresponding to each attribute are determined; In response to the selection of the target attribute, other sub-prompt words are retained, and the target sub-prompt words corresponding to the target attribute are replaced by the prompt word generation model; Based on the replaced target sub-prompt word and the other sub-prompt words, a reference prompt word related to the target initial prompt word is obtained.
6. The method according to claim 1, wherein when the prompt word expansion mode is the third mode, acquiring a reference prompt word related to the target initial prompt word according to the expansion strategy corresponding to the target initial prompt word and the prompt word expansion mode comprises: When the target initial prompt word is one and contains multiple sub-prompt words, in response to the selection of the target sub-prompt word, the target sub-prompt word is retained, and other sub-prompt words are replaced by the prompt word generation model; Based on the target sub-prompt word and the replaced other sub-prompt words, a reference prompt word related to the target initial prompt word is obtained.
7. The method according to claim 1, wherein the obtaining of the initial prompt word comprises: Obtain the initial prompt word in the prompt word generation model, wherein the prompt word generation model is a model obtained after initializing the initial prompt word generation model based on a randomness parameter, a diversity parameter, and an association parameter, the randomness parameter and the diversity parameter are parameters to be adjusted by the target user when setting the initial prompt word generation model, and the association parameter is a parameter required when loading the reference prompt word in the background.
8. The method according to claim 7, wherein the method further comprises: Acquire a first number of the reference prompt words, wherein the first number of the reference prompt words is determined according to a batch value in the association parameter; Sending feedback information on the first number of reference prompt words; A second number of the reference prompt words is obtained, wherein the second number of the reference prompt words is determined according to the number of reference prompt words generated by each request in the associated parameters and the number of requests collected in a single time.
9. A device for expanding prompt words, comprising: A first acquisition module is used to acquire an initial prompt word, wherein the initial prompt word is related to a target image to be acquired by a target user; A determination module, used for selecting a prompt word expansion mode, and determining a target initial prompt word from the initial prompt words according to the number of prompt words corresponding to the prompt word expansion mode; The second acquisition module is used to acquire a reference prompt word related to the target initial prompt word according to the target initial prompt word and the expansion strategy corresponding to the prompt word expansion mode, wherein the reference prompt word is an expanded prompt word of the target initial prompt word.
10. A computer device comprising: A memory and a processor, wherein the memory and the processor are connected to each other for communication, the memory stores computer instructions, and the processor executes the method for expanding prompt words according to any one of claims 1 to 8 by executing the computer instructions.
11. A computer-readable storage medium having computer instructions stored thereon, wherein the computer instructions are used to enable a computer to execute the method for expanding prompt words according to any one of claims 1 to 8.
Citation Information
Patent Citations
Text expansion method and device based on artificial intelligence, equipment and storage medium
CN114385791A
Intelligent cue word optimization method and system for generating images through characters
CN116012492A
Image generation method and device, electronic equipment and computer readable storage medium
CN116580127A
Image generation method and device, electronic equipment and storage medium
CN117170559A
Prompt word expansion method and device, computer equipment and storage medium
CN117764068A
Cited By
AI dialogue cue word optimization method and system
CN121029959A
Method and device for reducing token consumption of large model and storage medium
CN122489730A