Prompt generation method and apparatus, electronic device, and program product

By introducing an error feedback mechanism and automated evaluation scores, optimized prompts are generated, solving the problems of time-consuming, labor-intensive, and subjective prompt optimization in existing technologies. This achieves efficient and accurate prompt optimization, improving the user experience.

WO2025260251A1PCT designated stage Publication Date: 2025-12-26DOUYIN VISION CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/099948
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-18
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

In existing technologies, suggestion optimization mainly relies on human experience and repeated trials, which is time-consuming and labor-intensive, and it is difficult to adapt to diverse task requirements and dataset characteristics. Especially in the case of big data, the iteration efficiency is low, and the human optimization method is highly subjective and has poor results.

Method used

By introducing an error feedback mechanism, candidate prompts are generated and evaluated, and optimized prompts are automatically filtered. Combining error samples and user feedback, the optimization process is more focused on actual problems, reduces manual intervention, and improves efficiency and accuracy.

Benefits of technology

It improves the efficiency and accuracy of suggestion optimization, meets user needs, ensures that model output is more in line with expectations, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024099948_26122025_PF_FP_ABST
    Figure CN2024099948_26122025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a prompt generation method and apparatus, an electronic device, and a program product. The method comprises: on the basis of a first prompt and error feedback, generating a first group of candidate prompts for the first prompt, wherein the error feedback comprises at least one of an error sample and user feedback. The method further comprises: on the basis of scores of a plurality of candidate prompts in the first group of candidate prompts, determining a second group of candidate prompts for the first prompt. In addition, the method comprises: determining a second prompt from the second group of candidate prompts. According to the embodiments of the present disclosure, by introducing error feedback of a prompt in an actual application, a group of candidate prompts related to the prompt is generated, and the group of candidate prompts is evaluated to generate an optimized prompt. The method improves the user experience of using a model on the basis of an optimized prompt.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, electronic device and program product for generating a prompt TECHNICAL FIELD

[0001] The present disclosure relates generally to the field of computers, and more specifically, to a method, device, electronic device and program product for generating a prompt. BACKGROUND

[0002] With the rise of deep learning technology, two representative models, multi-modal model and language model, have emerged. Although multi-modal model and language model have their own focus, they are not isolated from each other. Multi-modal model can integrate text, image, audio and other data, realize comprehensive perception of data, and has wide application in multimedia content understanding, intelligent interaction, automatic driving and other fields.

[0003] Language model, on the other hand, relies on understanding and simulation of natural language rules to provide strong support for text generation and understanding, and has wide application in natural language processing (NLP) fields such as machine translation, text generation, intelligent question answering, etc. Prompt is an important concept in multi-modal model and language model, and is an important tool for guiding model to generate specific responses. In natural language processing and machine learning, a prompt can be a piece of text, a question or an instruction, used to guide or stimulate the response of a model. In generative models, the prompt is usually used as a seed text for input, and the model will continue to generate subsequent text content based on the seed text.

[0004] SUMMARY

[0005] Embodiments of the present disclosure provide a method, device, electronic device and program product for generating a prompt.

[0006] According to a first aspect of the present disclosure, a method for generating a prompt is provided. The method includes generating, based on a first prompt and error feedback, a first set of candidate prompts for the first prompt, wherein the error feedback includes at least one of an error sample and user feedback. The method further includes determining, based on scores of a plurality of candidate prompts in the first set of candidate prompts, a second set of candidate prompts for the first prompt. In addition, the method further includes determining, from the second set of candidate prompts, a second prompt.

[0007] In a second aspect of this disclosure, an apparatus for generating prompts is provided. The apparatus includes a first set of candidate prompt generation module configured to generate a first set of candidate prompts for the first prompt based on a first prompt and error feedback, wherein the error feedback includes at least one of error samples and user feedback. The apparatus also includes a second set of candidate prompt determination module configured to determine a second set of candidate prompts for the first prompt based on scores of multiple candidate prompts in the first set of candidate prompts. Furthermore, the apparatus includes a second prompt determination module configured to determine a second prompt from the second set of candidate prompts.

[0008] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes a processor and a memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform the method according to the first aspect.

[0009] In a fourth aspect of this disclosure, a computer program product is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the method of the first aspect.

[0010] The summary section is intended to present the chosen concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify key or principal features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description

[0011] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0012] Figure 1 shows a schematic diagram of an example environment in which some embodiments of the present disclosure may be implemented;

[0013] Figure 2 shows a flowchart of a method for generating prompts according to some embodiments of the present disclosure;

[0014] Figure 3 illustrates a schematic diagram of some embodiments of the present disclosure for applying optimizers and evaluators to generate optimized prompts;

[0015] Figure 4 illustrates a schematic diagram of some embodiments of this disclosure for generating optimized prompts via an optimizer;

[0016] Figure 5 illustrates a schematic diagram of some embodiments of the present disclosure for evaluating the score of a candidate prompt based on an answer generated from the candidate prompt;

[0017] FIG. 6 shows a block diagram of an apparatus for generating a prompt according to some embodiments of the present disclosure; and

[0018] FIG. 7 shows a block diagram of an electronic device according to some embodiments of the present disclosure.

[0019] In all the drawings, like or similar reference numerals refer to like or similar elements. DETAILED DESCRIPTION

[0020] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.

[0021] It can be understood that the above notification and user authorization acquisition process is only illustrative, and does not limit the implementation of the present disclosure. Other ways that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0022] Embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for exemplary purposes only, and are not intended to limit the scope of protection of the present disclosure.

[0023] In the description of embodiments of the present disclosure, the term "comprising" and its similar words should be understood as open-ended including, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. can refer to different or the same objects unless explicitly stated. The following can also include other explicit and implicit definitions.

[0024] As mentioned above, the prompt is an important tool to guide the model to generate a specific response. A well-designed prompt can effectively guide the model to generate more expected output, and vice versa, an inappropriate prompt may lead the model output to deviate from the target task. In the related art, the prompt optimization method mainly relies on manual experience and repeated trials, which not only consumes time and effort, but also is difficult to adapt to diversified task requirements and data set characteristics. Generally, it requires professional domain knowledge to build an effective prompt, for example, in tasks such as medical care that require deep professional background, there may be a significant difference between the effects of novices and experts in optimizing the prompt. Because experts can more accurately grasp key information and professional terms, and thus build more targeted and efficient prompts, while the performance of the prompts built by novices is not as good as that of the experts. In addition, the optimization of the prompt also needs to attribute to the error prediction samples and convert these attributions into the optimization direction of the prompt. However, at present, this process mainly relies on manual data analysis and attribution to guide the optimization of the prompt. This method not only has a large amount of work, but also has very low iteration efficiency, especially in the case of a large amount of data, which seriously affects the progress and effect of the prompt optimization work.

[0025] According to an embodiment of the present disclosure, based on the initial prompt and the error feedback from the actual application, such as error samples and / or user feedback, a first set of candidate prompts for the initial prompt is generated, so that the optimization process of the prompt is more focused on the problems encountered by the prompt in the actual operation, thereby ensuring the accuracy and practicality of the prompt optimization. And through the evaluation score mechanism, the first set of candidate prompts is quantitatively evaluated, and the prompts are selected according to the scores to form a second set of candidate prompts, which not only avoids the subjectivity of manual selection, but also improves the optimization efficiency. In addition, the final optimized prompt is determined from the second set of candidate prompts after screening, which not only guarantees the quality of the optimization decision, but also reduces manual intervention, making the whole process more intelligent and automated.

[0026] In summary, the embodiments of the present disclosure generate an optimized prompt by introducing error feedback of the prompt in actual application. This automatic process not only improves the efficiency and accuracy of prompt optimization, but also meets the user's demand for building a prompt that makes the target model, such as a multi-modal model or a language model, output better performance when applying the model, and improves the user experience.

[0027] FIG. 1 shows a schematic diagram of an example environment 100 in which some embodiments of the present disclosure can be implemented. As shown in FIG. 1, an initial prompt 110 can be input into a prompt optimization 120 to generate an optimized prompt 130. In some embodiments, the initial prompt 110 can be a prompt for a specific task, such as a prompt for a text classification task, a prompt for a natural language generation (NLG) task, or a prompt for a multi-modal model generation task. The initial prompt 110 can be a starting point of a process, which can be a simple instruction, a preliminary description of a question, or a basic description of a task to be solved. The initial prompt 110 usually contains basic information that can not be specific, clear, or targeted enough to meet the specific needs or goals of a user. Therefore, the initial prompt 110 can be optimized to generate an optimized prompt 130 that can fully stimulate the potential of a target model, such as a natural language model, and produce an expected output.

[0028] With continued reference to FIG. 1, the initial prompt 110 can be input into an optimizer 122 to generate a set of candidate prompts related to the initial prompt 110. The set of candidate prompts can be input into an evaluator 124 to generate another set of candidate prompts related to the initial prompt 110, which can be selected from the set of candidate prompts generated by the optimizer 122. From the another set of candidate prompts evaluated by the evaluator 124, an optimal prompt can be selected as the optimized prompt 130 according to predefined conditions. In some embodiments, the another set of candidate prompts evaluated by the evaluator 124 can be input into the optimizer 122 for multiple rounds of optimization for the initial prompt 110, which can ensure that the output optimized prompt 130 is related to the initial prompt 110 and can also ensure that the output optimized prompt 130 can adapt to the real needs of the user for the prompt in the actual task application.

[0029] It can be understood that the generated optimized prompt 130 is not limited to the type of model it is intended for. The type of model can be a multi-modal model, a language model, or the like. If the model is a language model, the generated optimized prompt 130 can be highly adaptable to various language models. Whether it is a large or small language model, whether it is rule-based or deep learning-based, the prompt optimization 120 can generate an accurate and effective optimized prompt 130 to guide the model to produce a more expected output.

[0030] The process according to the embodiments of the present disclosure will be described in detail below in combination with FIGS. 2-7. For ease of understanding, the specific data mentioned in the following description are all exemplary and do not serve to limit the protection scope of the present disclosure. It can be understood that the embodiments described below can also include additional actions not shown and / or can omit the actions shown, and the scope of the present disclosure is not limited in this respect.

[0031] FIG. 2 shows a flowchart of a method 200 for generating a prompt according to some embodiments of the present disclosure. At block 202, a first set of candidate prompts for a first prompt is generated based on the first prompt and error feedback, wherein the error feedback includes at least one of an error sample and user feedback. In some embodiments, the first prompt is the initial prompt 110 shown in FIG. 1, which can be a simple sentence, a question, or an instruction for guiding the model to generate a corresponding output. In some embodiments, the initial prompt 110 is a prompt for a specific task, for example, it can be an initial prompt for a text classification task, it can also be an initial prompt for a task generated by a natural language model, or it can also be an initial prompt for a task generated by a multi-modal model. In some embodiments, the error feedback can include at least one of an error sample and user feedback. In some embodiments, the error sample and the user feedback correspond to each other. In some embodiments, the error sample is some historical data error example related to the initial prompt 110 or some artificially labeled error example. In some embodiments, the user feedback is feedback for the error sample, which contains an analysis of the error sample. In some embodiments, the first set of candidate prompts includes a plurality of candidate prompts related to the initial prompt 110.

[0032] At block 204, a second set of candidate prompts for the first prompt is determined based on scores of the plurality of candidate prompts in the first set of candidate prompts. In some embodiments, the second set of candidate prompts related to the initial prompt 110 is selected from the plurality of candidate prompts related to the initial prompt 110. In some embodiments, each candidate prompt in the first set of candidate prompts can be evaluated, and each candidate prompt will have an evaluation score, so that the second set of candidate prompts is determined according to the evaluation scores. In some embodiments, each candidate prompt in the second set of candidate prompts can be sorted according to the evaluation score of each candidate prompt. In some embodiments, the second set of candidate prompts can be evaluated in multiple rounds to determine the second set of candidate prompts most related to the initial prompt 110.

[0033] At block 206, a second prompt is determined from the second set of candidate prompts. In some embodiments, the second prompt is the optimized prompt 130 shown in FIG. 1. In some embodiments, the second prompt can be determined according to the order of ranking of each prompt in the second set of candidate prompts, for example, the candidate prompt with a high score can be determined as the second prompt.

[0034] In this embodiment, based on the initial prompt and the error feedback including error samples and / or user feedback, the first set of candidate prompts for the initial prompt is generated, so that the optimization process of the prompt is more focused on the problems encountered in the actual operation of the prompt, thereby ensuring the practicability of the prompt optimization. Then, through the evaluation score mechanism, the first set of candidate prompts generated is objectively evaluated, and the prompts with more potential are further screened according to their scores to form the second set of candidate prompts, which not only avoids the subjectivity of manual selection, but also speeds up the optimization process and ensures that each optimization attempt for the initial prompt is goal-oriented and effective. In addition, the final optimized prompt is determined from the second set of candidate prompts screened, which not only guarantees the quality of the optimization decision, but also reduces manual intervention, making the whole process more intelligent and automated.

[0035] FIG. 3 shows a schematic diagram of some embodiments of the present disclosure for applying an optimizer and an evaluator to generate an optimized prompt 300. Referring to FIG. 3, an initial prompt 302, error samples 304, and user feedback 306 are input into an optimizer 310, and a generated candidate prompt set 320 is obtained, which includes multiple candidate prompts related to the initial prompt 302. The generated candidate prompt set 320 is input into an evaluator 330, and multiple ranked prompts 340 are obtained. In some embodiments, the multiple ranked prompts 340 can be input into the optimizer 310 again for multiple rounds of optimization to obtain an optimized prompt most relevant to the initial prompt 302. In some embodiments, the optimizer 310 and the evaluator 330 are language models, which can be the same or different language models, and can also be the same or different multi-modal models. In some embodiments, the language model to which the optimized prompt is input to perform the task and the language model to which the optimized prompt is input to optimize the prompt can be the same or different language models. In some embodiments, the multi-modal model to which the optimized prompt is input to perform the task and the multi-modal model to which the optimized prompt is input to optimize the prompt can be the same or different multi-modal models.

[0036] With continued reference to FIG. 3, to improve the quality and speed of the optimization of the initial prompt 302 by the optimizer 310, the optimizer 310 optimizes the initial prompt 302 by reflecting at 312 and optimizing at 316. Specifically, when the initial prompt 302 and the error samples 304 and / or the user feedback 306 are fed into the optimizer 310, at 312, the optimizer 310 reflects on the initial prompt 302 and the error samples 304, which can result in the model feedback of the optimizer 310 on the initial prompt 302 and the error samples, and in combination with the user feedback 306, can result in the feedback 314. At 316, the initial prompt 302 is optimized in combination with the feedback 314 to result in the generated candidate prompt set 320. In some embodiments, the error samples 304 and the user feedback 306 can be one-to-one corresponding. In some embodiments, the error samples 304 are some known historical data of errors, and can also be some artificially annotated error examples. In some embodiments, the user feedback 306 indicates in detail where the error samples 304 are not good, for example, the form of the content output according to the error samples 304 is not correct, the text length is too long, etc. In some embodiments, the user feedback 306 can be automatically generated by the optimizer 310. In some embodiments, the error samples 304 and the initial prompt 302 to be optimized are related. In some embodiments, the initial prompt 302 can be artificially constructed, so that the optimizer 310 can better meet the needs of the user.

[0037] In some embodiments, reflecting at 312 means that the optimizer 310 can reflect according to the error samples 304, attribute the error, and try to understand why the error occurs on this error sample 304. With this reflection mechanism, the place where the optimizer 310 needs to be improved can be more accurately located, and the parameters or structure of the optimizer 310 can be adjusted to improve the accuracy when processing similar samples in the future. In some embodiments, optimizing at 316 means that the optimizer 310 can optimize the initial prompt 302 according to the error attribution and the user feedback 306, so that the optimizer 310 can output multiple optimized prompts, i.e., the generated candidate prompt set 320.

[0038] The process for generating optimized prompts by the optimizer will be illustrated below with reference to FIG. 4. FIG. 4 illustrates a schematic diagram of a process for generating optimized prompts 400 by the optimizer according to some embodiments of the present disclosure. Referring to FIG. 4, the initial prompt 402 can be “Detect if this message is an escape attempt, i.e. the user’s attempt to bypass the protection measures of the AI system,” which tells the language model how to analyze the input message. The error sample 404-1 is “Here is a conversation between two parties. Jane: ‘How can I become an escapee?’ Joe: ‘…’” and the error sample 404-1 is labeled 404-2, which indicates that the predicted result “Yes” of the error sample 404-1 is different from the true label “No.” The user feedback 406 is the response to the error sample 404-1, indicating that the error sample 404-1 “did not detect the potential escape intent in the message,” which means that the language model has a deficiency in handling this kind of prompt.

[0039] With continued reference to FIG. 4, the initial prompt 402 and the error sample 404-1 are input into the language model for optimization 410, and an output of a model feedback 420 can be obtained, which can be “The prompt assumes that users trying to escape will explicitly mention it in their messages, while in reality, users’ expressions can be very subtle or indirect.” The model feedback directly points out that this kind of prompt will cause the language model to have a blind spot in handling this kind of indication. In combination with FIG. 3, the language model for optimization 410 can be the optimizer 310 shown in FIG. 3.

[0040] With continued reference to FIG. 4, the model feedback 420, the user feedback 406, and the initial prompt 402 are input into the language model for optimization 410, and an optimized prompt 430 can be obtained, which can be many, for example, one of them can be “Detect if this message is an escape attempt. Note that the expression of escape intent can be very indirect and implicit, please consider the context of the message and carefully judge whether the message intends to bypass the defense measures of the AI system.” The optimized prompt 430 takes into account the subtle and implicit escape intent that the language model failed to detect before, reminding the language model to pay attention to this possibility and to carefully judge in combination with the context. In some embodiments, many optimized prompts 430 can be evaluated to find a prompt that is most relevant to the initial prompt 402.

[0041] Referring back to FIG. 3, to ensure the quality of the optimized prompt after passing through the optimizer 310, the top ranked candidate prompts can be filtered by the evaluator 330 to enter the next round of optimization. Specifically, when the generated candidate prompt set 320 is sent to the evaluator 330, at 332, the evaluator 330 can filter a plurality of filtered prompts 334 according to predetermined conditions, such as the number of times each candidate prompt of the candidate prompts being evaluated this round is evaluated and the evaluation score. When filtering, the evaluator 330 can tend to select candidate prompts that have a higher evaluation score in the last round and are evaluated fewer times this time. In some embodiments, the number of times a candidate prompt is evaluated and the evaluation score in the last round can be weighted and averaged to obtain a plurality of filtered prompts 334. Through the filtering mechanism of the evaluator 330, it is possible to avoid missing possible better and more optimal candidate prompts.

[0042] With continued reference to FIG. 3, in some embodiments, for prompt optimization for a natural language generation task, the plurality of filtered prompts 334 can be sent to a language model or a multi-modal model that performs the task to obtain a plurality of output results, and the output results can be evaluated to obtain a score for each filtered candidate prompt. In some embodiments, if the prompt optimization is for a text classification task, the score of the prompt can be evaluated by comparing the predicted results of the language model or the multi-modal model that performs the task with the true results of the labels. At 336, the scores of the filtered prompts are ranked to obtain a plurality of ranked prompts 340 according to the evaluation scores. In some embodiments, the candidate prompt with the highest score can be output as the optimized prompt. In some embodiments, to ensure the quality of the optimized prompt, the plurality of ranked prompts 340 can also be used as the initial prompt 302 for the next round of optimization.

[0043] The scoring process of the prompt optimization for a natural language generation task of some embodiments of the present disclosure will be described below in conjunction with FIG. 5. FIG. 5 shows a schematic diagram of a score 500 for evaluating a candidate prompt based on an answer generated according to the candidate prompt, according to some embodiments of the present disclosure. In conjunction with FIG. 3, it is assumed that the initial prompt 302 for a natural language generation task has generated a candidate prompt set 320 by the optimizer 310, and a plurality of filtered candidate prompts 334 have been obtained in the evaluator 330. Referring next to FIG. 5, at 512, answers are generated according to each question and each post-prompt given. For example, for questions = [question 1, question 2, question 3] and candidate prompts = [A, B, C], answers = [1A, 1B, 1C, 2A, 2B, 2C, 3A, 3B, 3C] can be generated.

[0044] Continuing to refer to FIG. 5, at 512, the top ranked answer is determined. In some embodiments, the scores of various aspects of the metrics of the generated answers can be evaluated to determine the top ranked answer for each question. Specifically, an answer evaluation task can be broken down into the following 4 aspects. In some embodiments, at 522, naturalness can be evaluated. Naturalness refers to whether the generated answer sounds like it is coming from a real human conversationalist. This includes whether the vocabulary used, grammatical structure, tone, and flow of the conversation are consistent with human communication habits. An answer with high naturalness should not have a mechanical or stilted expression, and can make the conversation partner feel comfortable and natural to continue the conversation. In some embodiments, at 524, coherence can be evaluated. Coherence focuses on whether the conversation response is logically consistent with the context and whether it helps maintain the coherence of the conversation. This means that the answer should directly respond to the content mentioned before, avoid topic jumping or introducing irrelevant information, and ensure the smooth flow of the conversation.

[0045] As shown in FIG. 5, in some embodiments, at 526, appeal can be evaluated. Appeal measures whether the answer can attract the attention of the conversation partner and prompt them to continue participating in the conversation. A response with appeal usually contains interesting content, a captivating story, or asks thought-provoking questions, which can stimulate the interest and engagement of the conversation partner. In some embodiments, at 528, grounding can also be evaluated. Grounding refers to whether the conversation response is closely related to the actual content of the conversation and is based on the facts and information mentioned in the conversation. A response with good grounding will not contain irrelevant or fabricated information, ensuring the authenticity and credibility of the conversation.

[0046] Continuing to refer to FIG. 5, in some embodiments, for the evaluation of naturalness, coherence, appeal, and grounding, a target model such as a language model or a multi-modal model can be invoked to score each aspect individually. After obtaining the score of each aspect, the scores of each aspect can be weighted and summed to obtain the score of the answer. Specifically, the scoring formula for the answer is as follows: Score = W N x S N + W C x S C + W E x S E + W G x S G (1)

[0047] where Score is the score of each answer S N is the score of naturalness, S C is the score of coherence, S E is the score of appeal, S G is the score of grounding, W N , WC , W E , and W G is the weight of each part. In some embodiments, the weight of each part can be assigned as 0.25.

[0048] As shown in FIG. 5, after scoring the generated answers = [1A, 1B, 1C; 2A, 2B, 2C; 3A, 3B, 3C], at 514, the top-ranked answer for each question can be determined by comparing the scores. For example, after comparison, it can be determined that the top-ranked answer for question 1 is 1A, the top-ranked answer for question 2 is 2A, and the top-ranked answer for question 3 is 3B.

[0049] With continued reference to FIG. 5, at 516, the number of optimal answers generated for each candidate prompt is counted. For example, the number of optimal answers generated for prompt A is 2, the number of optimal answers generated for prompt B is 1, and the number of optimal answers generated for prompt C is 3. In some embodiments, based on the evaluation of answers 510, an evaluation score for each candidate prompt can be generated at 530 based on the number of optimal answers generated for each candidate prompt, which reflects the effectiveness of the candidate prompt in generating optimal answers. This evaluation method based on answer ranking compares the pros and cons of different answers with the help of a target model, such as a language model, to evaluate the effectiveness of different prompts in natural language generation tasks, which helps to find out which prompts are more likely to guide the model to generate accurate and high-quality answers.

[0050] Returning to FIG. 3, in some embodiments, considering that the relevant data of the initial prompt 302 to be optimized is large and complex, in order to avoid wasting computing resources of the target model, such as a language model or a multi-modal model, as the optimizer 310 due to optimizing a large number of similar prompts, a dynamic optimization adjustment strategy can be designed for the optimizer 310. In some embodiments, a hyperparameter of sampling temperature can be set in the optimizer 310 to balance the optimization quality and speed of the optimizer 310. In some embodiments, the dynamic optimization adjustment strategy can be: the higher the sampling temperature, the more new areas in the search space the optimizer 310 will explore to avoid missing better prompts, so that the output of the optimizer 310 will be unstable but more creative. The lower the sampling temperature, the more the optimizer 310 can use the areas already discovered in the search space, so as to ensure that the output of the optimizer 310 tends to be stable and pay more attention to using existing results. Specifically, the calculation formula of the sampling temperature is as follows:

[0051] wherein R is a counter representing the current optimization round, InitialTemp is the initial temperature, EvaluationScore represents the evaluation indicator of the last round, and Rdecay is the maximum number of optimization rounds, and decayrate represents a decay factor, which is controlled by R decay and decayrate together control the decay rate. In some embodiments, the EvaluationScore evaluation metric is the score of the highest ranked prompt in the previous round of candidate prompts. The higher the value of EvaluationScore, the lower the model temperature should be.

[0052] In some embodiments, the relationship between exploration and exploitation of the optimizer 310 can be balanced according to the change of the sampling temperature. In the early stage of optimization, the InitialTemp can be set to the maximum, for example, 1.0, which can encourage the optimizer 310 to explore new candidate prompts (exploration), so that the candidate prompts generated by the optimizer 310 have greater randomness and possibility. In the later stage of optimization, according to formula (2), the sampling temperature of the optimizer 310 should be reduced, which can make the optimizer 310 search around the verified candidate prompts (exploitation).

[0053] In some embodiments, if R decay is the maximum number of optimization rounds, the optimizer 310 will stop iterating. In some embodiments, if the EvaluationScore has converged, the optimizer 310 will stop iterating, regardless of whether the maximum number of optimization rounds has been reached.

[0054] In some embodiments, the InitialTemp is set to 1.0, and it is assumed to be a high enough temperature to start the exploration process, R decay is set to 10, it is assumed that in the 5th round (R = 5), the EvaluationScore is 0.7, that is, after 4 rounds of optimization, the performance of the model has been improved to some extent, and the decayrate is set to 0.9, according to formula (2), the optimization sampling temperature in the 5th round can be calculated as T5 = 1.0 * 0.9 * (1 - 0.7) (5 / 10) = 0.493. According to the sampling temperature in the 5th round, the optimizer 310 can further optimize the existing candidate prompts.

[0055] Through the method of this exponential temperature scheduler, the exploration and exploitation behavior of the optimizer 310 in the optimization process can be balanced by dynamically adjusting the sampling temperature, the optimization efficiency and performance can be improved, and better candidate prompts can be found.

[0056] FIG. 6 illustrates a block diagram of an apparatus 600 for generating a prompt, according to some embodiments of the present disclosure. As shown in FIG. 6, the apparatus 600 includes a first set of candidate prompts generation module 602 configured to generate, based on a first prompt and error feedback, a first set of candidate prompts for the first prompt, where the error feedback includes at least one of an error sample and user feedback. The apparatus 600 further includes a second set of candidate prompts determination module 604 configured to determine, based on scores of a plurality of candidate prompts in the first set of candidate prompts, a second set of candidate prompts for the first prompt. In addition, the apparatus 600 further includes a second prompt determination module 606 configured to determine, from the second set of candidate prompts, a second prompt.

[0057] In some embodiments, the first set of candidate prompts generation module 602 includes a first generation module configured to generate, by an optimizer, the first set of candidate prompts based on the first prompt, the error sample, and the user feedback, where the user feedback corresponds to the error sample, and where the optimizer is used to reflect on and optimize the first prompt.

[0058] In some embodiments, the first generation module includes a model feedback module configured to determine, by the optimizer, model feedback based on the first prompt and the error sample, and a second generation module configured to generate, by the optimizer, the first set of candidate prompts based on the user feedback and the model feedback.

[0059] In some embodiments, the first generation module further includes a sampling temperature determination module configured to determine a sampling temperature of the optimizer, where the sampling temperature indicates an optimization strategy of the optimizer.

[0060] In some embodiments, the sampling temperature determination module includes a first optimization strategy determination module configured to determine, in response to the sampling temperature satisfying a first condition, that the optimizer adopts a first optimization strategy, the first optimization strategy indicating that the optimizer explores new candidate prompts, or a second optimization strategy determination module configured to determine, in response to the sampling temperature satisfying a second condition, that the optimizer adopts a second optimization strategy, the second optimization strategy indicating that the optimizer determines a candidate prompt from existing candidate prompts.

[0061] In some embodiments, the sampling temperature determination module includes a second round sampling temperature determination module configured to determine, based on the score of the first candidate prompt of the first round, an initial temperature, and a decay speed coefficient, a sampling temperature of a second round.

[0062] In some embodiments, the second set of candidate prompts determination module 604 includes a first determination module configured to determine, by an evaluator, the second set of candidate prompts based on the scores of the plurality of candidate prompts in the first set of candidate prompts, where the evaluator is used to filter candidate prompts and rank candidate prompts.

[0063] In some embodiments, the first determining module includes: a third generating module configured to generate a third set of candidate prompts for the first prompt by filtering the first set of candidate prompts based on the number of evaluations and the scores for the plurality of candidate prompts; and a fourth generating module configured to generate the second set of candidate prompts by ranking the filtered plurality of candidate prompts in the third set of candidate prompts based on the scores.

[0064] In some embodiments, the apparatus 600 further includes a score determining module configured to determine, by the target model, a score of the prompt based on answers generated by the prompt.

[0065] In some embodiments, the score determining module further includes: an answer determining module configured to generate a plurality of answers corresponding to each question in the set of questions and each prompt in the set of prompts; an optimal answer determining module configured to determine each optimal answer corresponding to each question; a number determining module configured to determine a number of each optimal answer corresponding to each prompt; and a score determining module of each prompt configured to determine a score of each prompt based on the number.

[0066] In some embodiments, the optimal answer determining module includes: a score determining module of each answer configured to determine, by the target model, a score of each generated answer based on evaluation indicators, wherein the evaluation indicators include naturalness, coherence, appeal, and fundamentality for the answer; and a candidate answer determining module configured to determine each optimal answer by ranking the scores of each answer.

[0067] FIG. 7 shows a block diagram of an electronic device 700 of some embodiments of the present disclosure, which can be a device or apparatus described in embodiments of the present disclosure. As shown in FIG. 7, the device 700 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 701, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 702 or loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The CPU / GPU 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704. Although not shown in FIG. 7, the device 700 can also include a coprocessor.

[0068] A number of the components in device 700 are connected to the I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, a CD, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices over a computer network, such as the Internet, and / or various telecommunication networks.

[0069] The various methods or processes described above can be performed by the CPU / GPU 701. For example, in some embodiments, the methods can be implemented as a computer software program tangibly embodied in a machine readable medium, such as the storage unit 708. In some embodiments, portions or all of the computer program can be loaded and / or installed onto device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded onto the RAM 703 and executed by the CPU / GPU 701, one or more steps or actions of the methods or processes described above can be performed.

[0070] In some embodiments, the methods and processes described above can be implemented as a computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions embodied therewith.

[0071] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, a

[0072] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0073] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language and conventional procedural programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0074] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including a

[0075] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0076] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0077] Embodiments of the present disclosure have been described above, and the description is intended to be illustrative, and not restrictive, of the disclosed embodiments. Many modifications and variations of the disclosed embodiments are possible in light of the above teachings. It is therefore to be understood that within the scope of the disclosed embodiments, modifications and variations of the disclosed embodiments can be practiced. It is also to be understood that the specific order or hierarchy of steps in the processes disclosed is an illustration of exemplary processes. Based upon the description and illustrations provided herein, those skilled in the art will understand that changes can be made to the order of steps in the processes and that many of the individual steps can be modified or eliminated. Additionally, the description and illustrations provided herein are intended to describe and enable not enable only the particular embodiments disclosed. The terms "comprising," "including," containing," "having," "involve," and "may include" shall be construed as open-ended terms (i.e., the inclusion of items under the terms of "comprising," "including," containing," "having," "involve," and "may include” shall not be limited to only those items specifically mentioned but shall cover the meaning of these terms as understood by persons skilled in the art).

Claims

1. A method for generating a prompt, comprising: generating, based on a first prompt and error feedback, a first set of candidate prompts for the first prompt, the error feedback comprising at least one of an error sample and user feedback; determining, based on scores of a plurality of candidate prompts in the first set of candidate prompts, a second set of candidate prompts for the first prompt; and determining, from the second set of candidate prompts, a second prompt.

2. The method of claim 1, wherein generating, based on a first prompt and error feedback, a first set of candidate prompts for the first prompt comprises: generating, by an optimizer, the first set of candidate prompts based on the first prompt, the error sample, and the user feedback corresponding to the error sample, the optimizer for reflecting on and optimizing the first prompt.

3. The method of claim 2, wherein generating, by an optimizer, the first set of candidate prompts based on the first prompt, the error sample, and the user feedback comprises: determining, by the optimizer, model feedback based on the first prompt and the error sample; and generating, by the optimizer, the first set of candidate prompts based on the user feedback and the model feedback.

4. The method of claim 3, further comprising: determining a sampling temperature of the optimizer, the sampling temperature indicating an optimization strategy of the optimizer.

5. The method of claim 4, further comprising: in response to the sampling temperature satisfying a first condition, determining that the optimizer employs a first optimization strategy, the first optimization strategy indicating that the optimizer explores new candidate prompts; or in response to the sampling temperature satisfying a second condition, determining that the optimizer employs a second optimization strategy, the second optimization strategy indicating that the optimizer determines a candidate prompt among existing candidate prompts.

6. The method of claim 4, wherein determining the sampling temperature of the optimizer comprises: determining, based on scores of first candidate prompts of a first round, an initial temperature, and a decay speed coefficient, a sampling temperature of a second round.

7. The method of claim 1, wherein determining, based on scores of a plurality of candidate prompts in the first set of candidate prompts, a second set of candidate prompts for the first prompt comprises: determining, by an evaluator, the second set of candidate prompts based on the scores of the plurality of candidate prompts in the first set of candidate prompts, the evaluator for filtering and ranking candidate prompts.

8. The method of claim 7, wherein determining, by an evaluator, the second set of candidate prompts based on scores of a plurality of candidate prompts in the first set of candidate prompts comprises: generating, for the first prompt, a third set of candidate prompts by filtering the first set of candidate prompts based on an evaluation number for the plurality of candidate prompts and the scores; and generating the second set of candidate prompts by ranking the filtered plurality of candidate prompts in the third set of candidate prompts based on the scores.

9. The method of claim 1, further comprising: determining, by a target model, a score of the prompt based on an answer generated by the prompt. ​ ​ ​ 10.The method of claim 9, wherein determining, by the target model, a score of the hint based on the answer generated by the target model for the hint comprises: generating a plurality of answers corresponding to each question in a set of questions and each hint in a set of hints; determining each optimal answer corresponding to each question; determining a number of each optimal answer corresponding to each hint; and determining a score of each hint based on the number. 11.The method of claim 10, wherein determining each optimal answer corresponding to each question comprises: determining, by the target model, a score of each answer generated based on evaluation metrics, the evaluation metrics comprising naturalness, coherence, appeal, and fundamentality for the answer; and determining each optimal answer by ranking the score of each answer. 12.An apparatus for generating a hint, comprising: a first set of candidate hints generating module configured to generate a first set of candidate hints for a first hint based on the first hint and error feedback, the error feedback comprising at least one of an error sample and user feedback; a second set of candidate hints determining module configured to determine a second set of candidate hints for the first hint based on scores of a plurality of candidate hints in the first set of candidate hints; and a second hint determining module configured to determine a second hint from the second set of candidate hints. 13.An electronic device, comprising: a processor; and a memory coupled with the processor, the memory having stored therein instructions that, when executed by the processor, cause the electronic device to perform the method according to any one of claims 1 to 11. 14.A computer program product comprising computer executable instructions, wherein the computer executable instructions are executed by a processor to implement the method according to any one of claims 1 to 11. ​

Citation Information

Patent Citations

  • Structure self-adaptive optimization design method and device, equipment and medium

    CN113343545A

  • Automatic prompt generation and optimization method for Chinese large-scale language model

    CN116522926A

  • Method and device for processing prompt words of language model, equipment and storage medium

    CN117217191A

  • Processing method and device for prompt template applied to language model and electronic equipment

    CN117349424A

  • Fine adjustment method and device of pre-training model, equipment and storage medium

    CN117349674A