Large Model Automatic Prompt Tuning Method, Device and Medium Optimized by Gradient

The gradient-based optimization framework for prompt tuning addresses inefficiencies in APE by focusing on the weakest rule, enhancing convergence and effectiveness in complex tasks with a unified scoring system, ensuring explainability.

CN119962681BActive Publication Date: 2025-07-15SOUTHWESTERN UNIV OF FINANCE & ECONOMICS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510041210.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-07-15
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

Traditional APE technology lacks unified task evaluation criteria in generative tasks, and the long prompt word optimization process is complex and lacks interpretability, resulting in the optimization process being inaccurate and efficient enough.

Method used

The gradient optimization method is adopted to score the prompt words and question-and-answer model output of each test data by constructing a unified evaluation model, modify only the worst rules, use examples to guide optimization, realize the automated tuning of long propts, and form an interpretable optimization process.

Benefits of technology

The pertinence and effectiveness of the long-prompt optimization process are improved, the convergence problem is solved, and efficient and accurate prompt word optimization is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962681B_ABST
    Figure CN119962681B_ABST
Patent Text Reader

Abstract

The present invention discloses an automatic prompt tuning method, device, and medium using a large model optimized by gradients. By evaluating the execution rules in the prompt words of each test data and the answers output by the corresponding question-and-answer model, the total evaluation score of the prompt words and the evaluation scores of each execution rule are optimized by gradient descent. During the tuning process, it is determined whether the current average score and round meet the set threshold; if not, the prompt words of the current round are output as the optimal prompt words; if so, the optimization direction of the optimal prompt words is selected, the optimization direction is refined, and new prompt words are generated based on the refined optimization direction. This step is repeated until the set threshold is met, the optimization round is stopped, and the optimal prompt words are output. It can efficiently and accurately solve the convergence problem in the optimization process of long prompt words, and improve the pertinence and effectiveness of prompt word tuning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of model question-answering evaluation, and specifically to an automatic prompt tuning method, device, and medium for large models optimized by gradients. Background Art

[0002] With the rapid development of artificial intelligence, large language models (LLMs) have become an important research direction in the field of natural language processing. Large language models (LLMs) refer to language models containing hundreds of billions (or more) of parameters and trained on a large amount of text data. These models have powerful language understanding and generation capabilities and show good results in tasks such as machine translation, text summarization, and dialogue generation. Large language models are also gradually playing a crucial role in various industries. A prompt is the input text or instruction provided to a generative AI model. It can be a question, a sentence, a paragraph, or a complete passage, used to guide the model to generate a corresponding answer or output. The role of the prompt is to tell the model what task needs to be completed or what result needs to be produced.

[0003] The prior art has proposed the concept of Chain of Thoughts Prompting, which enhances complex reasoning capabilities by introducing intermediate reasoning steps. This method can be combined with few-shot prompting to achieve better results, especially suitable for tasks that require more complex reasoning. That is, describe a series of intermediate reasoning steps and manually construct the reasoning process, enabling large models to achieve more ideal results in tasks such as arithmetic reasoning, common sense reasoning, and symbolic reasoning through learning these reasoning processes. It can be seen that the task execution effects of the same base model may vary greatly under different prompts. However, the process of constructing suitable prompts is a time-consuming and laborious process. Therefore, Zhou et al. (2022) proposed Automatic Prompt Engineer, which is a framework for automatic instruction generation and selection. The instruction generation problem is formulated as a natural language synthesis problem, and LLMs are used as the solution to optimize the problem as a black box to generate and search for candidate solutions. After that, many people have made certain optimizations to the deficiencies of this work. Add instructions and corresponding scores during the guidance process to let the model understand the process of score improvement to improve the optimization of the prompt; add an analysis of the deficiencies of the original instructions during the guidance process and let it guide the optimization of the new prompt; solve the non-convergence problem of long prompt optimization by splitting the original prompt into several sentences and only modifying one sentence each time.

[0004] However, after multiple optimizations, there are still some obvious deficiencies in the APE technology, which are mainly manifested in the following aspects:

[0005] When traditional APE technology solves non-standard tasks such as generative tasks, such as translation tasks, text summarization, dialogue systems, etc., different evaluation rules often need to be set separately for different tasks during the optimization process, that is, there is a lack of a unified task evaluation standard.

[0006] Secondly, when performing complex tasks, the input prompts are usually long and complex. Although previous technologies have alleviated the convergence problem in long optimization to a certain extent, during the optimization process, since the model optimizes each sentence separately, not only is the process complex and cumbersome, bringing unnecessary overhead, but the examples for guiding modifications are not specifically targeted at the problematic sentences, but more of a general and broad analysis, weakening the guiding ability for the optimal prompts.

[0007] Finally, the optimization process of traditional APE is actually a black-box process, resulting in the lack of interpretability of our optimization process. Summary of the Invention

[0008] The technical problem to be solved by the present invention is the lack of interpretability in the traditional optimization process. The purpose is to provide a method, device, and medium for automatically tuning large model prompts using gradient optimization. By proposing an automated prompt tuning framework for long prompts in complex scenarios, under this framework, the evaluation model scores the execution rules in the prompts of each test data and the corresponding answers output by the question-answering model, and performs gradient descent tuning on the total evaluation score of the prompts and the evaluation scores of each execution rule. In each modification process, we only modify the worst rule and further guide it using specific examples of this rule. Therefore, the convergence problem in the long prompt optimization process can be solved more efficiently and accurately. The framework constructed by this solution is an interpretable process for the ape process, improving the pertinence and effectiveness of prompt tuning.

[0009] The present invention is achieved through the following technical solutions:

[0010] The first aspect of the present invention provides a method for automatically tuning large model prompts using gradient optimization, including the following specific steps:

[0011] Initialize parameter settings;

[0012] Obtain the execution rules in the prompts of each test data and the corresponding answers output by the question-answering model;

[0013] Build an evaluation model to score the execution rules in each test data prompt and the corresponding answers output by the Q&A model, obtaining the total evaluation score of the prompt and the evaluation scores of each execution rule;

[0014] Perform gradient descent optimization on the total evaluation score of the prompt and the evaluation scores of each execution rule:

[0015] Obtain the average score corresponding to the prompt of the current round, and determine whether the current average score and round meet the set threshold;

[0016] If not, output the prompt of the current round as the optimal prompt;

[0017] If so, select the optimization direction of the optimal prompt, refine the optimization direction, generate a new prompt based on the refined optimization direction, repeat this step until the set threshold is met, stop the optimization round, and output the optimal prompt.

[0018] Furthermore, the initialization parameter settings specifically include: initializing the current round, the maximum number of iteration rounds, the initial prompt effect score, the target threshold, the step size for each optimization, the number of boundary attempt times, and the initial prompt.

[0019] Furthermore, the target threshold is used to describe the effect score that the user expects the prompt to reach, the step size is used to represent the number of words added, deleted, or modified each time, the number of boundary attempt times is used to represent the number of rounds when the score remains unchanged and reaches the model upper limit, and the initial prompt includes background introduction, task description, execution rules, and task subject.

[0020] Furthermore, when the evaluation model scores, the evaluation model scores according to the answers output by the Q&A model and the execution rules in the prompt. The scoring process includes: calculating the total score of each test data under the prompt, taking the average value of n data, and obtaining the total evaluation score of the prompt and the evaluation scores of each rule based on the average value of n data.

[0021] Furthermore, when the evaluation model scores the execution rules in each test data prompt and the corresponding model output answers, it also includes:

[0022] Judge whether the average score and round of the evaluation score of the current round meet the set threshold. If the average score does not change for k consecutive rounds, it is determined that the optimal boundary of the model is reached, and the prompt of the current round is output as the optimal prompt.

[0023] Furthermore, the selection of the optimization direction of the optimal prompt specifically includes: obtaining all the rules of the current prompt, generating multiple forward optimization spaces according to all the rules, and selecting the rule with the lowest score among all the rules as the current optimization direction.

[0024] Furthermore, the refinement and optimization directions specifically include:

[0025] Advance according to the current optimization direction, and obtain the number of words that can be modified each time as the step size for each step;

[0026] Conduct a score analysis on the prompt words corresponding to the rule with the lowest score;

[0027] Select the 3 pieces of test data with the lowest scores for the rule and the corresponding answers as examples;

[0028] The evaluation model analyzes based on the examples, and the Q&A model optimizes the step size of the rule based on the analysis results to generate N new prompt words.

[0029] Furthermore, each time the step size of the rule is optimized, only the item with the lowest average score in the rule in the prompt is modified.

[0030] The second aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements a method for automatically tuning prompts of a large model using gradient optimization.

[0031] The third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements a method for automatically tuning prompts of a large model using gradient optimization.

[0032] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0033] 1. The present invention proposes an automated prompt tuning framework for long prompts in complex scenarios. Our framework has a unified format for all tasks, and a unified evaluation method is proposed for our framework.

[0034] 2. Under the framework of the present invention, only the key part of the execution rules that guides whether the model can accurately generate answers is tuned. And in each modification process, we only modify the worst rule and use specific examples of this rule to further guide it. Therefore, the convergence problem in the long prompt optimization process can be solved more efficiently and accurately.

[0035] 3. The ape process under the framework of the present invention is an interpretable process, further improving the pertinence and effectiveness of prompt tuning. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings. In the drawings:

[0037] Figure 1 is the tuning process in the embodiments of the present invention;

[0038] Figure 2 is the prompt optimization process in the embodiments of the present invention;

[0039] Figure 3 is the prompt optimization and simplification process in the embodiments of the present invention. Detailed implementation manners

[0040] To make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and the drawings. The illustrative embodiments of the present invention and their descriptions are only used to explain the present invention and are not used to limit the present invention.

[0041] As a possible implementation manner, as Figure 1 shown, to solve the problems existing in the existing APE optimization process, this embodiment provides a method for automatically prompting and tuning large models using gradient optimization. First, this embodiment proposes a basic framework for long prompts in complex scenarios and proposes a unified evaluation rule for this framework. In this embodiment, all prompts can be regarded as four parts, namely background introduction (env), task description (task), execution rules (rules), and task subject (object). We believe that rules are the key to guiding whether the model can accurately generate answers. Therefore, when evaluating the quality of generative tasks, we use the evaluation model to score the satisfaction of the answer and the rule, and thus obtain a standardized evaluation rule.

[0042] Based on this, our specific optimization method is proposed. For the test set, we can obtain the overall average score and the average score of each rule. In each modification, we only modify the item with the lowest average score in the rules of the prompt, and input k test examples with the lowest score and score analysis of this rule to the model, so that the model can better understand the deficiencies and use them to modify this rule. In this process, we not only achieve the interpretability of modifying sentences, but also achieve the interpretability of our modification process.

[0043] Finally, we regard the above process as an optimization process of gradient descent and abstract it into the mathematical space to achieve the visualization and stability of the optimization process. If each rule is regarded as the forward direction of optimization in the optimization process, and the number of words that can be modified each time is regarded as the step size of each step, then this process is abstracted into a process of gradient descent of the score corresponding to each prompt towards the optimal point. At the initial point, we can obtain the average score corresponding to prompt0 and the rule that should be modified. Then, each autonomous modification of the model can be regarded as continuous attempts in this direction, and the prompt closest to the optimal point obtained by this modification is used as the new prompt. (Our gradient descent idea) The above process is repeated until the optimal boundary of this model is reached or the upper limit of the specified execution round is reached. The implementation process is as Figure 2 shown. If we confirm the forward direction at each step, Figure 2 the process shown can be transformed into the process as Figure 3 shown. After transformation, the performance of the model is improved. The closer the initial prompt0 is to the optimal boundary of this model, and the closer the optimal boundary of the model is to the optimal point.

[0044] Specifically, this embodiment provides a detailed optimization step:

[0045] 1. Initial parameter setting

[0046] Set the current round t = 0, the maximum number of iterations T, the initial prompt effect score aver-score_0t = 0, the target threshold threshold = P, the step size l for each optimization, the number of boundary attempt times k, and the initial prompt. Among them, the score of the overall effect of the initial prompt is aver-score_00 = 0. The threshold threshold describes the effect score that the user expects the prompt to reach. The step size represents the number of words added, deleted, or modified each time. The number of boundary attempt times represents the number of rounds when the score remains unchanged and has reached the upper limit of the model. The initial prompt structure includes background introduction, task description, execution rules, and task body.

[0047] 2. Testing and scoring

[0048] The prompt for each piece of test data (hereinafter denoted as promptt-1) and the corresponding output of the Q&A model are input into the evaluation model (the evaluation model used in this embodiment can be GPT-4). The evaluation model scores based on the answers output by the Q&A model and the execution rules (rules) in promptt-1. The full score for each rule is 10 points. Calculate the total score of each piece of test data under promptt-1, and take the average of n pieces of data to obtain the total evaluation score aver-score_0t of the evaluation model for promptt-1 and the evaluation scores aver-score_it of each rule.

[0049] 3. Gradient Descent Optimization Process

[0050] In each round of optimization, the following detailed steps are carried out:

[0051] 3.1 Condition Judgment

[0052] Analyze the score of the current round, and judge whether the current average score aver-score_t and the round meet the predetermined conditions. If satisfied, proceed to the next step; if not satisfied, output the prompt-t of the current round as the optimal prompt. Note that if the average score does not change for k consecutive rounds, we consider that the optimal boundary of the model has been reached (i.e., Figure 1 Figure 2 the optimal boundary that this model can reach), and also output the prompt-t of the current round as the optimal prompt

[0053] 3.2 Select the Optimization Gradient

[0054] Regard all rules as the forward optimization space for this task, and select the rule with the lowest score as the current optimization direction. That is, in gradient descent, select the direction with the steepest gradient for optimization.

[0055] 3.3 Use the evaluation to further refine the optimization direction and optimize according to the specified step size

[0056] Further analyze the test data corresponding to the rule with the lowest score. After analysis, select the 3 pieces of test data and answers with the lowest score of this rule among all test data as few-shot. The evaluation model gives corresponding analysis according to the input content, and the Q&A model uses the analysis to modify and optimize this rule N times according to the specified step size to generate N new prompt-t. That is, in gradient descent, by walking a certain step size along the gradient direction, new points are generated.

[0057] 3.4 Select Optimization

[0058] Test the N new prompt-t again, and evaluate and score them by the evaluation model. If none of the N new prompt-t shows improvement, return to modify them again until the score improves and then proceed to the next step. If the number of times of re-modification reaches the boundary attempt number k, it is considered that the optimal boundary within the attempt times has been reached, and the optimization is terminated. Among these N prompts, select the prompt with the highest score as the updated prompt-t.

[0059] 4. Update and loop:

[0060] Update the round t, and repeat the loop in step 3 until the termination condition is met or the maximum iteration round T is reached.

[0061] As a possible implementation manner, this embodiment provides a specific application scenario, which includes multiple experimental contents. The experiment focuses on a complex task, that is, recommending financial products according to the user profile. Each prompt is composed of four parts: background introduction (env), task description (task), execution rules (rules), and task subject (object). Experiments 1 to 3 aim to show that the prompt tuning scheme we proposed runs under the condition of known standard answers, and its effect is not inferior to the traditional scheme, and it can converge relatively quickly. Experiment 4 shows the applicability of our scheme in a complex background without standard answers.

[0062] Experiment 1: Apply the prompt tuning scheme we proposed to the traditional background of standard answers.

[0063] Experiment 2: Use the OPRO prompt tuning scheme mentioned by Chengrun et al. in "LARGE LANGUAGE MODELS AS OPTIMIZERS". Input the same original prompt and customer feature data into the large model, and let the large model select investment plans for each customer to obtain the accuracy rate under the standard answer. Input this prompt, the corresponding investment plan selected for each customer, the accuracy rate, and the customer's real choice into GPT4, and let it give a score and modification suggestions for the entire execution rules. Input this modification suggestion and score into the large model, let it optimize and generate a new prompt by itself, then apply the new prompt to the recommendation of investment plans to obtain a new accuracy rate, and add the newly generated prompt, investment plan, and accuracy rate in this iterative process to the set of the original prompt, investment plan, and accuracy rate. Input them together with the customer's real choice into GPT4 again, and let GPT4 score and give modification suggestions again. Repeating this process, there are mainly four major differences from the scheme we proposed: First, it can only be applied to tasks with standard answers; Second, the GPT4 scoring is for the entire rules, rather than scoring each rule separately; Third, GPT4 gives modification suggestions instead of the analysis mentioned in our scheme. Correspondingly, the large model only explores the process of modifying the prompt by itself once in each round, while we explore three times in each round and take the effective modification; Fourth, the dataset input to GPT4 each time is the set of prompts, investment plan selections, and accuracy rates generated in all historical iterative processes.

[0064] Experiment 3: Adopt the two-step task prompt tuning scheme for meta-prompting LLMs proposed by Qinyuan Ye et al. in "Prompt Engineering a Prompt Engineer". Input the same original prompt and customer feature data into the large model, let the large model select an investment plan for each customer, and obtain the accuracy rate under the standard answer. Then, proceed in two steps. First, input this prompt, the corresponding investment plan selected for each customer, the accuracy rate, and the customer's actual choice into GPT4, and ask GPT4 whether this prompt needs to be modified. Second, if modification is needed, let GPT4 give a score for the entire execution rules and positive modification suggestions. Next, input this modification opinion and score into the large model, let it optimize and generate a new prompt by itself, then apply the new prompt to the selection of investment plans to obtain a new accuracy rate, and input this new prompt, investment plan, accuracy rate, and the customer's actual choice into GPT4 again, let GPT4 score again and give modification opinions. Repeat this process. There are mainly four differences from the scheme we proposed: First, it can only be applied to tasks with standard answers; Second, GPT4 scores the entire rules, rather than scoring each rule separately; Third, GPT4 gives positive modification suggestions instead of the analysis mentioned in our scheme. Correspondingly, the large model only explores the process of modifying the prompt by itself once in each round, while we explore three times in each round and take the effective modification; Fourth, the analysis process of GPT4 is carried out in two steps. First, it is necessary to ask whether this prompt needs to be modified.

[0065] Experiment 4: Apply the prompt tuning scheme we proposed to the background of non-traditional tasks without providing standard answers.

[0066] Table 1: Accuracy rate of each round in the iteration process of this scheme and the control scheme under the background of providing standard answers

[0067]

[0068] Table 2: Score of each rule in the iteration process of this scheme under the background of not providing standard answers

[0069] Rules The first round The second round The third round Total score 77.8 78.4 85 The first rule 85 85 90 The second rule 78 80 85 The third rule 82 70 80 The fourth rule 74 82 88 The fifth rule 70 75 82

[0070] According to the above content, it can be seen that the method adopted in this embodiment improves the pertinence and effectiveness of prompt tuning.

[0071] As a possible implementation manner, this embodiment provides a tuning process in a specific scenario corresponding to the experimental process of the above experiment, such as Figure 1As shown, in the context of not providing the standard answer, the maximum number of iteration rounds is set to 8, the number of boundary attempt times is 3, and the maximum threshold score is 85. If the result does not improve after three attempts or the effect decreases in reverse, or the maximum number of iteration rounds is reached or the maximum threshold score is reached, the test will be terminated.

[0072] Corresponding to the experimental process of Experiment 4:

[0073] The first round:

[0074] First, input the original prompt "According to the customer's financial situation, investment goals, and risk preferences, select a suitable investment plan; ensure that the selected investment plan matches the customer's long-term financial plan and life goals; the selected investment plan should include considerations of market dynamics and economic indicators to achieve diversified asset allocation; according to the customer's investment goals, select a plan that matches the customer's expected investment return and time frame; for the customer's personalized financial situation and liquidity needs, develop a flexible and adjustable investment plan" and customer characteristic data (e.g., User 1: "Age": "30", "Gender": "Male", "Annual Income": "100,000", "Occupation": "Programmer", "Marital Status": "Unmarried", "Whether there is a personal housing": "No", "Risk tolerance": "Medium", "Investment goal": "Long-term appreciation") to the large model, and let the large model formulate an investment plan for each customer. Input this original prompt and the investment plan formulated by the large model to GPT4, and let GPT4 score each execution rule of the original prompt and give corresponding analysis (e.g., The 1st execution rule:

Score: 85; Analysis: The investment plan generally considers the customer's financial situation, investment goals, and risk preferences, but in some users' plans, the matching degree between risk tolerance and the investment portfolio is slightly insufficient.

[0075] Next, enter the process of modifying the execution rules. Take the scores and analysis given by GPT4 for each execution rule, and then use GPT4 to propose corresponding optimization suggestions according to the analysis. Finally, input the analysis and optimization suggestions into the large model, and let the large model modify the prompt by itself accordingly. The requirement is to only modify the execution rule with the lowest score given by GPT4.

[0076] Finally, let the large model formulate the customer's investment plan again according to the new prompt obtained after modification, and input this plan and the new prompt into GPT4 again, and let GPT4 score this new plan. Judge whether the overall score has improved. If it has improved and neither the maximum number of iteration rounds nor the maximum threshold score has been reached, iterate to the second round. Otherwise, judge whether the termination condition is reached. If it is reached, terminate.

[0077] Second round: First, let the large model propose modification suggestions and modify the prompting words according to the analysis of GPT4. Then, let the large model formulate an investment plan based on the modified prompting words. Finally, input the new prompting words and the formulated investment plan into GPT4 to let it score each execution rule. Judge whether the overall score has improved. If not, let the large model modify again, and judge whether the termination condition is reached and whether to enter the third round of iteration.

[0078] Third round:......

[0079] Repeat this process to obtain the optimal prompting words and formulate the optimal investment plan.

[0080] When comparing this solution with the control solution in the context of providing the standard answer, set the maximum number of iteration rounds to 11, the number of boundary attempt times to 3, and the maximum threshold score to 0.9. If the result does not improve or the effect decreases in reverse after more than three times, or the maximum number of iteration rounds is reached or the maximum threshold score is reached, terminate the test.

[0081] Corresponding to the experimental process of Experiment 1:

[0082] First round:

[0083] First, input the original prompt "Based on the customer's financial situation, investment goals, risk preferences, etc., select a suitable investment plan; ensure that the selected investment plan matches the customer's long-term financial plan and life goals; the selected investment plan should include considerations of market dynamics and economic indicators to achieve diversified asset allocation; according to the customer's investment goals, select a plan that matches the customer's expected investment return and time frame; for the customer's personalized financial situation and liquidity needs, formulate a flexible and adjustable investment plan" and customer characteristic data (e.g., User 1: "Age": "30", "Gender": "Male", "Annual Income": "100,000", "Occupation": "Programmer", "Marital Status": "Unmarried", "Whether there is a personal housing": "No", "Risk Tolerance": "Medium", "Investment Goal": "Long-term appreciation") to the large model, and let the large model select an investment plan for each customer (including: A: 90% stock investment, 10% bond investment; B: 70% stock investment, 30% bond investment; C: 50% stock investment, 50% bond investment; D: 30% stock investment, 70% bond investment; E: 10% stock investment, 90% bond investment.), and compare it with the customer's actual choice to obtain the recommendation accuracy rate. Input this original prompt, investment plan selection, accuracy rate, and the customer's actual choice to GPT4, let GPT4 score each execution rule of the original prompt and give corresponding analysis (e.g., The first execution rule: [Score: 85; Analysis: The investment plans of most users match their risk tolerance and investment goals, but the plan selections of some users are not precise enough.]), and then input the analysis into GPT4 to propose optimization suggestions.

[0084] Next, enter the process of modifying the execution rules. According to the analysis and optimization suggestions feedback by GPT4, the large model proposes three modification opinions and makes three modifications to the prompt, selects the modification with the highest accuracy rate as the effective modification, discards the other two, and during this process, discards those with low scores or no improvement in accuracy rate during the modification process, and increases the exploration times. Input the scores and analysis given by GPT4 for each execution rule into the large model, and let the large model modify the prompt by itself accordingly. The requirement is to only modify the execution rule with the lowest score given by GPT.

[0085] Finally, let the large model recommend the customer's investment plan again based on the new prompt obtained from the effective modification by itself, and obtain the accuracy rate. Determine whether this accuracy rate has been improved. If it has been improved and neither the maximum iteration round nor the maximum threshold score has been reached, then iterate to the second round. Otherwise, determine whether the termination condition has been reached. If it has, then terminate.

[0086] Second round: First, input the new prompt, the investment plan options given by the large model, the accuracy rate, and the customer's actual choice into GPT4 together. Let GPT4 score each execution rule of the new prompt and give corresponding analysis. Next, let the large model propose modification suggestions by itself based on the analysis of GPT4 and modify the prompt, and test the modification. Select one of the three modifications (the one with the highest accuracy rate) as the effective modification. If the accuracy rate does not improve, return for re-modification until there is an improvement or the maximum number of attempts, 3 times, is reached. Finally, check whether the termination condition is met to determine whether to enter the third-round iteration.

[0087] Third round:......

[0088] Repeat this process to obtain the optimal prompt and the recommended accuracy rate.

[0089] As a possible implementation manner, this embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements a method for automatically tuning the prompts of a large model using gradient optimization.

[0090] As a possible implementation manner, this embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements a method for automatically tuning the prompts of a large model using gradient optimization.

[0091] The above specific implementation manners further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific implementation manners of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An automatic prompt tuning method for large models optimized by gradient, characterized in that It includes the following specific steps: Initialize parameter settings; Obtain the execution rules and the corresponding Q&A model output answers in each test data prompt; Construct an evaluation model to score the execution rules and the corresponding Q&A model output answers in each test data prompt, obtaining the total evaluation score of the prompt and the evaluation scores of each execution rule; Perform gradient descent optimization on the total evaluation score of the prompt and the evaluation scores of each execution rule: Obtain the average score corresponding to the prompt in the current round, and determine whether the current average score and round meet the set threshold; If not, output the prompt in the current round as the optimal prompt; If so, select the optimization direction of the optimal prompt, refine the optimization direction, generate a new prompt based on the refined optimization direction, repeat this step until the set threshold is met, stop the optimization round, and output the optimal prompt; Among them, the refinement of the optimization direction specifically includes: Advance according to the current optimization direction, and obtain the number of words that can be modified each time as the step size for each step; Conduct a score analysis on the prompt corresponding to the rule with the lowest score; Select the 3 test data with the lowest scores for the rule and the corresponding answers as examples; The evaluation model analyzes according to the examples, and the Q&A model optimizes the step size of the rule based on the analysis results to generate N new prompts.

2. The method for automatically prompting and optimizing a large model using gradient optimization according to claim 1, wherein, The initialization of parameter settings specifically includes: initializing the current round, the maximum number of iteration rounds, the initial prompt effect score, the target threshold, the step size for each optimization, the number of boundary attempts, and the initial prompt.

3. The method for automatically prompting and optimizing a large model using gradient optimization according to claim 2, wherein The target threshold is used to describe the effect score that the user expects the prompt to reach. The step size is used to represent the number of words added, deleted, or modified each time. The number of boundary attempts is used to indicate how many rounds the score remains unchanged and has reached the model upper limit. The initial prompt includes background introduction, task description, execution rules, and task subject.

4. The method for automatically prompting and optimizing a large model using gradient optimization according to claim 2, wherein When the evaluation model scores, the evaluation model scores according to the answer output by the Q&A model and the execution rules in the prompt. The scoring process includes: calculating the total score of each test data under the prompt, taking the average value of n data, and obtaining the total evaluation score of the prompt and the evaluation scores of each rule by the evaluation model based on the average value of n data.

5. The method for automatically prompting and optimizing a large model using gradient optimization according to claim 1, wherein When the evaluation model scores the execution rules and the corresponding model output answers in each test data prompt, it also includes: Judge whether the average score and round of the evaluation scores in the current round meet the set threshold. If the average score does not change for k consecutive rounds, it is determined that the optimal boundary of the model is reached, and the prompt in the current round is output as the optimal prompt.

6. The automatic prompt tuning method for large models optimized by gradient according to claim 1, wherein The selection of the optimization direction of the optimal prompt specifically includes: obtaining all the rules of the current prompt, generating multiple forward optimization spaces according to all the rules, and selecting the rule with the lowest score among all the rules as the current optimization direction.

7. The method for automatically prompting and optimizing a large model using gradient optimization according to claim 1, characterized in that During each optimization of the step size of the rule, only the item with the lowest average score in the rule in the prompt is modified.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for automatically tuning large model prompts using gradient optimization according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method for automatically prompting and optimizing a large model using gradient optimization as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and system for optimizing cue word, electronic equipment and storage medium

    CN118428492A