Large language model automatic prompt optimization method based on sample feedback

Through the automatic prompt optimization method of large language model based on sample feedback, prompt optimization, streamlining and local search modules are built, which solves the problem of high computing power and reliance on human tuning in traditional methods, and realizes automated and efficient prompt optimization, improving the performance of the model on target tasks.

CN120067241AActive Publication Date: 2025-05-30HARBIN INST OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510107858.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

Traditional fine-tuning methods consume a lot of computing power to train large language models, which may lead to the forgetting of pre-training knowledge. The traditional prompt tuning methods rely on manpower and experience, and cannot achieve rapid and automated optimization.

Method used

The automatic prompt optimization method of large language model based on sample feedback is adopted. By constructing prompt optimization, streamlining and local search modules, input prompts are automatically optimized to reduce human intervention.

Benefits of technology

The ability to automatically optimize prompts for large language models is realized, which reduces manpower investment and improves task efficiency. The optimized prompts are closer to the global optimal solution of the target task, and improves the performance of the model on the target task.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067241A_ABST
    Figure CN120067241A_ABST
Patent Text Reader

Abstract

The invention discloses a large language model automatic prompt optimization method based on sample feedback, and belongs to the technical field of large language model prompt optimization. The problem that in the prior art, a traditional prompt optimization method is difficult to achieve automatic adjustment and optimization is solved. The method comprises the following steps: constructing a prompt optimization module based on a large language model, inputting preprocessed input data, and performing prompt optimization based on sample feedback on the preprocessed input data to obtain a modified prompt; constructing a prompt simplification module based on a large language model, simplifying and rewriting a super-long prompt in the modified prompt to obtain an updated prompt, and transmitting the updated prompt to a prompt optimization module for iteration to obtain an optimized prompt; and constructing a local search module based on a large language model, and carrying out local search and optimization on the optimized prompt to obtain an optimal prompt. According to the method, the performance of the large language model for prompt optimization is effectively improved, and the method can be applied to automatic prompt optimization by adopting the large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for optimizing automatic prompts of large language models, and particularly to a method for optimizing automatic prompts of large language models based on example feedback, belonging to the technical field of large language model prompt optimization. Background Art

[0002] Large Language Models (LLMs) are one of the current research hotspots in the field of artificial intelligence. With the development of deep neural network structures and computing power, neural networks have been widely applied in various industries. Due to the serial modeling scheme based on transformers and the good parallel computing efficiency brought by the model structure, large language models have emerged. The large number of parameters and training corpora of large language models endow them with good generalization ability, and they can also achieve good results when facing complex downstream tasks that have not been trained.

[0003] Due to the large number of parameters of large language models, training large language models using traditional fine-tuning methods will cause extremely high computing power consumption and may also bring the problem of pre-trained knowledge forgetting, reducing the comprehensive performance of the model; traditional prompt tuning methods rely heavily on the input of manpower and experience, and cannot achieve a fast and automated optimization process when facing complex problems.

[0004] In summary, there is a need for a method for automatically optimizing prompts of large language models based on example feedback. Summary of the Invention

[0005] A brief overview of the present invention is given below to provide a basic understanding of certain aspects of the present invention. It should be understood that this overview is not an exhaustive overview of the present invention. It is not intended to identify the key or important parts of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is only to present certain concepts in a simplified form as a prelude to the more detailed description to follow.

[0006] In view of this, to solve the problem that traditional prompt optimization methods in the prior art are difficult to achieve automatic tuning, the present invention provides a method for automatically optimizing prompts of large language models based on example feedback.

[0007] The technical solution is as follows: A method for automatically optimizing prompts of large language models based on example feedback, comprising the following steps:

[0008] S1. Construct a prompt optimization module based on the large language model, input the preprocessed input data, perform prompt optimization based on example feedback on it, and obtain the modified prompt;

[0009] S2. Build a prompt refinement module based on the large language model to refine and rewrite the overly long prompts in the modified prompts, obtain the updated prompts, and transmit them to the prompt optimization module for iteration to obtain the optimized prompts;

[0010] S3. Build a local search module based on the large language model to perform local search and optimization on the optimized prompts to obtain the optimal prompts.

[0011] Furthermore, in S1, it specifically includes the following steps:

[0012] S11. Preprocess the input data and input the preprocessed input data into the prompt optimization module based on the large language model;

[0013] S12. Perform forward propagation through the prompt optimization module to obtain the prediction error set S Gi ;

[0014] S13. According to the prediction error set S Gi , calculate the gradient through the prompt optimization module to obtain the gradient set G i ;

[0015] S14. According to the gradient set G i , perform backpropagation through the prompt optimization module and output the modified prompts;

[0016] In S11, the input data includes the initial prompt and the training data. The initial prompt is prompt_0, which is expressed as a piece of text and is used to clearly describe the target task. The training data is supervised data;

[0017] In S12, based on the large language model, use the prompt i to perform inference on the data S in batch i i , where i = 1, 2,..., n, to obtain the model inference result ModelOut s , based on the supervision label of the data S in batch i i and the preset threshold, use the large language model to evaluate the inference output, and incorporate the data below the preset threshold or with inference errors into the prediction error set S Gi ;

[0018] The prediction error set S Gi is expressed as:

[0019] S Gi ={answer s ≠ModelOut s |s∈S i}

[0020] where s is the input of the model, answer sThe label corresponding to the model input;

[0021] In S13, use a large language model to evaluate the prediction error set S Gi Analyze the input features of the training data and the gap between the model inference result ModelOut s and the supervision label, use the large language model to automatically analyze the reasons for the inference errors of the current prompt _i, and output in text with a fixed format, and calculate the gradient set G i ;

[0022] The gradient set G i is expressed as:

[0023] G i ={ModelOut s,ModelOuts,answers |s∈S Gi}

[0024] In S14, use the large language model to expand the current prompt _i, add the features extracted from the gradient set G i to the prompt i to obtain a new prompt i+1, and based on the new prompt i+1 and batch i+1, perform the next round of loop until the preset maximum number of iteration steps is reached, and output the current new prompt _i+1, that is, the modified prompt.

[0025] Furthermore, in S2, it specifically includes the following steps:

[0026] S21. Through the prompt refinement module based on the large language model, cluster and merge the current new prompt _i+1 to obtain the classification result after merging by class;

[0027] S22. According to the classification result after merging by class and the original prompt, recombine to obtain the updated prompt for the iteration of the prompt optimization module to obtain the optimized prompt;

[0028] In S21, for the rule description text in the current new prompt _i+1, classify according to semantics and output the classification result. For each class of classification results, based on the large language model, merge and delete the redundant parts in the same class of classification results to obtain the classification result after merging by class.

[0029] Furthermore, in S3, it specifically includes the following steps:

[0030] S31. Select two prompts from the optimized prompts and randomly rewrite them through the local search module to obtain n rewritten prompts;

[0031] S32. According to the obtained n rewritten prompts, evaluate them through the local search module to obtain the optimal prompt on the validation set;

[0032] In S31, for the k rules of the rule description text in the two selected prompts, which are the classification results, randomly select some rules, use the large language model to randomly rewrite them, and repeat n times to obtain n rewritten prompts, regarded as performing n local searches around the current optimal prompt;

[0033] In S32, according to the obtained n rewritten prompts, based on the predetermined rules and the large language model, evaluate on the evaluation set, select the two best-performing prompts to repeat the iteration of the local search module until the preset maximum number of iteration steps is reached, and obtain the optimal prompt on the validation set.

[0034] The beneficial effects of the present invention are as follows: The present invention proposes a self-feedback mechanism for building a large language model based on a small number of supervised examples. Based on a small number of manually annotated examples, it analogizes the interaction logic of building a large language model by the method of training a model based on backpropagation, and uses the manually annotated examples to automatically optimize the input prompts; The present invention can enable the large language model to automatically optimize and update the prompts, liberate the manpower from the cumbersome prompt modification tasks, improve the task efficiency, and at the same time, compared with the manually optimized prompts, the results of the present invention can be closer to the global optimal solution of the prompts concerned by the target task and achieve better performance on the target task. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The drawings described herein are used to provide a further understanding of the present invention and form a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0036] Figure 1 It is a schematic flow chart of an automatic prompt optimization method for a large language model based on example feedback;

[0037] Figure 2 It is a schematic flow chart of an embodiment of an automatic prompt optimization method for a large language model based on example feedback;

[0038] Figure 3 It is a schematic flow chart of an embodiment of simplifying and rewriting using a prompt simplification module;

[0039] Figure 4 It is a schematic flow chart of an embodiment of searching and tuning using a local search module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] To make the technical solutions and advantages in the embodiments of the present invention clearer and more understandable, the following further elaborates on the exemplary embodiments of the present invention with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than an exhaustive list of all embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0041] Reference Figures 1-4 Specifically describe this embodiment, an automatic prompt optimization method for large language models based on example feedback, which specifically includes the following steps:

[0042] S1. Construct a prompt optimization module based on the large language model, input the preprocessed input data, perform prompt optimization based on example feedback on it, and obtain the modified prompt;

[0043] S2. Construct a prompt refinement module based on the large language model, refine and rewrite the overly long prompts in the modified prompts, obtain the updated prompts, transmit them to the prompt optimization module for iteration, and obtain the optimized prompts;

[0044] S3. Construct a local search module based on the large language model, perform local search and optimization on the optimized prompts, and obtain the optimal prompts.

[0045] Furthermore, in S1, it specifically includes the following steps:

[0046] S11. Preprocess the input data, and input the preprocessed input data into the prompt optimization module based on the large language model;

[0047] S12. Perform forward propagation through the prompt optimization module to obtain the prediction error set S Gi ;

[0048] S13. According to the prediction error set S Gi , perform gradient calculation through the prompt optimization module to obtain the gradient set G i ;

[0049] S14. According to the gradient set G i , perform backpropagation through the prompt optimization module and output the modified prompt;

[0050] In S11, the input data includes the initial prompt and the training data. The initial prompt is prompt_0, which is presented as a brief piece of text to clearly describe the target task, such as "You are a Chinese-English translator and need to translate the input English into Chinese and output". The training data is generally on the order of hundreds of supervised data. For example, in the translation task, the training data is presented as hundreds of parallel corpora;

[0051] In S12, based on the large language model, use prompt i to perform inference on the data S in batch i i where i = 1, 2, ..., n, to obtain the model inference result ModelOut s , and based on the supervision label of the data S in batch i i and a preset threshold (if not present, no need to consider), use the large language model to evaluate the inference output, and incorporate the data below the preset threshold or with incorrect inferences into the prediction error set S Gi ;

[0052] The prediction error set S Gi is expressed as:

[0053] S Gi = {answer s ≠ModelOut s | s ∈ S i}

[0054] where s is the input of the model and answer s is the label corresponding to the model input;

[0055] In S13, use the large language model to evaluate the prediction error set S Gi , analyze the input features of the training data, as well as the gap between the model inference result ModelOut s and the supervision label, use the large language model to automatically analyze the reason for the incorrect inference of the current prompt _i, and output it in a fixed format of text, and calculate the gradient set G i ;

[0056] The gradient set G i is expressed as:

[0057] G i = {ModelOut s,ModelOuts,answers | s ∈ S Gi}

[0058] In S14, use the large language model to expand the current prompt _i, add the features extracted from the gradient set G i to the prompt i to obtain a new prompt i + 1, and based on the new prompt i + 1 and batch i + 1, perform the next round of loop until the preset maximum number of iteration steps is reached, and output the current new prompt _i + 1, that is, the modified prompt.

[0059] Specifically, referring to Figure 3 , the main purpose of step S1 is to optimize the input prompt based on a small amount of supervised examples, and achieve automatic prompt optimization of the large language model based on a small-scale training data;

[0060] The prompt optimization module, as the core module, is similar to training a deep neural network based on the backpropagation algorithm. The prompt optimization module is divided into three algorithm steps: (1) Forward propagation: Reason about a mini-batch of data based on the current prompt; (2) Calculate gradients: Based on the mini-batch of data with inference errors, use the large language model to generate the features of error examples in text form; (3) Backpropagation: Based on the features generated by calculating gradients, modify the current prompt and perform the next round of iteration.

[0061] Further, in S2, it specifically includes the following steps:

[0062] S21. Through the prompt refinement module based on the large language model, cluster and merge the current new prompt_i+1 to obtain the classification result after merging by class;

[0063] S22. According to the classification result after merging by class and the original prompt, recombine to obtain the updated prompt for the iteration of the prompt optimization module to obtain the optimized prompt;

[0064] In S21, for the rule description text in the current new prompt_i+1, classify it according to semantics and output the classification result. For each class of classification results, based on the large language model, merge and delete the redundant parts in the same-class classification results to obtain the classification result after merging by class, improving the information density of the current prompt.

[0065] Specifically, referring to Figure 4 , the main purpose of step S2 is to optimize the overly long prompts that appear during the training process of step S1, automatically perform refinement and rewriting, improving the information density of the optimized prompt and the inference efficiency of the large language model. For the rule description text in the prompt, cluster and merge it based on the thinking chain of the large language model;

[0066] In the prompt optimization module, since each optimization step requires expanding the current prompt, in the case of a large amount of training data, there may be a phenomenon of overly long prompts. Overly long prompts will cause many additional problems, such as increased computational overhead and loss of context. When the prompt is too long, the generation of the large language model may have problems such as inconsistent content or lack of context consistency. The large language model may lose focus in long texts and generate content that is inconsistent with the previous text or irrelevant to the question, reducing the quality of the output.

[0067] Further, in S3, it specifically includes the following steps:

[0068] S31. Select two prompts from the optimized prompts and randomly rewrite them through the local search module to obtain n rewritten prompts;

[0069] S32. Based on the obtained n rewritten prompts, evaluate them through the local search module to obtain the optimal prompt on the validation set;

[0070] In S31, when the local search module starts iterating, the current number of prompts is 1. For the k rules of the rule description text in the two selected prompts, that is, the classification results, randomly select some rules, use the large language model to perform random rewriting, and repeat n times to obtain n rewritten prompts, regarded as performing n local searches around the current optimal prompt;

[0071] In S32, based on the obtained n rewritten prompts, and based on predetermined rules and the large language model, evaluate them on the evaluation set, select the two best-performing prompts to repeat the iteration of the local search module until the preset maximum number of iteration steps is reached, and obtain the optimal prompt on the validation set.

[0072] Specifically, refer to Figure 2 , the two currently selected prompts are prompt_fst, that is, prompt_1, and prompt_sed, that is, prompt_2. The main purpose of step S3 is to prevent the overfitting phenomenon that may occur in the prompt optimization algorithm, improve the generalization performance of the prompt. After the training of the prompt optimization module is completed, search near the currently obtained optimal prompt to obtain the optimal prompt on the validation set;

[0073] In the automatic prompt optimization technology, the overfitting phenomenon means that the generated prompts perform well on the training set, but perform poorly on unseen data, that is, the optimization results of the prompts are overfitted to the training data involved in the optimization process and lack generalization ability. For the prompt automatic optimization technology with less training data and the optimization process directly depending on example features, the overfitting phenomenon will be more serious. The purpose of the local search module is to alleviate this overfitting phenomenon. The specific steps are divided into random rewriting of prompts and prompt evaluation;

[0074] In this embodiment, first, about a hundred pieces of supervised data and the initial prompts to be optimized need to be collected. Then, based on the prompt optimization module and the prompt refinement module, complete the optimization of the input prompts. Finally, based on the local search module, improve the generalization ability of the prompts optimized based on examples;

[0075] In the commodity tagging task in the e-commerce field, if the method of manually optimizing prompts is used, it will not only cause a large amount of human loss, but also the manually optimized prompts cannot fully cover the many rules required by the commodity compliance task. Using the method proposed by the present invention for prompt optimization, the prompts automatically optimized by the large language model can better utilize the capabilities of the large language model and achieve better results on the target task.

[0076] Although the present invention has been described in terms of a limited number of embodiments, those skilled in the art, having the benefit of the foregoing description, will appreciate that other embodiments can be contemplated within the scope of the invention as thus described. Further, it should be noted that the language used in this specification has been principally selected for readability and instructional purposes and not to limit or circumscribe the inventive subject matter. Accordingly, many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the appended claims. For the scope of the present invention, the disclosure herein is illustrative and not restrictive, the scope of the invention being defined by the appended claims.

Claims

1. A large language model automatic prompt optimization method based on sample feedback, characterized in that: The following steps are involved: S1. Build a prompt optimization module based on a large language model, input the preprocessed input data, optimize the prompt based on sample feedback, and obtain the modified prompt; S2. construct a prompt simplification module based on a large language model, simplify and rewrite the super-long prompts in the modified prompts, obtain updated prompts, and transmit them to the prompt optimization module for iteration to obtain optimized prompts; S3. Build a local search module based on the large language model, perform local search and tuning on the optimized prompts, and obtain the optimal prompts.

2. The large language model automatic prompt optimization method based on sample feedback according to claim 1, characterized in that: The S1 specifically includes the following steps: S11. Preprocess the input data, and input the preprocessed input data into a prompt optimization module based on a large language model; S12. Forward propagation is performed through the prompt optimization module to obtain a set of prediction errors S13. Based on the prediction error set The gradient set G is obtained by performing gradient calculation through the prompt optimization module. i ; S14. According to the gradient set G i , back-propagates through the prompt optimization module and outputs the modified prompt; In S11, the input data includes an initial prompt and training data, the initial prompt is prompt_0, which is a sentence of text used to clearly describe the target task, and the training data is supervised data; In S12, based on the large language model, the hint i is used to analyze the data S in batch i. i Perform inference, i = 1, 2, ..., n, and obtain the model inference result ModelOut s , based on the data S in batch i i The supervised labels and pre-set thresholds are used to evaluate the inference output using a large language model, and the data below the pre-set threshold or with inference errors are incorporated into the prediction error set Prediction Error Collection It is expressed as: Among them, s is the input of the model, answer s Enter the corresponding label for the model; In S13, a large language model is used to predict the error set. Evaluate and analyze the input features of the training data and the model inference results ModelOut s The gap between the supervised label and the current prompt_i is automatically analyzed using a large language model to cause the reasoning error, and the output is in a fixed format text to calculate the gradient set G i ; Gradient set G i It is expressed as: In S14, the current prompt _i is expanded using the large language model, and the gradient set G i The features extracted are added to the prompt i to obtain the new prompt i+1. Based on the new prompt i+1 and batch i+1, the next cycle is performed until the preset maximum number of iterations is reached, and the current new prompt _i+1, i.e. the modified prompt, is output.

3. The large language model automatic prompt optimization method based on sample feedback according to claim 2 is characterized in that: The S2 specifically includes the following steps: S21. Cluster and merge the current new prompt _i+1 through the prompt simplification module based on the large language model to obtain the classification result after merging by category; S22. According to the classification results after merging by category and the original prompt, an updated prompt is obtained by recombining, and used for iterating the prompt optimization module to obtain an optimized prompt; In S21, the rule descriptive text in the current new prompt_i+1 is classified according to semantics and the classification results are output. For each type of classification result, the redundant parts in the same type of classification results are merged and deleted based on the large language model to obtain the classification results after being merged by category.

4. The large language model automatic prompt optimization method based on sample feedback according to claim 3 is characterized in that: The S3 specifically includes the following steps: S31. Select two prompts from the optimized prompts, and randomly rewrite them through the local search module to obtain n rewritten prompts; S32. Based on the obtained n rewriting hints, evaluate them through the local search module to obtain the best hint on the validation set; In S31, for the k rules of the rule descriptive texts in the two selected prompts, i.e., the classification results, some rules are randomly extracted, randomly rewritten using the large language model, and repeated n times to obtain n rewriting prompts, which is regarded as performing n local searches around the current optimal prompt; In S32, according to the obtained n prompts, based on predetermined rules and a large language model, an evaluation is performed on the evaluation set, and the two prompts with the best performance are selected to repeat the iteration of the local search module until a preset maximum number of iterations is reached, thereby obtaining the optimal prompt on the verification set.

Citation Information

Patent Citations

  • Automatic optimization method and device for cue word, equipment and storage medium

    CN117520507A

  • Model optimization method and device, equipment and storage medium

    CN117829244A

  • Medical record information extraction method for large model prompt design based on particle swarm optimization

    CN118966228A

  • Semantic relationship prediction method based on multi-section prompt optimization of large language model

    CN119294399A

  • Revising large language model prompts

    US20240362422A1