Automatic thinking chain prompt generation method based on black box optimization and vulnerability quantification
By filtering the subset of high-difficulty problems, generating and optimizing the reasoning chain, and performing diversified style processing, the problem of inefficient thinking chain prompt technology is solved, and efficient training and accurate answers of large language models in complex reasoning tasks are achieved.
Patent Information
- Application Number
- CN202510806024.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-17
AI Technical Summary
Existing thinking chain cues techniques are inefficient and insufficiently scalable in complex inference tasks, requiring manual careful design of step-by-step reasoning examples, resulting in high training costs and time, and there is cues vulnerability.
The difficulty judgment module filters out the subset of high-difficulty problems, generates labeled data sets, uses a large language model to generate multiple inference chains, uses the variance reduction strategy gradient estimator to optimize the inference chain selection strategy, and uses the novel format prompt module to diversify styles to generate automatic thinking chain prompts that resist vulnerability.
It reduces the training cost and time of large language models, improves the accuracy and stability of answering inference questions, reduces the fragility of prompts, and improves the model's reasoning ability in different fields.
Smart Images

Figure CN120336491A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of black-box optimization, and particularly relates to an automatic chain of thought prompt generation method based on black-box optimization and vulnerability quantification. Background Art
[0002] In recent years, with the rapid development of large language models, a variety of innovative prompt engineering techniques have emerged one after another. Among them, the zero-shot prompting technique can guide the model to complete tasks by only providing task descriptions without relying on any training data. In contrast, the few-shot prompting technique goes a step further by incorporating a small number of carefully designed input-output examples into the prompt, significantly improving the model's ability to understand task intentions and boundaries. These two breakthrough techniques have greatly reduced the application threshold of large language models, enabling them to quickly adapt to various new task scenarios without time-consuming fine-tuning training. In the solution of complex reasoning tasks, currently two advanced prompt engineering techniques are mainly adopted: chain of thought prompting and automatic chain of thought prompting. The chain of thought prompting technique significantly improves the model's performance in logical reasoning tasks by guiding the large language model to show a complete reasoning process. However, this method requires manual careful design of step-by-step reasoning examples, with limitations of low efficiency and insufficient scalability. Summary of the Invention
[0003] Object of the Invention: The object of the present invention is to provide an automatic chain of thought prompt generation method based on black-box optimization and vulnerability quantification, which screens training samples to reduce the training cost and time of the large language model for automatically generating prompts, generates detailed chains of thought prompts for reasoning problems in different fields, and sets different prompt styles to reduce prompt vulnerability, helping the large language model understand the prompts and questions and give more reasonable and appropriate answers.
[0004] Technical Solution: An automatic chain of thought prompt generation method based on black-box optimization and vulnerability quantification described in the present invention includes the following steps: (1) Obtain an unlabeled training data set; screen a subset of high-difficulty problems through a difficulty judgment module to generate a labeled data set; (2) Use the large language model to generate multiple reasoning chains for each problem in the labeled data set, and retain the reasoning chains that are consistent with the correct answers to form an example library; (3) Adopt a variance reduction strategy gradient estimator to optimize the reasoning chain selection strategy to generate a normal style prompt; (4) Perform style diversification processing on the normal style prompt through a novel format prompt module to generate an anti-fragile automatic chain of thought prompt.
[0005] Further, the operation of the difficulty judgment module includes: calculating the entropy value of the prediction probability distribution of the large language model generating k answers Calculate the entropy value: ; Among them, represents the parametric model generates a specific answer when given a problem of the prediction probability; represents the parameters of the language model; tunable hyperparameters and are used to control the probability scaling ratio; select the top n questions with the largest entropy values as high-difficulty questions.
[0006] Furthermore, when the number of high-difficulty questions exceeds G, randomly select M questions; the construction of the labeled dataset is achieved through automatic labeling by the large language model; among them, G is the set threshold for the number of high-difficulty questions.
[0007] Furthermore, the generation and pruning of the reasoning chain include: generating l reasoning chains for each question; comparing the generated answer with the labeled answer, and only retaining the correctly derived reasoning chains to pair the questions with the correct reasoning chains to form an initial example set.
[0008] Furthermore, the optimization process of the variance reduction policy gradient estimator includes: defining a latent variable that follows a categorical distribution ; among them, is the probability vector of N candidate examples; update through the gradient estimation formula : ; where T represents the complete sample example set, is the i-th example, is the probability that the i-th example is selected; the set of examples independently sampled from ; ; is the gradient estimation used to update ; is the gradient of the probability vector ; represents the difference between the loss of the current example set and the average loss of all example sets; Adopt projected stochastic gradient descent to update the example distribution: ; where, is the learning rate, I is the size of the example set, is the projection calculation.
[0009] Furthermore, during the optimization process: the loss function Calculation of cross-entropy between the output of the black-box model and the true label; the number of context examples is dynamically adjusted to 3-6.
[0010] Furthermore, the operations of the novel format prompt module include: designing unique style templates for each few-shot example, and the template types include: paraphrasing non-semantic markup formats, adjusting visual structures, optimizing hierarchical reconstructions, and nesting ordinary style prompts into style templates to generate final prompts.
[0011] An electronic device according to the present invention includes a memory and a processor, the memory stores a computer program, and when the processor executes the program, it implements the method described in any one of the above.
[0012] A computer-readable storage medium according to the present invention stores a computer program, and when the program is executed by a processor, it implements the steps of the method described in any one of the above.
[0013] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: (1) Using difficulty as an indicator to select sample problems that are helpful for improving the reasoning ability of large language models, reducing the training cost and time of the model. (2) Using the variance reduction strategy gradient estimator and projection calculation to let the large language model autonomously generate correct thought chain prompts, avoiding the appearance of toxic and harmful prompt content. (3) Building a new style prompt template, and the large language model uses the novel prompts generated by the template, which well reduces the vulnerability of the prompts themselves. (4) The novel prompts improve the Q&A effect of the large language model in the field of reasoning problems, and the accuracy and stability of the answers are significantly improved. Description of the Drawings
[0014] Figure 1 is a flowchart of the present invention; Figure 2 is a reasoning process diagram of difficult problem samples of the present invention; Figure 3 is a performance diagram of random prompts and novel prompts of the present invention on 6 reasoning datasets; Figure 4 is a comparison diagram of the effects of novel prompts, automatic thought chain prompts, and self-improvement under different pool sizes of the present invention; Figure 5 is a comparison diagram of the performance differences between ordinary style prompts and novel prompts of the present invention. Detailed Embodiments
[0015] The technical solutions of the present invention will be further described below with reference to the drawings.
[0016] As Figure 1 shown, an embodiment of the present invention provides an automatic thought chain prompt generation method based on black-box optimization and vulnerability quantification, including the following steps: Given an unlabeled training dataset (where each q is a question without reasoning steps), construct a new dataset , which contains n labeled samples selected from . It will be described how to screen the most challenging n questions based on question difficulty and annotate them. After obtaining the dataset E, use a large language model to generate l reasoning chains for each sample and retain the reasoning chains that are consistent with the correct answers. By pairing the questions with the reasoning chains to form new prompts, these prompts can be used as efficient demonstration samples. Adopt the "novel format" strategy, incorporate diverse formats into these automatic thinking chain prompts, and quantify the style sensitivity difference by calculating the performance fluctuation degree (i.e., the accuracy difference between the optimal and worst prompt formats). Among them, (1) The judgment of difficulty is as follows: When screening a subset of questions from a large-scale dataset, the difficulty score of the large language model needs to be used as a measure to evaluate the difficulty of each question. Based on the thinking chain prompt framework, require the large language model to generate k answers for each question, and then calculate the importance index by predicting the entropy value of the answers. The specific calculation formula is as follows: ; where, represents the predicted probability of the parametric model generating a specific answer when given the question ; represents the parameters of the language model; the adjustable hyperparameters and are used to control the probability scaling ratio; select the top n questions with the largest entropy value as high-difficulty questions. By increasing the weight of the head prediction or weakening the influence of low-frequency predictions, adjust the logarithmic contribution degree of the prediction results. This design can balance high-confidence and low-confidence predictions, thus achieving a more accurate difficulty grading for the selection of questions for different complexity reasoning tasks.
[0017] (2) The selection and annotation are as follows: After evaluating the difficulty of each question, generate a ranking according to the difficulty score and screen out the top n questions with the highest difficulty for annotation. If there are a large number of high-difficulty questions (the number exceeds G, G = 500), then randomly select m questions from this subset. Subsequently, the correct answers of these selected questions will be annotated by the large language model to form a new dataset . E contains the finally selected sample cases. This dataset will be used by the large language model to generate thinking chain prompts.
[0018] (3)Generate and prune auto-generated chain-of-thought prompts as follows: Based on the sample set E, use a large language model to generate l chains of reasoning for each question q. Considering cost and efficiency, l = 3 is set in the experiment. By comparing the generated answers with the standard answers, filter out the incorrect chains of reasoning and only retain the chains that can derive the correct answer. Finally, n high-quality chains of reasoning are obtained. These chains of reasoning and the sample set E together constitute an example library. As Figure 2 shown, the reasoning process of the large language model on difficult questions and the accuracy of the obtained answers.
[0019] (4)Chain-of-reasoning optimization; specifically as follows: After constructing a high-quality example set, the next step is to apply it to the large language model, but the following limitations need to be considered: (41) The input context length of the LLM is limited and it is impossible to use all examples at once; (42) The presentation order of examples may affect the model performance; (43) Different downstream tasks may require different examples to achieve the best results. A reasonable solution is to automatically select 3 - 6 examples according to the specific context and task. This process can be regarded as optimizing a supervised model with latent variables: For each chain-of-thought index i, initialize a latent variable that follows a categorical distribution , and this random variable samples on N candidate example indices according to the probability distribution . Among them, the probability vector belongs to the set C, represents the probability of selecting each candidate example, and moreover must satisfy the constraints defined by C. That is . Since the distributions are independent, the joint probability of the overall input examples is , where T represents the complete sample example set, is the i-th example, represents that the i-th example is the -th selected from N candidate examples, is the probability that the i-th example t i is selected. The corresponding loss function is defined as , G represents the generative model, S is the current question, represents the composition of the prompt, and y is the true label of the question. Since it is impossible to obtain gradient information using a black box model, the prompt cannot be directly updated by gradient backpropagation ( is the gradient of the probability vector ), and a variance reduction strategy gradient estimator is adopted to optimize the loss function through forward propagation: ; Among them, represents the expectation of the loss function with respect to the example set T. Then the gradient estimation formula of ; a set of examples independently sampled from P(T) , is the gradient estimation, used to update .
[0020] represents the difference between the loss of the current example set and the average loss of all example sets. Subsequently, projected stochastic gradient descent is used to update the example distribution: ; where is the learning rate, I is the scale of the example set, is the projection calculation. For a black-box model (such as GPT-3.5) for which the gradient cannot be obtained, the user's question is used as the input, and by iteratively updating , the model automatically selects the optimal example set, which consists of the reasoning chain and the question, achieving the optimization effect of the reasoning chain. The example set can be used as a normal style prompt.
[0021] (5) Prompt style variation. In the present invention, the thought chain prompts generated by the large language model still have prompt vulnerability caused by style differences, which is similar to the phenomenon observed in the field of computer vision - minor changes in image style (such as color or background modification) may affect the prediction accuracy of the model. Therefore, the present invention proposes a "novel format" prompt strategy. The characteristic of the strategy is to adopt different style formats for each small sample example in the existing thought chain prompts, where the style template contains the initial sample example and the prompt words for rewriting the content of the sample example. For example, rewrite the example by only changing the format of non-semantic markers, reorganize the visual structure of the example without changing the semantic expression, and optimize the layering of the example (such as title grading, indentation adjustment) while maintaining the integrity of the information. To further enhance the model's understanding, the default model nests the prompt content of the normal style prompt on the style template to generate novel prompts.
[0022] The automatic thought chain prompts designed by the "novel format" can be applied to the reasoning problem tasks of the large language model and compared with the existing normal prompt techniques. For example, automatic thought chain prompts, self-consistency, self-improvement, thought growth, self-check. In the present invention, the novel prompts will be compared with the existing prompt techniques. The effects of the prompts will be tested on 6 different reasoning tasks, and then the generated answers will be compared with the true answers to obtain the answer accuracy of the large language model.
[0023] To ensure the richness and effectiveness of the experiment, experiments were conducted on 6 reasoning tasks, including 4 arithmetic reasoning problem datasets: GSM8K, ASDiv, SVAMP, and AQuA. In addition, a commonsense reasoning dataset: StrategyQA and a symbolic reasoning dataset: Letter were used. Table 1 gives the detailed statistics of these datasets, and the evaluation metric for each task is the accuracy of the model's answers.
[0024] In the experiment, the present invention was compared with four existing prompting techniques: Automatic Chain of Thought Prompting, Self-Consistency, Self-Improvement, and Thought Growth. The experiment mainly used gpt-3.5-turbo, and the API calls were made through OpenAI's services. The Llama2-70b-chat model was also used to ensure the adaptability of the method to the model.
[0025] The experimental results are shown in Table 1. The present invention is significantly better than other prompting techniques. In the 6 benchmark datasets, the large language models applying the novel prompt always show excellent performance. On the gpt-3.5-turbo model, the average value of the novel prompt is 3.6% and 5.3% higher than Self-Consistency and Automatic Chain of Thought Prompting respectively. Similarly, on the Llama2-70b-chat model, the average value of the novel prompt is 0.8% and 1.9% higher than Self-Improvement and Thought Growth respectively. The results of arithmetic reasoning, commonsense reasoning, and symbolic reasoning will be discussed separately. It is worth noting that the novel prompt has improved the answer accuracy on both gpt-3.5 and Llama2-70b, and the improvement on Llama2-70b is greater, indicating that the method of the present invention is applicable to large language models with different parameter scales.
[0026] Table 1 Performance comparison between the novel prompt and other prompting methods on 6 reasoning tasks ; In the four arithmetic reasoning tasks, the novel prompt showed a significant improvement compared to the Automatic Chain of Thought Prompting, with an average gain of 5.325%. Compared with the strong baseline method of Self-Improvement, the answer accuracy still increased by 1.875%. In the test of the Llama2-70b-chat model, the novel prompt also achieved the best average result. Compared with Self-Consistency, it increased by 2.625% on average in the four tasks. Among them, the novel prompt increased by 6.6% compared to the Automatic Chain of Thought Prompting in the ASDIV reasoning task. Since ASDIV relies more on fine-grained semantic representations and explicit standard reasoning paths, the novel prompt generated appropriate reasoning chains using the language model in the multi-round reasoning path generation, which helps to be used as sample examples for prompting.
[0027] In common sense and symbolic reasoning tasks, novel prompting has achieved significant improvements compared to self-consistency and auto-thinking chain prompting. On the gpt-3.5-turbo-1106 model, novel prompting improved by 2.0% compared to thought growth in the common sense reasoning task STRATEGY and by 1.0% compared to self-improvement in the symbolic reasoning task LETTER. When tested using Llama2-70b-chat, the improvements of novel prompting compared to auto-thinking chain prompting were 2.7% and 6.3% on STRATEGY and LETTER respectively. It also outperformed other strong baseline methods in the performance of these two types of reasoning tasks. This fully demonstrates that the method of the present invention is effective in different types of reasoning tasks.
[0028] Next, a series of ablation experiments were conducted to examine the impact of each part in the method design proposed by the present invention. First, the contribution of the proposed difficult question selection strategy was explored. Then, the influence of the answer pool size was studied. Finally, how the "novel format" prompting strategy enhances the understanding ability of large language models was analyzed.
[0029] In Table 2, the accuracies of language models of random prompting, self-consistency, self-improvement, and novel prompting on three reasoning tasks were compared. It was observed that the performance of novel prompting was significantly better than that of random prompting, while the effect of random prompting was only slightly higher than that of self-consistency and lower than that of self-improvement. As Figure 3 shown, the answer accuracies of random prompting in six reasoning tasks were all inferior to those of novel prompting. This indicates that the performance improvement does not come from the advantages in the annotation process, but from the strategy of selecting challenging questions. This also highlights the need for further exploration of automatic prompting engineering in terms of difficulty, diversity, and style.
[0030] Table 2 Accuracies of Random Prompting, Self-Consistency, Self-Improvement, and Novel Prompting in 3 Reasoning Tasks ; In the process of judging the difficulty of sample questions in the present invention, k answers are generated for each input question to establish a prediction pool. It is assumed that the size of k will affect the difficulty judgment and performance of the large language model in downstream tasks. To verify this, a series of experiments were conducted to test different pool sizes. The experiments were jointly carried out on six reasoning tasks. As Figure 4As shown, the accuracy of different numbers of predicted answers is plotted (k = 1, 5, 10, 20). It can be clearly seen from the results that when the pool size is only 1, the performance of the novel prompt is inferior to that of the auto-thinking-chain prompt. This indicates that when the pool size is small, high-quality reasoning steps cannot be selected. However, when the pool size increases to 10 or more, the performance of the selected prompt can exceed that of the self-improving method. This trend shows that as the pool size grows, more complex and diverse reasoning steps can be selected, leading to continuous improvement in the performance of large language models. The results also converge at k = 20, and further increase in the pool size will result in diminishing returns. Intuitively, smaller k values introduce noise in the reasoning selection process, while larger k values make the selection for difficult problems more stable. Due to computational limitations, the maximum pool size was limited to 20 in the experiment. This shows that while a larger pool size helps with more robust difficulty judgment, there is a practical upper limit beyond which the performance will plateau.
[0031] During the experiment, the self-improving method was used as the baseline, where the large language model utilized the auto-thinking-chain prompt optimized by the reasoning chain, and at this time the prompt had not adopted the "novel format". This prompt is called the ordinary-style prompt. The performance of the large language model when applying the ordinary-style prompt and the novel prompt was compared. The main focus was on whether the novel prompt could narrow the performance difference of the large language model when the prompt style changed. The difference is a measure quantifying the vulnerability of style-induced prompts, and the performance difference mainly refers to the difference between the prompt with the highest correct rate and the prompt with the lowest correct rate. In Table 3, the results of the best and worst prompts in the ordinary-style prompt and the novel prompt are reported, and it is observed that the novel prompt not only reduces the difference but also improves the minimum and maximum accuracies. As Figure 5 shown, in most datasets, the performance of the novel prompt is better than that of the ordinary-style prompt, while in AQUA and LETTER, the ordinary-style prompt performs better.
[0032] Table 3 Results of the best and worst prompts for the ordinary-style prompt and the novel prompt ; In the experiment, the training data (denoted as ) and the test data (denoted as ). In different types of reasoning tasks, the number of difficult questions selected is as follows: 6 difficult questions are selected from each of GSM8K, ASDiv, SVAMP, and StrategyQA, and 4 difficult questions are selected from each of AQuA and Last Letter. In the reasoning stage, the temperature parameter is set to 0.68, and each question is run through 10 inferences. The maximum number of tokens for the input-output is set to 256. The default version of GPT-3.5-turbo used is gpt-3.5-turbo-1106. To ensure the generalization of the method, Llama2-70b-chat will also be tested.
[0033] In the present invention, problem samples manually annotated are first used to ensure the reliability of the results in the difficulty judgment stage. These problem samples usually come from the source dataset itself. Even in the absence of manually annotated problem samples, the experiment can still proceed. During the experiment, the size of the candidate problem pool is limited to 500 - if the original training dataset exceeds 500, 500 questions are randomly selected; if it is less than 500, all data is used. Through testing and verification with data pools of different sizes, 500 questions can achieve a better balance in performance. Generally speaking, a larger data pool will bring a higher performance gain. When the number of answer generations k = 15, the model performance tends to converge. The main experimental report is based on the entropy-based difficulty judgment method.
[0034] In the sample set E, 200 problem samples are respectively selected as the training set and the validation set to balance performance and cost. The cross-entropy loss between the API-generated answer and the true answer is calculated, and the AdamW optimizer is used to perform multiple rounds of iterative optimization on the latent variables - the learning rate is set to , and the batch size is 5. After the reasoning chain optimization is completed, the normal-style prompts are stored in the default file. It is also necessary to reduce the performance fluctuations (i.e., dispersion) caused by changes in the prompt format style for the prompts. The value range of the dispersion is between 0.0 and 1.0. The closer it is to 0.0, the stronger the model robustness and the less sensitive it is to style changes; conversely, the closer it is to 1.0, the more sensitive the model is to prompt style changes. We use the designed "novel format" prompt template to require the large language model to rewrite each normal-style prompt as the final novel prompt. By default, when asking the large language model a reasoning question, the large language model will query the novel prompt similar to the question content, then combine the novel prompt with the question and pass it to the large language model, and then compare the generated answer with the true answer to calculate the answer accuracy of the large language model.
Claims
1. An automatic thought chain prompt generation method based on black-box optimization and vulnerability quantification, characterized in that, It includes the following steps: (1) Obtain an unlabeled training dataset; screen a subset of high-difficulty questions through a difficulty judgment module to generate a labeled dataset; (2) Use a large language model to generate multiple reasoning chains for each question in the labeled dataset, and retain the reasoning chains that are consistent with the correct answers to form an example library; (3) Adopt a variance reduction strategy gradient estimator to optimize the reasoning chain selection strategy and generate common-style prompts; (4) Perform style diversification on the common-style prompts through a novel format prompt module to generate anti-fragile automatic thinking chain prompts.
2. The automatic thinking chain prompt generation method based on black-box optimization and vulnerability quantification according to claim 1, wherein The operations of the difficulty judgment module include: the predicted probability distribution of the large language model generating k answers Calculate the entropy value: ; Among them, represents the parametric model generates the prediction probability of a specific answer for a given question ; represents the parameters of the language model; tunable hyperparameters and are used to control the probability scaling ratio; select the top n questions with the largest entropy value as high-difficulty questions.
3. The automatic thought chain prompt generation method based on black-box optimization and vulnerability quantification according to claim 2, characterized in that When the number of high-difficulty questions exceeds G, randomly select m questions; the construction of the labeled dataset is realized through automatic annotation by the large language model; where G is the set threshold for the number of high-difficulty questions.
4. The automatic thinking chain prompt generation method based on black-box optimization and vulnerability quantification according to claim 1, wherein, The generation and pruning of the reasoning chains include: generating l reasoning chains for each question; comparing the generated answers with the labeled answers, and only retaining the correctly derived reasoning chains to pair the questions with the correct reasoning chains to form an initial example set.
5. An automatic thinking chain prompt generation method based on black-box optimization and vulnerability quantification according to claim 1, characterized in that The optimization process of the variance reduction policy gradient estimator includes: defining a latent variable subject to a categorical distribution ; where the probability vector of N candidate examples; updating via the gradient estimation formula : ; where T represents the complete set of sample examples, is the i-th example, is the i-th example is the probability of being selected; from the set of examples sampled independently ; is the gradient estimate used to update ; is the gradient of the probability vector ; Indicates the difference between the loss of the current example set and the average loss of all example sets; Update the example distribution using projected stochastic gradient descent: ; Among them, is the learning rate, I is the scale of the example set, is the projection calculation.
6. The automatic thought chain prompt generation method based on black-box optimization and vulnerability quantification according to claim 1, wherein During the optimization process: loss function Calculated based on the cross-entropy between the output of the black-box model and the true labels; the number of context examples is dynamically adjusted to 3-6.
7. The automatic thought chain prompt generation method based on black box optimization and vulnerability quantification according to claim 1, wherein, The operations of the novel format prompt module include: designing a unique style template for each few-shot example, and the template types include: paraphrasing in a non-semantic markup format, visual structure adjustment, hierarchical optimization and reconstruction, nesting the common-style prompt into the style template to generate the final prompt.
8. An electronic device, characterized in that, It includes a memory and a processor, the memory stores a computer program, and when the processor executes the program, it implements the method described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, A computer program is stored, and when the program is executed by a processor, it implements the steps of the method described in any one of claims 1-7.
Citation Information
Patent Citations
Language model training method and device and computer readable storage medium
CN117669767A
Text processing model training method, text processing method and related equipment
CN119150982A
Alignment thinking chain-based science answering method and system
CN119293153A
Method and system for solving thinking chain reasoning mathematical problem based on feature classifier
CN119443267A
Natural language question and answer framework, method and device based on self-reflection
CN120011491A
Cited By
Large-scale language model enhanced table question and answer method based on thinking chain reasoning
CN120654835A
Question and answer large model training method, question and answer method and device, equipment and storage medium
CN120893511A