Knowledge-intensive task-oriented thinking chain prompt optimization method and system
By guiding the model to explicitly generate knowledge fragments and combining multi-dimensional evaluation loss functions to optimize thought chain prompts, the problem of knowledge gaps and errors in knowledge-intensive tasks is solved, and the accuracy of the model's knowledge generation and reasoning is improved.
Patent Information
- Application Number
- CN202511473921.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2026-01-16
AI Technical Summary
In knowledge-intensive tasks, existing technologies struggle to balance knowledge quality and the accuracy of predicted answers through thought chain prompts, resulting in issues such as missing, irrelevant, or erroneous reasoning performance in the models.
The system uses a knowledge generation module to explicitly generate knowledge fragments, and combines this with a knowledge evaluation module for multi-dimensional evaluation. It optimizes the thought chain prompts through candidate word initialization, gradient calculation, and prompt update modules, and introduces loss functions for knowledge format, relevance, and correctness to optimize prompt words and improve knowledge quality.
It improves the model's knowledge completeness, relevance, and accuracy in knowledge-intensive tasks, and enhances the knowledge generation ability and reasoning accuracy of the thought chain prompts.
Smart Images

Figure CN121352005A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of natural language processing technology, specifically relating to a method and system for optimizing thought chain prompts for knowledge-intensive tasks. It is primarily applied to large language models. Background Technology
[0002] Knowledge-intensive tasks are an important part of natural language processing. They typically rely on contextual knowledge, common sense, or domain-specific expertise, and require logical reasoning, information integration, or complex decision-making to complete. Compared to general language understanding tasks, knowledge-intensive tasks demand that language models possess stronger knowledge understanding, knowledge integration, and reasoning abilities. Typical tasks include common sense reasoning, medical question answering, and multi-hop question answering.
[0003] Chain-of-Thought Prompts effectively improve the interpretability and accuracy of model reasoning and decision-making by guiding the model to gradually generate intermediate reasoning basis, and have also achieved certain results in knowledge-intensive tasks. However, manually designed chain-of-thought prompts rely on human experience, have high construction costs, lack standardization and automatic optimization mechanisms, and are difficult to stimulate the model's optimal performance.
[0004] Therefore, researchers have proposed various automatic optimization methods for thought chain prompts. Among them, gradient-based thought chain prompting is one of the mainstream methods for prompt optimization. This method combines the downstream task objective with the design of a loss function, calculates the gradient information of this loss function with respect to the candidate words of the prompt, determines the update direction of the prompt based on the gradient information, and thus iteratively optimizes the prompt. This method has a high degree of automation, fine optimization granularity, and low computational resource requirements.
[0005] However, this method only focuses on the accuracy of predicted answers, neglecting the knowledge within the thought chain. This makes it difficult to handle issues such as knowledge gaps, irrelevance, and errors in the thought chain during knowledge-intensive tasks. Therefore, a method is urgently needed to optimize thought chain prompts by balancing the accuracy of predicted answers with the quality of knowledge within the thought chain, in order to more effectively enhance the reasoning performance of language models on knowledge-intensive tasks. Summary of the Invention
[0006] To address the issue that when models reason about knowledge-intensive tasks based on thought chain prompts, the generated thought chains often contain missing, irrelevant, or even incorrect knowledge, this invention provides a thought chain prompt optimization method and system for knowledge-intensive tasks. By combining knowledge generation and knowledge evaluation, the prompt optimization process is improved, resulting in thought chain prompts that pay more attention to knowledge quality. This enhances the completeness, relevance, and accuracy of the knowledge in the model's generated thought chains, ultimately improving the model's reasoning accuracy on knowledge-intensive tasks.
[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0008] A thought chain prompt optimization system for knowledge-intensive tasks, comprising a thought chain knowledge generation module, a knowledge evaluation module, a candidate word initialization module, a candidate word gradient calculation module, and a prompt update module;
[0009] The thought chain knowledge generation module: uses fixed prompt words to guide the model to explicitly output knowledge fragments in the thought chain;
[0010] The knowledge evaluation module extracts and evaluates knowledge fragments from the model-generated thought chain from multiple dimensions, and calculates the knowledge evaluation loss.
[0011] The candidate word initialization module calculates the replaceable candidate words at each position in the prompt based on each training example, and constructs a common candidate word set using a soft intersection strategy;
[0012] The candidate word gradient calculation module calculates the total loss and uses the total loss to calculate gradient information for the common candidate word set at each position in the prompt.
[0013] The prompt update module selects the best candidate word based on the common candidate word set and its gradient information, and replaces the words in each position of the original prompt to achieve optimized updating of the prompt words.
[0014] Furthermore, the thought chain knowledge generation module takes as input a concatenation of training examples, initial prompts, and fixed prompt words, and outputs a thought chain generated by the model that includes knowledge and reasoning answers.
[0015] The fixed prompt words are "Knowledge:" and "Answer:". The fixed prompt words remain unchanged during the update and optimization process of the initial prompt. Their function is to guide the model to output a knowledge fragment after "Knowledge:" and then output the reasoning answer after "Answer:".
[0016] This module aims to guide the model to output thought chains in a specific format through specific prompts, providing actionable textual evidence for subsequent explicit separation of knowledge fragments and reasoning answers, as well as automatic evaluation of knowledge completeness, relevance, and correctness.
[0017] Furthermore, the knowledge evaluation module takes knowledge fragments from the model-generated thought chain as input and outputs the knowledge evaluation loss obtained after evaluating the knowledge fragments in terms of knowledge format, knowledge relevance, and knowledge correctness.
[0018] The knowledge evaluation loss includes three loss items: knowledge format loss item, knowledge relevance loss item, and knowledge correctness loss item. Each loss item is measured by a corresponding function. The knowledge format loss item measures whether the knowledge fragments in the model's thinking chain are generated according to the format requirements. The knowledge relevance loss item measures the degree of relevance between the knowledge fragments in the model's thinking chain and the example problem. The knowledge correctness loss item measures whether the knowledge in the thinking chain is correct.
[0019] This module aims to provide a supervisory signal for the replacement and updating of prompt words through knowledge evaluation loss, so that the optimized thought chain prompts are more inclined to guide the model to generate knowledge fragments that are formatted correctly, semantically relevant and factually correct, thereby improving the knowledge quality of the model's generated thought chains and the accuracy of the final inference.
[0020] Furthermore, the candidate word initialization module takes as input an initial prompt and a set of training examples, and outputs a set of common candidate words that can be replaced at each position in the prompt;
[0021] The candidate word initialization module, based on the context information and partial hints of each training example question, uses the probability distribution of the language model to predict the top-k tokens with the highest probability of the hint word at the next position, which are then used as candidate words. Subsequently, a soft intersection strategy is used to construct a common candidate word set for all example questions. The common candidate word set is used for subsequent hint updates.
[0022] The soft intersection strategy is as follows: calculate the occurrence frequency of candidate words for all sample questions, set a frequency threshold t and a minimum common candidate word threshold n, select tokens with a frequency greater than the threshold t from the candidate words as common candidate words; if the number of common candidate words is less than the minimum common candidate word threshold n, select tokens with higher probabilities from the candidate words of all sample questions to supplement the common candidate words until the minimum common candidate word threshold n is met; after deduplication, a set of common candidate words is obtained.
[0023] This module aims to determine the feasible domain for suggestion optimization, that is, the limited range of alternative candidate words for each word in the original suggestion.
[0024] Furthermore, the candidate word gradient calculation module takes as input the predicted answer in the thought chain generated by the model for the training example question, the standard answer of the training example, and the common candidate word set based on the training example, and outputs the gradient information after differentiating the total loss with respect to the common candidate word set; this module includes answer prediction loss calculation, prompt confusion loss calculation, total loss combination, and gradient information calculation;
[0025] The answer prediction loss calculation calculates the cross-entropy loss between the standard answer and the predicted answer in the model-generated thought chain, which is used to measure the accuracy of the predicted answer based on the model thought chain.
[0026] The calculation of the prompt perplexity loss is performed by calculating the average negative log probability of each prompt word based on the sample question, thereby assessing the overall perplexity and measuring the linguistic fluency and semantic rationality of the prompt words.
[0027] The total loss combination is a weighted combination of the answer prediction loss, the hint confusion loss, and the knowledge evaluation loss to form the total loss function; the weight parameters of each loss item are set and optimized based on experiments.
[0028] The gradient information calculation involves differentiating the total loss with respect to the input embedding vector corresponding to the common candidate word at each position in the prompt words to obtain the gradient value of each common candidate word. This gradient value is used to measure the influence of different candidate words on the quality of the model's output thought chain knowledge and the accuracy of reasoning, and to guide the optimization and updating of subsequent prompt words.
[0029] Furthermore, the prompt update module takes a common candidate word set and its gradient information as input, and outputs an optimized prompt obtained by replacing the words at each position of the original prompt with the best candidate word; this module includes best candidate word selection, prompt word replacement, and optimization.
[0030] The optimal candidate word selection involves choosing the word that causes the target loss to decrease the most at each prompt position from the corresponding public candidate word set, i.e., the word with the largest negative gradient direction, as the optimal candidate word.
[0031] The prompt word replacement and optimization involves replacing the words in each position of the original prompt with the best candidate words to obtain the optimized prompt, thereby achieving the optimization and updating of prompt words.
[0032] A method for optimizing thought chain prompts for knowledge-intensive tasks, comprising the aforementioned system; the method includes the following steps:
[0033] Step 1, Knowledge Generation of Mind Chain: The sample question, initial hints and fixed hint words are concatenated and input into the language model to guide the model to output a mind chain containing knowledge fragments;
[0034] Step 2, Knowledge Evaluation Loss Calculation: Extract knowledge fragments from the thought chain generated by the model in Step 1, perform multi-dimensional knowledge evaluation on them, and calculate the knowledge evaluation loss.
[0035] Step 3: Initialize the common candidate word set of the prompts: Calculate the replaceable candidate words for each position in the prompts based on each sample question and part of the prompts, and then construct a common candidate word set for all sample questions using a soft intersection strategy;
[0036] Step 4, Calculate the total loss and candidate word gradient: Calculate the total loss, which includes the answer prediction loss, knowledge evaluation loss, and hint confusion loss. Take the gradient derivative of the total loss with respect to the common candidate words at each position in the hint to obtain the gradient information of the common candidate words.
[0037] Step 5, prompt update and optimization: Based on the common candidate word set obtained in Step 3 and the common candidate word gradient information calculated in Step 4, the best candidate word is automatically selected and the words in each position of the prompt are replaced to achieve the optimization and update of the prompt words.
[0038] Furthermore, in step 1, knowledge generation of the thought chain: the fixed prompt words are "Knowledge:" and "Answer:". The fixed prompt words remain unchanged during the update and optimization process of the initial prompt. Their function is to guide the model to output a knowledge fragment after "Knowledge:" and then output a reasoning answer after "Answer:".
[0039] Furthermore, in step 2, the knowledge evaluation loss is calculated as follows: the knowledge evaluation loss... It includes three loss terms: knowledge format loss, knowledge relevance loss, and knowledge correctness loss. Each loss term is measured by a corresponding function, and the specific formula is as follows:
[0040]
[0041] Among them, knowledge format loss item The formula for measuring whether knowledge fragments are generated in the model's thought process chain according to the required format is as follows:
[0042]
[0043] It is a very small positive number, such as 1e. -6 To avoid the denominator of the knowledge format loss item being 0;
[0044] This refers to the number of items that meet the format requirements, which have two aspects: whether "Knowledge:" is generated as required in the model's thought process; and whether there is a knowledge segment between "Knowledge:" and "Answer:" in the model's thought process. The string matching function is used to determine if the format requirements are met; if either requirement is met... When both conditions are met ,otherwise ;
[0045] Knowledge-related loss term The specific formula for measuring the relevance between knowledge fragments in the model's thought process chain and sample questions is as follows:
[0046]
[0047] and They are knowledge fragments And the vector representation of the sample problem, It is a knowledge fragment Similar to the sample problem Cosine similarity;
[0048] Knowledge correctness loss item To assess the correctness of knowledge within a thought process chain, a knowledge correctness scoring model is used. The correctness score of the output is used for measurement, and the specific formula is as follows:
[0049]
[0050] It is a knowledge fragment The vector representation of , It is a knowledge correctness scoring model fine-tuned based on the T5 base model. This refers to the correctness score of the knowledge fragment output using the Vera model. The range of the correctness score is... A higher score indicates that the knowledge fragment is more likely to be correct; therefore, the correctness loss for a single knowledge fragment is... ;
[0051] , , The weighting parameters for the loss terms are used to balance the importance of each knowledge loss term and can be set and optimized based on experiments. This is a format indicator factor, set to 1 if the thought chain generated by the model contains knowledge, and 0 otherwise.
[0052] Furthermore, step 3, which initializes the set of public candidate words provided, specifically includes the following sub-steps:
[0053] Step 3.1, Candidate word selection for a single sample: For the sample problem in the training set and some hints Using language models The probability distribution of prompt words is used to calculate the prompt words. The top k tokens with the highest probabilities are selected as candidate words. The details are as follows:
[0054]
[0055] in, Indicates text concatenation;
[0056] Step 3.2, Construction of the common candidate word set: A soft intersection strategy is used to construct a common candidate word set. The soft intersection strategy is as follows: calculate the candidate words for all sample questions. The frequency of occurrence is determined, and a frequency threshold t and a minimum number of common candidate words threshold n are set, from which candidate words... Select tokens with a frequency greater than a threshold t as common candidate words; if the number of common candidate words is less than the minimum threshold n for common candidate words, then select candidate words from all sample questions. Select tokens with higher probabilities to supplement the common candidate words until the minimum number of common candidate words threshold n is met; after deduplication, the set of common candidate words is obtained.
[0057] Furthermore, step 4, the calculation of the total loss and candidate word gradient, specifically includes the following sub-steps:
[0058] Step 4.1, calculate the answer prediction loss. Model acquisition based on problem and prompts The generated reasoning chain text r is compared with the answer extraction hints. The concatenated data is used as new input to allow the model to generate a predicted answer again. The computational model predicts the answer. Compared with the standard answer Cross-entropy loss This is used to measure the accuracy of the predicted answer obtained by the model based on the inference chain, and the specific formula is as follows:
[0059]
[0060]
[0061] in, This indicates text concatenation and provides hints for answer extraction. It's about splicing together the problem. ,hint The model generates a fixed prompt word after the inference chain text r, usually "Therefore, the final answer is", which guides the model to output a prediction of the answer after fully considering the generated knowledge and inference chain.
[0062] Step 4.2, calculate the cue confusion loss. Prompts based on word-by-word calculation Based on sample questions The negative logarithmic probability average is used to assess the overall perplexity of the prompt, which is used to measure the linguistic fluency and semantic plausibility of the prompt words. The specific formula is as follows:
[0063]
[0064] in, It's the length of the prompt. It's a hint The i-th word in the middle, It's a hint The hints before the i-th word The model is in a known problem and some hints Generate the i-th word under the following conditions The probability, This represents the average negative logarithmic probability. Represents an exponential function;
[0065] Step 4.3, Total Loss Calculation: Predict the loss of the answer. , and the loss of confusion. and the knowledge evaluation loss obtained in step (2) By performing a weighted combination, the total loss function is obtained. The specific formula is as follows:
[0066]
[0067] in These are the weighting parameters for each loss, which can be set and optimized based on experiments;
[0068] Step 4.4, Gradient information calculation: Total loss The input embedding vector corresponding to the common candidate word at each position in the prompt words. By taking the derivative, we obtain the gradient value of each common candidate word, which represents the influence of the model's output knowledge and reasoning process on the correct answer, as follows:
[0069] .
[0070] Furthermore, step 5, which prompts for updates and optimizations, specifically includes the following sub-steps:
[0071] Step 5.1, Optimal Candidate Word Selection: Based on the gradient information obtained in Step 4.4, for each position indicated, select the word that causes the largest decrease in the target loss from the corresponding public candidate word set, i.e., the word with the largest negative gradient direction, as the optimal candidate word. The specific formula is as follows:
[0072]
[0073] Step 5.2, Prompt word replacement and iterative optimization: Replace the original prompt words... Replace with the selected best candidate word The updated prompts are then used in the next round of optimization, thus achieving an iterative optimization process for the prompts.
[0074] Compared with the prior art, the present invention has the following advantages:
[0075] This invention improves the prompt optimization process by combining knowledge generation and knowledge evaluation, resulting in thought chain prompts that focus more on knowledge quality, thereby enhancing the completeness, relevance, and accuracy of the knowledge generated by the model. Knowledge generation refers to guiding the model to explicitly generate knowledge fragments within the thought chain; knowledge evaluation involves designing a loss function that includes answer accuracy and multi-dimensional knowledge indicators, and using its gradient information to select the best candidate words to update and optimize the prompts. The system of this invention includes: a thought chain knowledge generation module, a knowledge evaluation module, a candidate word initialization module, a candidate word gradient calculation module, and a prompt update module.
[0076] (1) This invention introduces the focus on the knowledge of the thinking chain, guides the model to explicitly generate knowledge fragments through fixed prompt words, and makes the knowledge inside the model explicit into operable and quantifiable text information, so that the thinking chain prompts can be optimized based on more granular information.
[0077] (2) This invention designs a multi-dimensional knowledge evaluation index that includes knowledge format, relevance and correctness, and combines the accuracy of the answer and the perplexity of the prompt words to form a composite loss function, which provides more accurate guidance for gradient calculation and thinking chain prompt optimization. Compared with other prompt optimization methods that are only guided by the correctness of the answer, it can more effectively improve the quality of thinking chain knowledge and reasoning accuracy of the model in knowledge-intensive tasks.
[0078] (3) In the candidate word initialization module, the present invention uses a soft intersection strategy to construct a common candidate word set for multiple sample candidate words. Compared with other methods that use a strict hard intersection strategy, it can better ensure the diversity of the candidate word set and the rationality of the gradient optimization feasible region. Attached Figure Description
[0079] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0080] Figure 1 The flowchart shows the thought chain prompting optimization method and system for knowledge-intensive tasks according to the present invention.
[0081] Figure 2 This is an overall structural diagram of all modules of the system described in this invention;
[0082] Figure 3 This is a flowchart illustrating the knowledge evaluation module of the system described in this invention.
[0083] Figure 4 This is a flowchart of the candidate word initialization module of the system described in this invention;
[0084] Figure 5 This is a flowchart of the candidate word gradient calculation module of the system described in this invention. Detailed Implementation
[0085] To gain a deeper understanding of this invention, we will provide a comprehensive and detailed description. However, this invention has various implementations and is not limited to the specific examples listed herein. These examples are presented to enhance a full understanding of the disclosure of this invention.
[0086] Example 1
[0087] like Figure 1 As shown in the figure, this embodiment presents a method and system for optimizing thought chain prompts for knowledge-intensive tasks. This system improves the prompt optimization process by combining knowledge generation and knowledge evaluation, achieving automatic optimization of thought chain prompts that focuses more on high-quality knowledge generation. This helps improve the completeness, relevance, and correctness of the knowledge generated by the model's thought chain, ultimately enhancing the accuracy of the model's reasoning. Knowledge generation refers to guiding the model to explicitly generate knowledge fragments within the thought chain; knowledge evaluation involves designing a loss function that includes answer accuracy and multi-dimensional knowledge indicators, and using its gradient information to select the best candidate words to update and optimize the prompts. The system mainly includes: a thought chain knowledge generation module, a knowledge evaluation module, a candidate word initialization module, a candidate word gradient calculation module, and a prompt update module.
[0088] Mind Chain Knowledge Generation Module: Uses fixed prompts to guide the model to explicitly output knowledge fragments in the mind chain.
[0089] Knowledge evaluation module: Extracts and evaluates knowledge fragments from the model-generated thought chain from multiple dimensions, and calculates the knowledge evaluation loss.
[0090] Candidate word initialization module: Calculates the replaceable candidate words at each position in the prompt based on each training example, and constructs a common candidate word set using a soft intersection strategy.
[0091] Candidate word gradient calculation module: Calculates the total loss and uses the total loss to calculate gradient information for the common candidate word set at each position in the prompt.
[0092] The prompt update module selects the best candidate word based on the common candidate word set and its gradient information, and replaces the words in each position of the original prompt to achieve optimized updating of the prompt words.
[0093] Furthermore, the knowledge evaluation module includes the following sub-modules: knowledge format loss, knowledge relevance loss, and knowledge correctness loss.
[0094] Furthermore, the candidate word initialization module includes the following sub-modules: candidate word selection for a single example and construction of a public candidate word set.
[0095] Furthermore, the candidate word gradient calculation module includes the following sub-modules: predicted answer loss calculation, prompt confusion loss calculation, total loss combination, and gradient information calculation.
[0096] Furthermore, the prompt update module includes the following sub-modules: best candidate word selection, prompt word replacement, and optimization.
[0097] like Figure 3 As shown, the knowledge evaluation module in this embodiment is implemented as follows:
[0098] The input to this module is the knowledge fragments in the model-generated thought chain, and the output is the knowledge evaluation loss obtained after evaluating the knowledge fragments in terms of knowledge format, knowledge relevance, and knowledge correctness.
[0099] The knowledge evaluation loss described in this module includes three loss items: knowledge format loss, knowledge relevance loss, and knowledge correctness loss. Each loss item is measured by a corresponding function, and the specific formula is as follows:
[0100]
[0101] Among them, knowledge format loss item The formula for measuring whether knowledge fragments are generated in the model's thought process chain according to the required format is as follows:
[0102]
[0103] It is a very small positive number, such as 1e. -6 To avoid the denominator of the knowledge format loss term being 0.
[0104] This refers to the number of items that meet the format requirements, which have two aspects: whether "Knowledge:" is generated as required in the model's thought process; and whether there is a knowledge segment between "Knowledge:" and "Answer:" in the model's thought process. The string matching function is used to determine if the format requirements are met; if either requirement is met... When both conditions are met ,otherwise .
[0105] Knowledge-related loss term The specific formula for measuring the relevance between knowledge fragments in the model's thought process chain and sample questions is as follows:
[0106]
[0107] and They are knowledge fragments And the vector representation of the sample problem, It is a knowledge fragment Similar to the sample problem The cosine similarity.
[0108] Knowledge correctness loss item To assess the correctness of knowledge within a thought process chain, a knowledge correctness scoring model is used. The correctness score of the output is used for measurement, and the specific formula is as follows:
[0109]
[0110] It is a knowledge fragment The vector representation of , It is a knowledge correctness scoring model fine-tuned based on the T5 base model. This refers to the correctness score of the knowledge fragment output using the Vera model. The range of the correctness score is... A higher score indicates that the knowledge fragment is more likely to be correct; therefore, the correctness loss for a single knowledge fragment is... .
[0111] , , The weighting parameters for the loss terms are used to balance the importance of each knowledge loss term and can be set and optimized based on experiments. This is a format indicator factor, set to 1 if the thought chain generated by the model contains knowledge, and 0 otherwise.
[0112] like Figure 4 As shown, the candidate word initialization module in this embodiment is implemented as follows:
[0113] The input to this module is an initial prompt and a set of training examples, and the output is a set of common candidate words that can be replaced at each position in the prompt. This module includes candidate word selection for individual examples and construction of the common candidate word set.
[0114] Candidate word selection for a single sample: a sample problem for the training set. and some hints Using language models The probability distribution of prompt words is used to calculate the prompt words. The top k tokens with the highest probabilities are selected as candidate words. The details are as follows:
[0115]
[0116] The common candidate word set is constructed using a soft intersection strategy, which involves calculating candidate words for all sample questions. The frequency of occurrence is determined, and a frequency threshold t and a minimum number of common candidate words threshold n are set, from which candidate words... Select tokens with a frequency greater than a threshold t as common candidate words; if the number of common candidate words is less than the minimum threshold n for common candidate words, then select candidate words from all sample questions. Select tokens with higher probabilities to supplement the common candidate words until the minimum number of common candidate words threshold n is met; after deduplication, the common candidate word set is obtained.
[0117] This module aims to determine the feasible domain for suggestion optimization, that is, the limited range of alternative candidate words for each word in the original suggestion.
[0118] After experimental optimization, the top-k parameter value of this module was set to 3, the frequency threshold t was set to 0.5, and the minimum number of common candidate words threshold n was set to 10. These parameter settings are merely illustrative and the invention is not limited thereto.
[0119] like Figure 5 As shown, the candidate word gradient calculation module in this embodiment is implemented as follows:
[0120] The input to this module is the predicted answer in the thought chain generated by the model for the training example question, the standard answer of the training example, and the common candidate word set based on the training example. The output is the gradient information after differentiating the total loss with respect to the common candidate word set. This module includes the calculation of answer prediction loss, the calculation of hint confusion loss, the combination of total loss, and the calculation of gradient information.
[0121] The answer prediction loss calculation: The model is obtained based on the question. and prompts The generated reasoning chain text r is compared with the answer extraction hints. The concatenated data is used as new input to allow the model to generate a predicted answer again. The computational model predicts the answer. Compared with the standard answer Cross-entropy loss This is used to measure the accuracy of the predicted answer obtained by the model based on the inference chain, and the specific formula is as follows:
[0122]
[0123]
[0124] Among them, the answer extraction hints It's about splicing together the problem. ,hint The model generates a fixed prompt following the inference chain text 'r', usually "Therefore, the final answer is '". Its purpose is to guide the model to output a prediction of the answer after fully considering the generated knowledge and inference chain.
[0125] The calculation of the confusion loss of the prompt is performed word-by-word. Based on sample questions The negative logarithmic probability average is used to assess the overall perplexity of the prompt, which is used to measure the linguistic fluency and semantic plausibility of the prompt words. The specific formula is as follows:
[0126]
[0127] in, It is the i-th word in the prompt. It provides hints for the part before the i-th word. The model is in a known problem and some hints Generate the i-th word under the following conditions The probability, This represents the average negative logarithmic probability. Represents an exponential function;
[0128] The total loss calculation includes the loss from predicting the answer. , and the loss of confusion. and the knowledge evaluation loss obtained in step (2) By performing a weighted combination, the total loss function is obtained. The specific formula is as follows:
[0129]
[0130] in These are the weighting parameters for each loss, which can be set and optimized based on experiments.
[0131] The gradient information is used to calculate the total loss. The input embedding vector corresponding to the common candidate word at each position i in the prompt words. The gradient value of each common candidate word is obtained by taking the derivative. This gradient value is used to measure the impact of different candidate words on the quality of the model's output thought chain knowledge and the accuracy of reasoning, and guides the optimization and updating of subsequent prompt words, as follows:
[0132] .
[0133] Contents not described in detail in this specification are prior art known to those skilled in the art. Although illustrative specific embodiments of the invention have been described above to facilitate understanding by those skilled in the art, it should be understood that the invention is not limited to the scope of the specific embodiments. Various modifications are readily apparent to those skilled in the art as long as they fall within the spirit and scope of the invention as defined and determined by the appended claims, and all inventions utilizing the concept of this invention are protected.
Claims
1. A thinking chain prompting optimization system for knowledge-intensive tasks, characterized by: The system comprises a thought chain knowledge generation module, a knowledge evaluation module, a candidate word initialization module, a candidate word gradient calculation module, and a prompt updating module. The thought chain knowledge generation module: uses a fixed prompt word guide model to explicitly output knowledge fragments in a thought chain. The knowledge evaluation module: extracts and evaluates knowledge fragments in the thought chain generated by the model in multiple dimensions, and calculates a knowledge evaluation loss. The candidate word initialization module: calculates replaceable candidate words for each position in the prompt according to each training example, and uses a soft intersection strategy to construct a public candidate word set. The candidate word gradient calculation module: calculates the total loss, and uses the total loss to calculate gradient information for the public candidate word set at each position in the prompt. The prompt updating module: selects the best candidate word from the public candidate word set and its gradient information, replaces the words at each position in the original prompt, and optimizes and updates the prompt words.
2. The thought chain prompt optimization system for knowledge-intensive tasks according to claim 1, wherein: The thought chain knowledge generation module: the input is the concatenation of the training example, the initial prompt, and the fixed prompt word, and the output is the thought chain generated by the model containing knowledge and reasoning answers. The fixed prompt word is "Knowledge:" and "Answer:". The fixed prompt word remains unchanged during the update and optimization process of the initial prompt, and its function is to guide the model to first output knowledge fragments after "Knowledge:", and then output reasoning answers after "Answer:". This module aims to guide the model to output thought chains in a specific format through specific prompts, and to provide operational text basis for subsequent explicit separation of knowledge fragments and reasoning answers, as well as automatic evaluation of knowledge completeness, relevance, and correctness.
3. The thought chain prompt optimization system for knowledge-intensive tasks according to claim 1, wherein: The knowledge evaluation module: the input is the knowledge fragment in the thought chain generated by the model, and the output is the knowledge evaluation loss obtained after evaluating the knowledge format, knowledge relevance, and knowledge correctness of the knowledge fragment. The knowledge evaluation loss includes three loss terms: knowledge format loss term, knowledge relevance loss term, and knowledge correctness loss term. Each loss term is measured by a corresponding function. The knowledge format loss term is used to measure whether the model generates knowledge fragments in the thought chain according to the format requirements, the knowledge relevance loss term is used to measure the relevance of the knowledge fragments in the thought chain to the sample question, and the knowledge correctness loss term is used to measure whether the knowledge in the thought chain is correct. This module aims to provide a supervision signal for the replacement and update of the prompt word through the knowledge evaluation loss, so that the optimized thought chain prompt is more inclined to guide the model to generate knowledge fragments that are format-specific, semantically relevant, and factually correct, thereby improving the knowledge quality of the model-generated thought chain and the final reasoning accuracy.
4. The thought chain prompt optimization system for knowledge-intensive tasks according to claim 1, wherein: The candidate word initialization module: the input is the initial prompt and a set of training examples, and the output is a public candidate word set that can be replaced at each position in the prompt. The candidate word initialization module predicts the top-k tokens with the maximum probability of the prompt word at the next position as candidate words based on the context information and part of the prompt of each training example question, using the probability distribution of the language model; then a soft intersection strategy is used to construct a common candidate word set for all example questions; and the common candidate word set is used for subsequent prompt updating; The soft intersection strategy is: calculating the frequency of the candidate words of all example questions, setting a frequency threshold t and a minimum number threshold n of common candidate words, and selecting tokens with a frequency greater than the threshold t from the candidate words as common candidate words; if the number of common candidate words is less than the minimum number threshold n of common candidate words, tokens with a larger probability are selected from the candidate words of all example questions to supplement the common candidate words until the minimum number threshold n of common candidate words is met; After deduplication, a common candidate word set is obtained; This module aims to determine the feasible region of prompt optimization, i.e., the limited range of replaceable candidate words for each word in the original prompt.
5. The thought chain prompt optimization system for knowledge-intensive tasks according to claim 1, characterized in that: The candidate word gradient calculation module inputs are the predicted answer in the thought chain generated by the model for the training example question, the standard answer of the training example, and the common candidate word set based on the training example, and the output is the gradient information of the total loss after derivation on the common candidate word set; this module includes answer prediction loss calculation, prompt perplexity loss calculation, total loss combination, and gradient information calculation; The answer prediction loss calculation calculates the cross-entropy loss between the standard answer and the predicted answer in the thought chain generated by the model, which is used to measure the accuracy of the predicted answer based on the model thought chain; The prompt perplexity loss calculation calculates the average value of the negative logarithmic probability of the prompt word based on the example question word by word, and evaluates the overall perplexity accordingly, which is used to measure the language fluency and semantic rationality of the prompt word; The total loss combination combines the answer prediction loss term, the prompt perplexity loss term, and the knowledge evaluation loss by weighting, to form a total loss function; the weight parameters of each loss term are set and optimized according to experiments; The gradient information calculation derives the input embedding vector corresponding to each position of the prompt word for each common candidate word based on the total loss, to obtain the gradient value of each common candidate word, which is used to measure the influence degree of different candidate words on the knowledge quality and reasoning accuracy of the model output thought chain, and guide the subsequent optimization and update of the prompt word.
6. The thought chain prompt optimization system for knowledge-intensive tasks according to claim 1, characterized in that: The prompt updating module inputs are the common candidate word set and its gradient information, and outputs the optimized prompt obtained by replacing the words at each position of the original prompt with the best candidate word; this module includes best candidate word selection, prompt word replacement, and optimization; The best candidate word selection selects the word with the maximum descending amplitude of the target loss, i.e., the word with the maximum negative gradient direction, from the corresponding common candidate word set for each position of the prompt, as the best candidate word. The prompt word replacement and optimization replace the words in each position of the original prompt with the best candidate words to obtain an optimized prompt, thereby realizing the optimization and update of the prompt words.
7. A method for optimizing a train-of-thought hint for a knowledge-intensive task, the method comprising: receiving a task description; receiving a train-of-thought hint; and determining a score for the train-of-thought hint based on the task description. The system comprises the thinking chain prompt optimization method according to any one of claims 1-6, and the thinking chain prompt optimization method comprises the following steps: Step 1, thinking chain knowledge generation: inputting the sample question, the initial prompt and the fixed prompt word into the language model after splicing to guide the model to output the thinking chain containing the knowledge fragment; Step 2, knowledge evaluation loss calculation: extracting the knowledge fragment from the thinking chain generated by the model in step 1, and performing multi-dimensional knowledge evaluation on the knowledge fragment to calculate the knowledge evaluation loss; Step 3, initialization of the public candidate word set of the prompt: calculating the replaceable candidate words in each position of the prompt according to each sample question and partial prompt, and then constructing the public candidate word set by using the soft intersection strategy on the candidate words of all sample questions; Step 4, total loss and candidate word gradient calculation: calculating the total loss comprising the answer prediction loss term, the knowledge evaluation loss term and the prompt perplexity loss term, and performing gradient derivation on the public candidate words in each position of the prompt to obtain the gradient information of the public candidate words; Step 5, prompt update and optimization: automatically selecting the best candidate word according to the public candidate word set obtained in step 3 and the public candidate word gradient information calculated in step 4, and replacing the words in each position of the prompt, thereby realizing the optimization and update of the prompt words.
8. The method of claim 7, wherein the method is characterized by: In step 1, the fixed prompt word is "Knowledge:" and "Answer:", and the fixed prompt word remains unchanged in the update and optimization process of the initial prompt, and the function is to guide the model to first output the knowledge fragment after "Knowledge:", and then output the reasoning answer after "Answer:"; The step 2, knowledge evaluation loss calculation: the knowledge evaluation loss , including three loss terms: knowledge format loss term, knowledge relevance loss term, knowledge correctness loss term, each loss term is measured by a corresponding function, and the specific formula is as follows: ; wherein the knowledge format loss term The following formula is used to measure whether the knowledge segment is generated according to the format requirement in the model thinking chain: ; is a sufficiently small positive number to avoid division by zero in the knowledge format loss term; is the number of items that meet the format requirements, which have two items: whether "Knowledge:" is generated in the thinking chain of the model according to the requirements; whether there is a knowledge fragment between "Knowledge:" and "Answer:" in the thinking chain of the model; whether the format requirements are met by the string matching function, when one of the two items is met , when both items are met , otherwise ; knowledge relevance loss term The relevance of the knowledge piece to the example question in the model thinking chain is measured, and the specific formula is as follows: ; and are vector representations of knowledge snippets and example questions, are vector representations of knowledge snippets and example questions are cosine similarities between knowledge correctness loss term To measure whether the knowledge in the thought chain is correct, a knowledge correctness scoring model is used The correctness score of the output is measured, and the specific formula is as follows: ; is a vector representation of a knowledge piece , is a knowledge correctness scoring model based on fine-tuning of the T5 base model, is the correctness score output by the Vera model for the knowledge piece, the correctness score ranges from , the greater the score, the more likely the knowledge piece is correct, so the correctness loss of a single knowledge piece is ; 、 、 is a weight parameter of the loss term, used to balance the importance of each knowledge loss term, which can be set and optimized according to experiments; is a format indicating factor, taking 1 when the thought chain generated by the model contains knowledge, otherwise taking 0. 9.The knowledge-intensive task-oriented thought chain prompting optimization method according to claim 7, characterized in that: In step 3, the initialization of the public candidate word set of the prompt specifically comprises the following sub-steps: Step 3.1, candidate word selection for individual examples: for example problems of the training set and partial cues , using the probability distribution of the language model , compute the top k tokens with the highest probability of being the cue word as candidate words , as follows: ; wherein represents text concatenation; Step 3.2, Construction of the common candidate word set: A soft intersection strategy is used to construct a common candidate word set. The soft intersection strategy is as follows: calculate the candidate words for all sample questions. The frequency of occurrence is determined, and a frequency threshold t and a minimum number of common candidate words threshold n are set, from which candidate words... Select tokens with a frequency greater than a threshold t as common candidate words; if the number of common candidate words is less than the minimum threshold n for common candidate words, then select candidate words from all sample questions. Select tokens with higher probabilities to supplement the common candidate words until the minimum number of common candidate words threshold n is met; after deduplication, the set of common candidate words is obtained. In step 4, the total loss and candidate word gradient calculation specifically comprises the following sub-steps: Step 4.1, calculate the answer prediction loss. Model acquisition based on problem and prompts The generated reasoning chain text r is compared with the answer extraction hints. The concatenated data is used as new input to allow the model to generate a predicted answer again. The computational model predicts the answer. Compared with the standard answer Cross-entropy loss This is used to measure the accuracy of the predicted answer obtained by the model based on the inference chain, and the specific formula is as follows: ; ; Among them, the answer extraction prompt is a fixed prompt word spliced after the question , prompt , model generation inference chain text r, generally "Therefore, the final answer is ", which guides the model to output the prediction of the answer after fully considering the generated knowledge and inference chain; Step 4.2, calculate the cue confusion loss. Prompts calculated word by word Based on sample questions The negative logarithmic probability average is used to assess the overall perplexity of the prompt, which is used to measure the linguistic fluency and semantic plausibility of the prompt words. The specific formula is as follows: ; wherein, is the length of the prompt, is the prompt in the i-th word, is the prompt up to the i-th word, is the probability that the model generates the i-th word in the prompt and the partial prompt given the known problem in the i-th word of the prompt, denotes the average negative log probability, denotes the exponential function; Step 4.3, total loss calculation: combine the answer prediction loss , the confusion loss , and the knowledge evaluation loss from step 2 to get the total loss function , which is given by the following formula: ; wherein are weight parameters of each loss, which can be set and tuned according to experiments; Step 4.4, Gradient information calculation: total loss input embedding vectors corresponding to the common candidate words for each position in the prompt word Take the derivative to get the gradient value of each common candidate word, which represents the influence of the knowledge and reasoning process of the model output on the correct answer of reasoning, as follows: 。 10. The method of claim 7, wherein the method is a method of thinking chain prompting optimization for knowledge-intensive tasks. In step 5, the prompt update and optimization specifically comprises the following sub-steps: Step 5.1, best candidate word selection: according to the gradient information obtained in step 4.4, for each position of the prompt, select the word with the largest decrease in target loss from the corresponding set of public candidate words, that is, the word with the largest negative gradient direction as the best candidate word The specific formula is as follows: ; Step 5.2, prompt word replacement and iterative optimization: the original prompt word is replaced by the selected best candidate word , and the updated prompt is further used for the next round of optimization, thus realizing the iterative optimization process of the prompt.
Citation Information
Cited By
Pancreatic cancer prediction method and system based on local and global confusion weighted pruning
CN121839088A