Metaevidence perception prompt optimization method for large language model

Through clustering and Bayesian inference, identify the causal relationship of the evidence output by large language models, construct DAG to eliminate redundant information, optimize initial prompts, solve the problem of difficulty in redundant information input and meta-evidence identification in the existing technology, and achieve more efficient prompt optimization effects.

CN119939294AActive Publication Date: 2025-05-06SOUTHEAST UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510101991.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-06
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The prompt optimization method of existing large language models has the problem of redundant information input, which leads to increased misleading and computational costs, and it is difficult to effectively identify meta-evidence.

Method used

The combination of clustering and Bayesian reasoning is used to carefully analyze the output evidence of the large language model, identify the causal relationship between the evidence, eliminate redundant information by constructing a directed acyclic graph (DAG), and optimize the initial prompt based on meta-evidence.

Benefits of technology

It effectively reduces the input of redundant information, improves the inference ability and computing efficiency of large language models, and improves the performance of prompt optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939294A_ABST
    Figure CN119939294A_ABST
Patent Text Reader

Abstract

The invention discloses a meta-evidence perception prompt optimization method of a large language model. According to the method, causal relationships in evidences are finely analyzed by utilizing clustering and Bayesian reasoning. Firstly, the MEPO method utilizes the feedback of a large model to generate preliminary evidence to describe the defect of a given prompt; secondly, clustering the preliminary evidences to eliminate redundant information, obtaining the probability of a clustering center by adopting a Bayesian inference method, and modeling a causal relationship between each pair of evidences to obtain a directed graph; in addition, a directed acyclic graph (DAG) is constructed by applying a maximum spanning tree algorithm, so that the reasoning capability of the LLMs is enhanced. And finally, the initial prompt is edited according to the obtained meta evidence, and the prompt with the best performance is selected by adopting beam search. Experimental results show that the MEPO method shows good performance in a plurality of tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of large language model optimization, and specifically is a meta-evidence perception prompt optimization method for a large language model. Background Art

[0002] Large language models (LLMs) are an important part of the deep learning field, specifically designed for natural language processing (NLP) tasks, and have achieved remarkable success in various downstream applications. To fully tap the potential capabilities of large language models, task-specific prompts need to be carefully designed to integrate detailed instructions and domain insights. Prompt optimization aims to construct effective prompts to fully exploit the capabilities of large language models. The key challenge of prompt optimization is to analyze the output of large language models, identify fine-grained errors as clear prompt defects, and guide reverse prompt correction. However, the existing mainstream methods input all evidence into the large language model, resulting in excessive input of redundant information that misleads the large language model. In addition, this redundant input may increase the computational cost. Therefore, this patent needs to explore new methods to reduce the input of redundant information and optimize the design of prompts for large language model tasks.

[0003] There are two main approaches to the research on hint optimization: soft hint optimization (SPO) and hard hint optimization (HPO). Figure 1 ,The initial prompt and task example are given as follows: Prompt – Determine whether the statement is a lie based on other information and external knowledge, Statement – ​​It will take years to build a wall on the border.

[0004] Soft hint optimization is an important existing method to address these challenges. These methods optimize hints by training continuous vectors (i.e., the representation of hints). Soft hints are continuous and learnable vectors that do not require manual design and can be automatically optimized for specific datasets through gradient optimization methods to adapt to different tasks. However, continuous vectors are not interpretable and have poor flexibility. Since the internal variables of the large language model need to be accessed, this optimization process may introduce noise and information loss. Figure 1 As shown in the lower left, the large language model generates evidence as the disadvantage of the initial prompt, and then edits the prompt according to the corresponding evidence. However, the output ignores the ambiguity of the definition of lies in the initial prompt and introduces noise (note the accuracy of the data), which affects the performance of the edited prompt.

[0005] Hard hint optimization uses the iterative feedback process of LLMs to improve hints. Compared with soft hint optimization methods, these methods do not need to access the internal variables of the large language model, thus reducing the risk of noise and information loss. However, hard hint optimization methods still have problems: they cannot find meta-evidence. Meta-evidence is obtained by eliminating redundant information from key hint defects. Figure 1As shown in the bottom right, the initial prompts Other Information and External Knowledge have similar meanings. First, the large language model generates broadly defined defects such as "Other Information" and "External Knowledge" in the initial prompt as evidence. Then, the hard prompt refinement method edits the prompt into corresponding evidence based on the given context and avoids including external details. The two new prompts are semantically similar, causing the large language model to waste computational resources. Summary of the invention

[0006] To solve the above problems, the present invention proposes a meta-evidence-aware prompt optimization method for a large language model, which uses clustering and Bayesian reasoning to perform a detailed analysis of the causal relationship within the evidence.

[0007] To achieve the above object, the technical solution adopted by the present invention is:

[0008] A meta-evidence-aware prompt optimization method for a large language model, characterized by the following specific steps:

[0009] Step 1: Generate preliminary evidence;

[0010] Step binary evidence identification;

[0011] Step 3: Optimize evidence-aware prompts.

[0012] As a further improvement of the present invention, the step 1 generates preliminary evidence, which is specifically as follows:

[0013] Preliminary evidence introduced As feedback from the model, the shortcomings of the current prompt are pointed out, and the mini-batch data is As a prompt p0 input into the large language model, thus obtaining the loss signal Initial evidence generation can be formalized as follows:

[0014]

[0015] The preliminary evidence was presented in a sequential format, and to enhance the relevance of the output, the amount of evidence was limited by adding prompts: Give five reasons why the prompts misunderstand the examples.

[0016] As a further improvement of the present invention, the step of binary evidence recognition is specifically as follows:

[0017] After generating preliminary evidence, clustering is used to remove similar evidence and Bayesian reasoning is used to capture causal relationships in the evidence;

[0018] Evidence group clustering;

[0019] Given preliminary evidence Using Embedding Model Each piece of evidence Embedded into vector v i ∈V, as shown below:

[0020]

[0021] in Represents the evidence embedding set, then, the K-means algorithm is applied to initialize the K empty clusters, expressed as Each cluster set All match an initial center u j , the set of all centers can be defined as For each v i , calculate its The Euclidean distance of all centers in v i Add to the nearest matching cluster After adding all points, update as follows Each center μ j :

[0022]

[0023] The iterative process is repeated until each cluster center set no longer changes. In the second stage, the preliminary evidence ((e1,…,e5)) is clustered into three cluster centers according to its distance from the cluster center, namely μ1, μ2 and μ3;

[0024] Identification of causal relationships in evidence;

[0025] In order to obtain the causal relationship between cluster centers, a set of cluster center pairs is constructed, denoted as Using the Bayesian method, a large language model is used to infer the causal relationship between cluster center pairs. For each cluster center pair (u a ,u b ), the prompt template is Is it possible to infer the feedback cluster centers (μ b ) is also true, answer yes or no, and then Input the large language model and query from μ a Infer μ b The probability is:

[0026]

[0027] Where LLM(·) represents the output of the large language model, I {·} is an indicator function that is equal to 1 if the condition is true. Finally, Find the probability of all center pairs in ;

[0028] Construct evidence maps;

[0029] For each cluster center (μ a ,μ b ), construct a directed cluster center graph The weight of the edge is (w(e a→b )=p(μ b |μ a )), use the maximum spanning tree algorithm MST to identify the subgraph with the largest sum of weights in the cluster center graph, which is in the form of;

[0030]

[0031] in represent All possible spanning trees in the, w(e) is the weight of edge e, and finally, the maximum spanning tree Convert to DAG, preserving the direction of edges and eliminating cycles in the graph, By vertex set and edge set E DAG Composition, which is defined as follows:

[0032]

[0033] in, and

[0034] In the second stage, the three cluster centers {μ1, μ2, μ3} are finally transformed into a DAG with the following causal relationship: μ1→μ2, μ1→μ3, μ2→μ3.

[0035] As a further improvement of the present invention, the step 3, the evidence perception prompt is optimized, specifically as follows;

[0036] Complexity metrics based on model parameter norms and outputs;

[0037] The large language model generates refined hints based on structured meta-evidence. First, the large language model generates refined hints based on Generate K candidate hints for iteration 1 from the current hint p0, denoted as Then, in iteration 2, for each Instruct the large language model to generate additional K candidates yes Next, we label the candidate hints for iteration 2 as Repeat the above process until the maximum number of iterations I is reached. max , for the i-th iteration, the candidate prompt P i It can be expressed as:

[0038]

[0039] where i∈[1,I max ]. Finally, this patent will produce the last iteration of As the final evidence perception cue;

[0040] Tips for optimization

[0041] Once the initial hint is expanded into multiple candidate hints, the optimization step selects the top b candidate hints to stay on the beam for the next iteration within T time steps. This step is simulated as a multi-armed bandit problem. In i iterations, the hint is selected at each time step t using the enhanced upper confidence bound algorithm. for:

[0042]

[0043] Where c is the exploration parameter, is the performance estimate for prompt p, is the number of queries from prompt p in the current time step t, then, in the sampled dataset Evaluate the selected prompt performance and calculate the reward Finally, after all iterations are completed, the tip b with the highest performance will be used as the final output return.

[0044] Beneficial effects:

[0045] This patent proposes a Meta-Evidence-Aware Prompt Optimization (MEPO) method for large language models. This method uses clustering and Bayesian reasoning to perform a fine analysis of the causal relationship within the evidence. First, MEPO generates preliminary evidence to describe the shortcomings of a given prompt. Then, the preliminary evidence is clustered to eliminate redundant information, and the probability of the cluster center is obtained by Bayesian inference to model the causal relationship between each pair of evidence to form a directed acyclic graph (DAG). In addition, this patent also applies the maximum spanning tree algorithm to construct the DAG, thereby enhancing the reasoning ability of LLMs. Finally, this patent edits the initial prompt based on the obtained meta-evidence and uses beam search to select the best performing prompt. Experimental results show that the MEPO method performs well in multiple tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 Optimized instructions for prompts of PO and HPO methods;

[0047] Figure 2 Schematic diagram of an overview of the MEPO method of the present invention. DETAILED DESCRIPTION

[0048] The following is a detailed description of the technical solution of the application in conjunction with the accompanying drawings. The described embodiments are only part of the embodiments involved in this patent. All non-innovative embodiments of other researchers in this field based on this embodiment are within the scope of protection of this patent.

[0049] Program Overview:

[0050] Assumptions represents a set of independent and identical training data, where x i and i Represent the i-th input and output information respectively, p0 represents the initial prompt, and all prompts p come from the language space L. The goal of the prompt optimization task is to improve the prompt p0 to produce the best prompt in is the metric function m(·).

[0051] Figure 2 The overall architecture of MEPO is shown. First, given the initial hint p0, this patent generates preliminary evidence E p , as the shortcomings of the prompt. Secondly, this patent designs a meta-evidence identification method to capture causal relationships. Finally, this patent applies the prompt optimization mechanism of evidence identification to screen out the best performing prompts.

[0052] Step 1: Generate preliminary evidence;

[0053] It is difficult for a large language model to produce the target result in one iteration when faced with a given task. To solve this problem, this patent introduces preliminary evidence As feedback to the model, it points out the shortcomings of the current prompt. Specifically, this patent converts small batches of data As a prompt p0 input into the large language model, thus obtaining the loss signal Initial evidence generation can be formalized as follows:

[0054]

[0055] like Figure 2 As shown, preliminary evidence is presented in a sequential format, for example: the prompt may not be specific enough, not clear enough, etc. To enhance the relevance of the output, this patent limits the amount of evidence by adding prompts: Give five reasons why the prompts misunderstand these examples.

[0056] Step binary evidence identification;

[0057] After generating preliminary evidence, similar evidence can be removed by clustering. Large language models tend to converge to local optimality in multiple iterations due to the lack of causal analysis of preliminary evidence. This patent uses Bayesian reasoning to capture causal relationships in evidence.

[0058] Evidence group clustering

[0059] Given preliminary evidence This patent adopts the embedded model Each piece of evidence Embedded into vector v i ∈V, as shown below:

[0060]

[0061] in Represents the evidence embedding set. Then, this patent applies the K-means algorithm to initialize K empty clusters, expressed as Each cluster set All match an initial center u j The set of all centers can be defined as For each v i , this patent calculates its The Euclidean distance of all centers in v i Add to the nearest matching cluster After adding all points, this patent is updated as follows Each center μ j :

[0062]

[0063] This patent repeats the iterative process until each cluster center set no longer changes. Figure 2 As shown in stage 2, the preliminary evidence ((e1,…,e5)) is clustered into three cluster centers, namely μ1, μ2 and μ3, according to their distances from the cluster centers.

[0064] Evidence causal relationship identification

[0065] In order to obtain the causal relationship between cluster centers, this patent constructs a set of cluster center pairs, denoted as Using the Bayesian method, this patent uses a large language model to infer the causal relationship between cluster center pairs. Specifically, for each cluster center pair (u a ,u b ), the prompt template designed by this patent is Is it possible to infer the feedback cluster centers (μ b ) is also true? Answer yes or no. Then, this patent will Input the large language model and query from μa Infer μ b The probability is:

[0066]

[0067] Where LLM(·) represents the output of the large language model, I {·} is an indicator function that is equal to 1 if the condition is true. Find the probability of all center pairs in .

[0068] Building the Evidence Graph

[0069] For each cluster center (μ a ,μ b ), this patent constructs a directed cluster center graph The weight of the edge is (w(e a→b )=p(μ b |μ a )). Use the maximum spanning tree algorithm (MST) to identify the subgraph with the largest sum of weights in the cluster center graph. Its form is

[0070]

[0071] in represent All possible spanning trees in , w(e) is the weight of edge e. Finally, this patent will be the maximum spanning tree Convert to a DAG, preserving the direction of edges and eliminating cycles in the graph. By vertex set and edge set E DAG Composition, which is defined as follows:

[0072]

[0073] in, and

[0074] like Figure 2 As shown in stage 2, the three cluster centers {μ1, μ2, μ3} are finally transformed into a DAG with the following causal relationships: μ1→μ2, μ1→μ3, μ2→μ3

[0075] Step 3: Optimize evidence-aware prompts

[0076] Complexity metrics based on model parameter norms and outputs

[0077] The large language model generates refined hints based on structured meta-evidence. Generate K candidate hints for iteration 1 from the current hint p0, denoted as Then, in iteration 2, for each This patent instructs the large language model to generate additional K candidates They are A variant of , which keeps the semantics similar but has different wording. Next, this patent marks the candidate hints of iteration 2 as This patent repeats the above process until the maximum number of iterations I is reached max (This is a hyperparameter.) For the i-th iteration, the candidate prompt P i It can be expressed as:

[0078]

[0079] where i∈[1,I max ]. Finally, this patent will produce the last iteration of As a final evidence perception cue.

[0080] Tips for optimization

[0081] Once the initial hint is expanded into multiple candidate hints, the optimization step selects the first b candidate hints to stay on the beam for the next iteration within T time steps. This patent simulates this step as a multi-armed bandit problem. Specifically, in the i iterations, this patent uses an enhanced upper confidence bound algorithm to select hints at each time step t for:

[0082]

[0083] Where c is the exploration parameter, is the performance estimate for prompt p, is the number of queries from prompt p in the current time step t. Then, in the sampled dataset Evaluate the selected prompt performance and calculate the reward Finally, after all iterations are completed, the tip b with the highest performance will be used as the final output return.

[0084] The above description is only a preferred embodiment of the present invention and does not constitute any other form of limitation to the present invention. Any modification or equivalent change made based on the technical essence of the present invention still falls within the scope of protection required by the present invention.

Claims

1. A meta-evidence-aware hint optimization method for a large language model, characterized in that: The specific steps are as follows: Step 1: Generate preliminary evidence; Step 2: Meta-evidence identification; Step 3: Optimize evidence-aware prompts.

2. The method for optimizing meta-evidence-aware prompts for a large language model according to claim 1, characterized in that: The step 1 generates preliminary evidence, which is as follows: Preliminary evidence introduced As feedback from the model, the shortcomings of the current prompts are pointed out, and the mini-batch data As a prompt p0 input into the large language model, thus obtaining the loss signal Initial evidence generation can be formalized as follows: The preliminary evidence was presented in a sequential format, and to enhance the relevance of the output, the amount of evidence was limited by adding prompts: Give five reasons why the prompts misunderstand the examples.

3. The method for optimizing meta-evidence-aware prompts for a large language model according to claim 1, characterized in that: The steps of binary evidence recognition are as follows: After generating preliminary evidence, clustering is used to remove similar evidence and Bayesian reasoning is used to capture causal relationships in the evidence; Evidence group clustering; Given preliminary evidence Using Embedding Model Each piece of evidence Embedded into vector v i ∈V, as shown below: in Represents the evidence embedding set, then, the K-means algorithm is applied to initialize the K empty clusters, expressed as Each cluster set All match an initial center u j , the set of all centers can be defined as For each v i , calculate its The Euclidean distance of all centers in v i Add to the nearest matching cluster After adding all points, update as follows Each center μ j : The iterative process is repeated until each cluster center set no longer changes. In the second stage, the preliminary evidence ((e1, ..., e5)) is clustered into three cluster centers, namely μ1, μ2 and μ3, according to its distance from the cluster center; Identification of causal relationships in evidence; In order to obtain the causal relationship between cluster centers, a set of cluster center pairs is constructed, denoted as Using the Bayesian method, a large language model is used to infer the causal relationship between cluster center pairs. For each cluster center pair (u a ,u b ), the prompt template is (μ a ), can we infer the feedback cluster center (μ b ) is also true, answer yes or no, and then Input the large language model and query from μ a Infer μ b The probability is: Where LLM(·) represents the output of the large language model, I {·} is an indicator function that is equal to 1 if the condition is true. Finally, Find the probability of all center pairs in ; Construct evidence maps; For each cluster center (μ a , μ b ), construct a directed cluster center graph The weight of the edge is (w(e a→b )=p(μ b |μ a )), use the maximum spanning tree algorithm MST to identify the subgraph with the largest sum of weights in the cluster center graph, which is in the form of; in represent All possible spanning trees in the, w(e) is the weight of edge e, and finally, the maximum spanning tree Convert to DAG, preserving the direction of edges and eliminating cycles in the graph, By vertex set and edge set E DAG Composition, which is defined as follows: in, and In the second stage, the three cluster centers {μ1, μ2, μ3} are finally transformed into a DAG with the following causal relationships: μ1→μ2, μ1→μ3, μ2→μ3.

4. The method for optimizing meta-evidence-aware prompts for a large language model according to claim 1, characterized in that: The step 3 of optimizing the evidence-aware prompt is as follows: Complexity metrics based on model parameter norms and outputs; The large language model generates refined hints based on structured meta-evidence. First, the large language model generates refined hints based on Generate K candidate hints for iteration 1 from the current hint p0, denoted as Then, in iteration 2, for each Instruct the large language model to generate additional K candidates yes Next, we label the candidate hints for iteration 2 as Repeat the above process until the maximum number of iterations I is reached. max , for the i-th iteration, the candidate prompt P i It can be expressed as: where i∈[1,I max ]. Finally, this patent will produce the last iteration of As the final evidence perception cue; Tips for optimization Once the initial hint is expanded into multiple candidate hints, the optimization step selects the top b candidate hints to stay on the beam for the next iteration within T time steps. This step is simulated as a multi-armed bandit problem. In i iterations, the hint is selected at each time step t using the enhanced upper confidence bound algorithm. for: Where c is the exploration parameter, is the performance estimate for prompt p, is the number of queries from prompt p in the current time step t, then, in the sampled dataset Evaluate the selected prompt performance and calculate the reward Finally, after all iterations are completed, the tip b with the highest performance will be used as the final output return.

Citation Information

Patent Citations

  • Method for enhancing document processing flow based on large language model

    CN118551046A

  • Service combination method based on large language model

    CN118709784A

  • Semantic relationship prediction method based on multi-section prompt optimization of large language model

    CN119294399A

  • Intelligent system and method of optimizing natural language processing models

    US20240370662A1

  • Method and apparatus for generating common-sense after-class exercise in low-resource scenario

    WO2024197740A1