Text evaluator cue word optimization method and device and storage medium

By iterative optimization of multiple design factor selection strategies, better evaluation prompt words are generated, which solves the problem of insufficient correlation between the evaluation results of text evaluators of large language models and manual evaluation in the prior art, significantly improves the evaluation performance and makes it closer to human evaluation standards.

CN120124622AActive Publication Date: 2025-06-10BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510171659.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-10
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

In the prior art, text evaluators based on large language models have shortcomings in the correlation between evaluation results and manual evaluation, and cannot completely replace manual evaluation. The optimization method is limited to the optimization of a single design factor, which fails to fully stimulate the evaluation ability of large language models.

Method used

A new selection strategy is generated by initializing the selection strategy cluster, including design factor selection strategies for multiple evaluation prompt words, and imposing perturbations on the selection strategy in each iteration. The large language model is used to generate evaluation results on the verification set with manual evaluation, calculate the correlation coefficient between the evaluation results and manual evaluation, and update the selection strategy cluster to optimize the selection strategy of multiple design factors.

Benefits of technology

It significantly improves the evaluation performance of the text evaluator, makes the evaluation results more correlate with manual evaluation, ensures that the evaluation quality is more in line with human standards, avoids local optimal solutions, and expands the search range of prompt words.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124622A_ABST
    Figure CN120124622A_ABST
Patent Text Reader

Abstract

The invention relates to a cue word optimization method and device of a text evaluator and a storage medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: firstly, initializing a selection strategy cluster, wherein the selection strategy cluster comprises a plurality of design factor selection strategies of evaluation cues; in each round of iteration, disturbance is applied to each selection strategy in the selection strategy cluster, and a new selection strategy is generated; determining an evaluation cue word based on the new selection strategy, and generating an evaluation result on the verification set with manual evaluation by using a large language model; and calculating a correlation coefficient of the evaluation result and manual evaluation, and selecting and updating the selection strategy cluster from the current selection strategy cluster and the new selection strategy based on the correlation coefficient. According to the method, the selection strategy is optimized by adopting the iterative search method guided by the heuristic function, and meanwhile, the selection strategy of a plurality of design factors in the cue word is optimized, so that the search range of the cue word is expanded, and the evaluation performance of a text evaluator is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and more particularly, to a method, device, and storage medium for optimizing prompt words of a text evaluator. Background Art

[0002] Automatic text quality evaluation has always been an important and challenging research issue in the field of natural language processing. In recent years, with the rise of large language model technology, text quality evaluation using a text evaluator based on a large language model has been widely used in various scenarios. Thanks to the powerful language understanding and instruction-following capabilities of large models, researchers have transformed the evaluation task into an instruction-following task. By inputting prompt words with task descriptions, evaluation rules, etc., and the text to be evaluated into the large language model, the large language model can output accurate evaluation results of the text to be evaluated, showing obvious advantages in terms of flexibility, interpretability, task generalization, etc., compared with automatic evaluation methods based on calculating similarity with reference answers and automatic evaluation methods based on generating probabilities or results using small pre-trained language models.

[0003] However, the text evaluator based on a large language model still has the problem of low correlation between evaluation results and human evaluation, and cannot completely replace human evaluation as an authoritative standard. Therefore, there is a series of prior arts that optimize the input prompt words of the large language model to improve its performance as an evaluator. One approach is the prompt word paradigm of "AutoCoT", that is, first requiring the large language model to generate evaluation steps, and then generating evaluation results according to the self-generated evaluation steps; another method is to require the large language to explain the reasons while generating evaluation results in the prompt words, which improves the correlation between its evaluation results and human evaluation; there is also a method of iteratively modifying and optimizing the evaluation criteria in the prompt words according to a validation set with human evaluation.

[0004] The general shortcoming of this series of work is that it is limited to optimizing a single design factor of the prompt words (such as output format or evaluation criteria). Considering that the prompt words for evaluation usually contain multiple components, and each combined part contains multiple design factors, which simultaneously affect the performance of the text evaluator based on a large language model, optimizing only a single design factor is insufficient. Therefore, how to optimize multiple design factors simultaneously to fully stimulate the evaluation ability of the large language model is a research topic worthy of study in this field. Summary of the Invention

[0005] In view of the above analysis, embodiments of the present invention aim to provide a method, device, and storage medium for optimizing prompt words of a text evaluator to optimize multiple design factors simultaneously and fully stimulate the evaluation ability of the large language model.

[0006] In the first aspect of the present application, a method for optimizing the prompting words of a text evaluator is provided, including:

[0007] Initializing a selection strategy cluster, where the selection strategy cluster includes multiple design factor selection strategies for evaluating prompting words; the evaluating prompting words include multiple design factors, each design factor has multiple options, and the design factor selection strategy is a set formed by selecting one from the options for each design factor of the evaluating prompting words;

[0008] In each round of iteration, perturb each selection strategy in the selection strategy cluster to generate a new selection strategy;

[0009] Determine the evaluating prompting words based on the new selection strategy, and use a large language model to generate evaluation results on a validation set with human evaluation;

[0010] Calculate the correlation coefficient between the evaluation result and the human evaluation to measure the evaluation performance of the new selection strategy;

[0011] Based on the correlation coefficient, select and update the selection strategies in the selection strategy cluster from the current selection strategy cluster and the new selection strategies.

[0012] Optionally, the step of perturbing each selection strategy in the selection strategy cluster in each round of iteration to generate a new selection strategy includes:

[0013] In each round of iteration, with a probability of 1 - ρ, select a local perturbation method, and with a probability of ρ, select a global perturbation method to perturb each selection strategy in the selection strategy cluster to generate a new selection strategy.

[0014] Optionally, the execution steps of the local perturbation method include:

[0015] Select an adjacent selection strategy set T that differs from the current selection strategy in only one design factor of the evaluating prompting words adj ;

[0016] Calculate the difference between the sum of the advantages of the current selection strategy and each selection strategy in the adjacent selection strategy set T adj in all design factors

[0017] Calculate the selection probabilities of each selection strategy in the adjacent selection strategy set T adj through a softmax function with temperature control, and select a new selection strategy from the adjacent selection strategy set T adj based on the selection probabilities

[0018] Among them, the selection probability is calculated by the following formula:

[0019]

[0020] where A ij is the advantage of the j-th option f i of the i-th design factor F ij , is the advantage of the current option of the i-th design factor in the current selection strategy , t is the current search step, and M ij is the number of occurrences of the option f ij during the entire search process, and λ and τ are hyperparameters.

[0021] Optionally, the execution steps of the global perturbation method include:

[0022] Calculate the sum of the advantages of each design factor of all un-searched selection strategies in the current search process;

[0023] Select the selection strategy with the largest sum of advantages as the new selection strategy;

[0024] where the new selection strategy T max is determined by the following formula:

[0025]

[0026] where {F 1 , F 2 ,..., F n} represents the design factor, represents the advantage calculated for the design factor F i during the current search process.

[0027] Optionally, it further includes:

[0028] Calculate the advantages of each option in each design factor:

[0029] where the advantage A i of the j-th option f ij of the i-th design factor F ij is defined as follows:

[0030]

[0031] where represents the evaluation performance of the prompt word on the validation set, m i is the number of options of F i ; For the expectation calculation of other factors except F among all design factors i except;

[0032] Wherein, the advantage is used to adjust the probability of the selection strategy in the search process, so that the option with a higher advantage value has a higher probability of being selected in subsequent iterations than the option with a lower advantage value.

[0033] Optionally, after generating a new selection strategy using the local perturbation method, it further includes:

[0034] Calculate the evaluation performance gain brought by the j-th option f i in the design factor F ij : is the evaluation performance of the new selection strategy, is the evaluation performance of the current selection strategy ; is the advantage of the current option of the i-th design factor in the current selection strategy ;

[0035] Update the advantage of the j-th option f i in the design factor F ij :

[0036]

[0037] Perform normalization processing so that the sum of the advantages of all options within the same design factor F i is 0,

[0038] wherein, N ij is the number of times f ij is explored during the search process.

[0039] Optionally, the selection of updating the selection strategy in the selection strategy cluster from the current selection strategy cluster and the new selection strategy based on the correlation coefficient includes:

[0040] After the search steps reach the maximum search steps, the iterative search terminates, and the final selection strategy cluster is output.

[0041] Optionally, the design factors of the evaluation prompt words include: scoring range, evaluation example, evaluation criterion, reference answer, output format, evaluation steps, reference questions, placement order.

[0042] Optionally, the selection of updating the selection strategy in the selection strategy cluster from the current selection strategy cluster and the new selection strategy based on the correlation coefficient includes:

[0043] For the data set For evaluation dimension a, the large language model being evaluated is M, and the evaluation prompt used On the i-th sample d of the dataset i The evaluation result is The human evaluation result corresponding to the sample is By searching for the evaluation prompt To maximize the correlation between the large language model evaluation result and the human evaluation result:

[0044]

[0045] Among them, the evaluation prompt Includes n design factors {F 1 , F 2 , …, F n}, and the i-th design factor F i Has m i Options corr(·) is the correlation coefficient, Represents the evaluation performance of the evaluation prompt On the evaluation dimension a of the dataset D.

[0046] In the second aspect of the present application, a prompt optimization device for a text evaluator is provided, including a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, it implements the prompt optimization method for the text evaluator according to any one of the above.

[0047] In the third aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the prompt optimization method for the text evaluator according to any one of the above.

[0048] The method for optimizing the prompt words of the text evaluator provided by this application. The evaluation prompt words include multiple design factors. By recording a number of currently searched optimal design factor selection strategy clusters, in each round of iteration, perturbations are applied to the selection strategies in the cluster to obtain new selection strategies. For each selection strategy, after using it to determine the evaluation prompt words, the large language model is used to generate evaluation results on a validation set with human evaluation. The performance is measured by calculating the correlation coefficient between the generated evaluation results and the human evaluation, and based on this, the selection strategies retained in the cluster are updated. It can be seen that this application uses an iterative search method guided by a heuristic function to optimize the selection strategies of each design factor, automatically discovers the optimal prompt word combination, and improves the search effect. At the same time, this application is not limited to optimizing a single design factor of the prompt words, but optimizes the selection strategies of multiple design factors in the prompt words, expands the search scope of the prompt words, makes the correlation between the evaluation results and the human evaluation higher, ensures that the evaluation quality is more in line with human standards, and significantly improves the evaluation performance of the text evaluator.

[0049] In addition, this application also provides a device and a storage medium for optimizing the prompt words of the text evaluator with the above technical effects. Brief Description of the Drawings

[0050] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of this specification. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0051] Figure 1 It is a flowchart of a specific implementation manner of the method for optimizing the prompt words of the text evaluator provided by this application;

[0052] Figure 2 It is a flowchart of the execution steps of the local perturbation method adopted by this application;

[0053] Figure 3 It is a flowchart of the execution steps of the global perturbation method adopted by this application;

[0054] Figure 4 It is a flowchart of another specific implementation manner of the method for optimizing the prompt words of the text evaluator provided by this application;

[0055] Figure 5 It is a schematic diagram of another specific implementation manner of the method for optimizing the prompt words of the text evaluator provided by this application;

[0056] Figure 6Schematic diagram of the performance results of the method of the present invention (HPSS) and the baseline method (CloserLook+ICL) under different validation set sizes Figure 1 ;

[0057] Figure 7 Schematic diagram of the performance results of the method of the present invention (HPSS) and the baseline method (CloserLook+ICL) under different validation set sizes Figure 2 ;

[0058] Figure 8 Block diagram of the structure of the prompt optimization device of the text evaluator provided by the present application Detailed implementation manners

[0059] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. It should be noted that, without conflict, the implementation manners and features in the implementation manners in the present disclosure may be combined with each other, separated, interchanged, and / or rearranged. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application

[0060] The terms used herein are for the purpose of describing specific embodiments and are not intended to be limiting. As used herein, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are also intended to include the plural forms. In addition, when the terms "comprise" and / or "include" and their variants are used in this specification, it is indicated that there are the stated features, wholes, steps, operations, components, assemblies, and / or groups thereof, but do not exclude the existence or addition of one or more other features, wholes, steps, operations, components, assemblies, and / or groups thereof. It should also be noted that, as used herein, the terms "substantially", "about", and other similar terms are used as approximate terms and not as degree terms, so they are used to explain the inherent deviations of the measured values, calculated values, and / or provided values that those of ordinary skill in the art will recognize

[0061] The flowchart of a specific implementation manner of the prompt optimization method of the text evaluator provided by the present application is as shown in Figure 1 and specifically includes

[0062] S101: Initialize the selection strategy cluster

[0063] The selection strategy cluster includes multiple design factor selection strategies for evaluating prompt words; the evaluation prompt words include multiple design factors, and each design factor has multiple options. The design factor selection strategy is a set formed by selecting one from the options for each design factor of the evaluation prompt words.

[0064] Among them, the evaluation prompt words altogether include n design factors {F 1 , F 2 , …, F n}, and the i-th design factor F i has m i options for selection, that is Then the evaluation prompt words can be determined by the values of {F 1 , F 2 , …, F n}. The selection strategy of the design factor refers to a set of values of {F 1 , F 2 , …, F n}, that is, each design factor selects one corresponding option.

[0065] The design factor selection strategy cluster can be used to record several currently searched design factor selection strategies with optimal performance. The specific initialization process can be: selecting existing artificial design prompt words as the starting point, making different modifications to each design factor of the prompt words to generate multiple prompt word combinations (i.e., design factor selection strategies), testing the performance of the above prompt words on the validation set, and selecting the k prompt word combinations with the best effects as the initial design factor selection strategy cluster.

[0066] S102: In each round of iteration, perturb each selection strategy in the selection strategy cluster to generate a new selection strategy.

[0067] Specifically, in each round of iteration of the present application, with a probability of 1 - ρ, the local perturbation method is selected, and with a probability of ρ, the global perturbation method is selected to perturb each selection strategy in the selection strategy cluster to generate a new selection strategy.

[0068] Among them, the local perturbation method is to modify a certain design factor of the current selection strategy to generate an adjacent selection strategy to ensure local optimization. The global perturbation method is to calculate the total advantage of all un-searched options, select the optimal strategy, expand the search scope, and prevent falling into local optimum.

[0069] S103: Determine the evaluation prompt words based on the new selection strategy, and use the large language model to generate evaluation results on the validation set with human evaluation.

[0070] Run different prompts on the validation set with human evaluation using a large language model to generate evaluation results.

[0071] S104: Calculate the correlation coefficient between the evaluation results and the human evaluation to measure the evaluation performance of the new selection strategy.

[0072] Specifically, the large language model (LLM) is a pre-trained natural language processing model, and the performance is measured based on the Spearman correlation coefficient or Pearson correlation coefficient between the generated evaluation results and the human evaluation results.

[0073] S105: Based on the correlation coefficient, select and update the selection strategies in the current selection strategy cluster from the current selection strategy cluster and the new selection strategy.

[0074] According to the calculated correlation coefficient, retain the best-performing prompt combination and adjust the search direction to make the model gradually tend to the optimal prompt.

[0075] Specifically, set an upper limit on the cluster size, for example, retain at most k optimal selection strategies. Among the strategies generated by the current cluster + new perturbations, select the k strategies with the highest correlation coefficient and eliminate the inferior strategies. After reaching the given number of iteration rounds, finally select the selection strategy with the highest correlation coefficient in the selection strategy cluster as the optimization result.

[0076] The method for optimizing the prompts of the text evaluator provided by this application, the evaluation prompts include multiple design factors. By recording several clusters of selection strategies of the currently searched performance-optimal design factors, in each iteration, perturbations are applied to the selection strategies in the cluster to obtain new selection strategies. For each selection strategy, after using it to determine the evaluation prompts, use a large language model to generate evaluation results on a validation set with human evaluation, and measure its performance by calculating the correlation coefficient between the generated evaluation results and the human evaluation, and accordingly update the selection strategies retained in the cluster. It can be seen that this application uses an iterative search method to optimize the selection strategies of each design factor, automatically discovers the optimal prompt combination, and improves the search effect. At the same time, this application is not limited to optimizing a single design factor of the prompts, but optimizes the selection strategies of multiple design factors in the prompts, expands the search scope of the prompts, makes the correlation between the evaluation results and the human evaluation higher, ensures that the evaluation quality is more in line with human standards, and significantly improves the evaluation performance of the text evaluator.

[0077] In the above embodiment, as Figure 2 shown, the execution steps of the local perturbation method include the following steps:

[0078] S201: Select the one related to the current selection strategy The adjacent selection strategy set T with different design factors having only one evaluation prompt word adj .

[0079] S202: Calculate the current selection strategy and the difference in the total advantages of each selection strategy in the adjacent selection strategy set T adj on all design factors

[0080] S203: Calculate the selection probabilities of each selection strategy in the adjacent selection strategy set T adj through the softmax function with temperature control, and select a new selection strategy from the adjacent selection strategy set T adj based on the selection probabilities

[0081] wherein, the selection probability is calculated by the following formula:

[0082]

[0083] wherein, A ij is the advantage of the jth option f i of the ith design factor F ij , is the advantage of the current option of the ith design factor in the current selection strategy , t is the current search step, M ij is the number of occurrences of the option f ij during the entire search process, and λ, τ are hyperparameters

[0084] The local perturbation method searches for a better selection strategy near the currently searched strategies and gradually optimizes the prompt words. Modifying only one design factor helps to understand the key factors affecting the evaluation performance

[0085] In the above embodiment, as Figure 3 shown, the execution steps of adopting the global perturbation method include the following steps:

[0086] S301: Calculate the total advantages of each design factor of all unsearched selection strategies in the current search process

[0087] S302: Select the selection strategy with the largest total advantage as the new selection strategy

[0088] wherein, the new selection strategy T max is determined by the following formula:

[0089]

[0090] wherein, {F1 , F 2 , …, F n} represents design factors represents design factor F i The advantages calculated during the current search process

[0091] Directly select the strategy with the largest total advantage among all current factors in the way of global perturbation, so as to jump out of the local optimum and quickly find the possible global optimum solution

[0092] By combining local perturbation and global perturbation, this application can not only ensure the stability of local optimization, but also avoid falling into the local optimum, and finally find the optimal prompt word design

[0093] Based on any of the above embodiments, in the prompt word optimization method of the text evaluator provided by this application, the design factors for evaluating prompt words specifically include: scoring range, evaluation examples, evaluation criteria, reference answers, output format, evaluation steps, reference questions, and placement order

[0094] Correspondingly, the definitions and options of each design factor are as follows

[0095] (1) Scoring range: The score range required for the large language model to output. It includes 5 options: 1 - 3, 1 - 5, 1 - 10, 1 - 50, and 1 - 100

[0096] (2) Evaluation examples: Evaluation examples with human evaluation. It includes 4 options: no evaluation examples provided, 3 evaluation examples provided, 5 evaluation examples provided, and 10 evaluation examples provided

[0097] (3) Evaluation criteria: The definition of the evaluated dimension and the corresponding relationship between scoring and text quality. It includes 3 options: no evaluation criteria provided, evaluation criteria written manually provided, evaluation criteria generated automatically by the large language model provided

[0098] (4) Reference answers: The reference answers to the instructions in the samples to be evaluated. It includes 3 options: no reference answers provided, reference answers generated automatically by the large language model provided, require the large language model to generate reference answers first and then generate evaluation results

[0099] (5) Output format: The output format of the evaluation results. It includes 3 options: directly output the score, output the evaluation explanation first and then the score, output the score first and then the evaluation explanation

[0100] (6) Evaluation steps: The evaluation steps for the evaluated dimension. It includes 2 options: no evaluation steps provided, evaluation steps provided

[0101] (7) Reference question: The reference evaluation question for the evaluation samples. It includes 2 options: Do not provide a reference question, Provide a reference question.

[0102] (8) Placement order: The above 7 design factors belong to 3 main components of the evaluation prompt: 1) Task description: including the definition of the evaluation task and requirements for the output format (including two design factors: scoring range and output format); 2) Evaluation rules: including the evaluation criteria and steps (including two design factors: evaluation criteria and evaluation steps); 3) Input content: including the samples to be evaluated and additional information related to the samples (including three design factors: evaluation examples, reference answers, and reference questions). The placement order refers to the placement order of these 3 components in the evaluation prompt. It includes 6 options, that is, the full permutation of these 3 components.

[0103] Compared with the prior art which is limited to optimizing a single design factor (such as output format or evaluation criteria) among them, resulting in the defect of insufficient optimization of the prompt. This application proposes to use not limited to the above 8 key evaluation prompt design factors, and use an iterative search method to optimize the selection strategy of each design factor.

[0104] The optimization objective of this application is formally defined as follows: For the dataset and the evaluation dimension a, the large language model for evaluation is M, and the evaluation prompt on the i-th sample d i of the dataset, the evaluation result is The corresponding human evaluation result of the sample is By searching for the evaluation prompt to maximize the correlation between the large language model evaluation result and the human evaluation result:

[0105]

[0106] Among them, the evaluation prompt includes n design factors {F 1 , F 2 ,..., F n}, and the i-th design factor F i has m i options corr(·) is the correlation coefficient, represents the evaluation performance of the evaluation prompt on the evaluation dimension a of the dataset D. During the automatic optimization process, this application uses a validation set with human evaluation to test the performance of the evaluation prompt . Without loss of generality, this application abbreviates as to represent the evaluation prompt Performance on specific evaluation dimensions of the validation set.

[0107] In the embodiments provided by the present application, by recording a cluster containing several currently searched design factor selection strategies with the best performance, in each round of iteration, perturbations are applied to the selection strategies in the cluster to obtain new selection strategies. For each selection strategy, after using it to determine the evaluation prompt words, the large language model is used to generate evaluation results on the validation set, and the performance is measured by calculating the correlation coefficient between the generated evaluation results and the manual evaluation, and accordingly, the selection strategies retained in the cluster are updated. In this step of applying perturbations, the present application introduces a heuristic function to calculate the selection probability of each option in each design factor, giving a greater selection probability to the option that can bring a greater gain to the evaluation performance of the large language model, and improving the effectiveness of the search. The selection strategy cluster is updated based on the heuristic function, where the heuristic function is used to calculate the advantages of each design factor option and adjust the search direction according to the advantages to improve the optimization effect.

[0108] Specifically, the heuristic function is defined as the advantage of each option in each design factor, that is, the mathematical expectation of the performance gain of the evaluation prompt words when this option is selected compared to randomly selecting an option in this design factor. The advantage A of the j-th option f i in the i-th factor F ij is calculated according to the following formula: ij The following formula:

[0109]

[0110] Among them, represents the evaluation performance of the prompt words on the validation set, and m i is the number of options of F i ; is the expectation calculation of other factors except F i in all design factors.

[0111] This advantage is used to adjust the probability of the selection strategy during the search process, so that the option with a higher advantage value has a higher probability of being selected in subsequent iterations than the option with a lower advantage value.

[0112] The present application also provides another specific implementation manner of the prompt word optimization method of the text evaluator. Referring to Figure 4 the flowchart shown in Figure 5 and the schematic diagram shown in

[0113] S401: Initialize the selection strategy cluster and calculate the initial advantages of each option.

[0114] As a specific implementation, this application starts with commonly used evaluation prompts, such as those used in MT-Bench, and modifies the options of each design factor individually to construct several selection strategies. Subsequently, the performance of these selection strategies is tested on the validation set, and the k selection strategies with the best performance are selected as the initial clusters to enter the next step. At the same time, according to the definition of the above advantages, the initial advantages of each option in each design factor are calculated.

[0115] As Figure 5 shown, the design factor selection strategy of MT-Bench includes three design factor examples:

[0116] Score range (Factor 1): 1 - 3, 1 - 5, 1 - 10, 1 - 50, 1 - 100;

[0117] Score examples (Factor 2): No examples, 3 examples, 5 examples, 10 examples;

[0118] Scoring criteria (Factor 3): None, Manually written, Model generated.

[0119] Combine different design factors to form an initial selection strategy cluster including multiple selection strategies. S402: In each round of iterative search, for each selection strategy in the selection strategy cluster perform g perturbations to generate new selection strategies.

[0120] Specifically, when perturbing, local perturbation and global perturbation can be combined to avoid local optimality.

[0121] Among them, local perturbation is also called the "exploration" perturbation method, which selects the selection strategies adjacent to the current selection strategy , that is, the set T of selection strategies that differ from the current selection strategy by only 1 factor adj . This application calculates the difference between the sum of the advantages of all factors of the current selection strategy and each selection strategy in T adj , and calculates the selection probability of each selection strategy in T adj through a softmax function with temperature control.

[0122] Additionally, this application also introduces an exploration coefficient to encourage comprehensive exploration of all options. Formally, assume the current selection strategy is and the option of its i-th factor F i is T adj in the selection strategy where the value of the i-th factor F i is the same as is different and is f ij then The selection probability of is calculated according to the following formula:

[0123]

[0124] where t is the current search step, and M ij is the number of occurrences of option f ij during the entire search process, and λ, τ are hyperparameters.

[0125] However, a major limitation of the "exploration" perturbation method is that its search range is limited to the vicinity of the searched selection strategies. To expand the search range and avoid getting stuck in local optimal points, the present application also uses the "exploitation" perturbation method (i.e., the global perturbation method) in addition to "exploration", that is, directly selects the selection strategy with the greatest advantage among all factors that have not been searched currently:

[0126]

[0127] In each perturbation, the present invention has a probability of ρ to select the "exploitation" perturbation method and a probability of 1 - ρ to select the "exploration" perturbation method. After reaching the maximum search step, the present application selects the selection strategy that has been searched currently and has the best performance on the validation set as the final search optimization result. As a specific example, ρ can take a value of 0.2.

[0128] S403: Update the cluster that records the currently k best-performing selection strategies according to the performance of the new selection strategy on the validation set.

[0129] Determine the evaluation prompt words based on the new selection strategy, and use the large language model to generate evaluation results on the validation set with human evaluation; calculate the correlation coefficient between the evaluation result and the human evaluation to measure the evaluation performance of the new selection strategy; based on the correlation coefficient, select and update the selection strategies in the current selection strategy cluster from the current selection strategy cluster and the new selection strategy.

[0130] S404: Update the advantage function.

[0131] After selecting a new selection strategy the present application tests its performance on the validation set, and accordingly calculates the performance gain brought by modifying the option of factor F i to f ij After that, if the new selection strategy is obtained by the local perturbation method, the present application updates the advantage A ij of f ij by the way of taking the moving average with the previous advantage, and factor F iThe mean of the advantages of all options is normalized to 0 to meet the performance requirement of an expected sum of 0. Formally, the advantage A is updated as follows, where N ij is the number of times option f ij has been explored.

[0132]

[0133] By dynamically adjusting the advantage values of each option, the search process is ensured to be stably optimized. After the update, options with low contribution values are weakened, and options with high contribution values are strengthened, which can improve the search effect.

[0134] This application uses two large language models, Qwen2.5-14B-Instruct and GPT-4o-mini, to verify the optimization effect on 5 natural language generation evaluation datasets, namely the Summeval dataset for text summarization generation, the Topical-Chat dataset for dialogue generation, the SFHOT and SFRES datasets for data-to-text generation, and the HANNA dataset for story generation. Each dataset is randomly divided into a 50% validation set and a 50% test set, and the Spearman correlation coefficient between the evaluation results and human evaluations is used as the performance measurement metric for the evaluator.

[0135] Three types of baselines were selected for the experiment: 1) Non-large language model evaluators: including BLEU-4, BERTScore, and UNIEVAL; 2) Human-prompt-optimized large language model-based text evaluators: including the starting point optimized in this application, i.e., the evaluation prompts used in MT-Bench, as well as the prompts used in the two works of G-EVAL and CloserLook, and the prompt CloserLook+ICL formed by adding 3 scoring examples to the CloserLook prompt; 3) Other automatic prompt optimization methods for large language model-based text evaluators: including APE, OPRO, and Greedy. For the two prompts of G-EVAL and CloserLook, the same decoding method as the original work is adopted, that is, the large model is required to generate 20 times, the sampling temperature is set to 1.0, and the average score is calculated as the final evaluation result; for other prompts, greedy search is used for decoding to ensure reproducibility. For all automatic prompt optimization methods, the maximum number of search times is set to 71.

[0136] In the selection of hyperparameters for this application, the strategy cluster size k = 5, the number of perturbations g for each selection strategy in the cluster = 2, the probability ρ of choosing the "exploitation" perturbation method each time = 0.2, the temperature τ when calculating the sampling probability using the softmax function = 5, and the weight λ of the exploration coefficient = 4.

[0137] The results on 5 datasets are shown in Table 1 and Table 2. From the content in the tables, it can be seen that the method of this application significantly improves the evaluation performance of the text evaluator based on large language models: compared with the optimized starting point, that is, the prompt used in MT-Bench, under the same number of generations, an average relative improvement of 29.4% can be achieved in the correlation with human evaluation, significantly exceeding other baseline methods; under 1 / 20 of the number of generations, this application can still achieve evaluation performance significantly better than that of the text evaluator based on large language models optimized by human prompts.

[0138] Table 1

[0139]

[0140] Table 2

[0141]

[0142]

[0143] In addition, this application also tested the performance of this application method and the text evaluator based on large language models optimized by the best-performing human prompts (CloserLook+ICL) with Qwen2.5-14B-Instruct as the evaluation model on the Summeval and Topical-Chat datasets when using different validation set sizes. As Figure 6 and Figure 7 shown, the abscissa in the attached figure is the validation set size, with the unit of %; the ordinate is the Spearman correlation coefficient. The experiment found that the method of this application (shown as HPSS in the figure) still outperforms the baseline method when the validation set size only accounts for 10% of the total dataset, proving that this application still has good performance when the validation set with human evaluation is small.

[0144] In addition, this application also provides a device for optimizing the prompts of a text evaluator. As Figure 8 shown in the structural block diagram of the device for optimizing the prompts of the text evaluator provided by this application, this device includes a memory 81 and a processor 82. The memory 81 stores a computer program, and when the computer program is executed by the processor 82, it realizes the method for optimizing the prompts of the text evaluator according to any one of the above.

[0145] In addition, this application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the method for optimizing the prompts of the text evaluator according to any one of the above.

[0146] A computer-readable storage medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0147] Those skilled in the art should also be able to further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of function in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans may use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0148] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be implemented in hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0149] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of this application. It should be understood that the above description is only the specific embodiments of this application and is not used to limit the protection scope of this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included in the protection scope of this application.

Claims

1. A prompt word optimization method for a text evaluator, characterized in that: include: Initializing a selection strategy cluster, wherein the selection strategy cluster includes a plurality of design factor selection strategies for evaluation prompt words; The evaluation prompt word includes a plurality of design factors, each design factor has a plurality of options, and the design factor selection strategy is to select a set of one of the options for each design factor of the evaluation prompt word; In each round of iteration, a perturbation is applied to each selection strategy in the selection strategy cluster to generate a new selection strategy; Determine evaluation prompt words based on the new selection strategy, and use the large language model to generate evaluation results on the validation set with manual evaluation; Calculating the correlation coefficient between the evaluation result and the manual evaluation to measure the evaluation performance of the new selection strategy; Based on the correlation coefficient, a selection strategy in the selection strategy cluster is updated by selecting from the current selection strategy cluster and the new selection strategy.

2. The prompt word optimization method of the text evaluator according to claim 3 is characterized in that: In each round of iteration, a perturbation is applied to each selection strategy in the selection strategy cluster to generate a new selection strategy, including: In each round of iteration, a local perturbation method is selected with a probability of 1-ρ, and a global perturbation method is selected with a probability of ρ. Perturbations are applied to each selection strategy in the selection strategy cluster to generate new selection strategies.

3. The prompt word optimization method of the text evaluator according to claim 2, characterized in that: The execution steps of the local disturbance mode include: Selection and current selection strategies A set of adjacent selection strategies T with only one evaluation prompt word with different design factors adj ; Calculate the current selection strategy and the adjacent selection strategy set T adj The difference between the sum of the advantages of each selection strategy on all design factors The adjacent selection strategy set T is calculated by the softmax function with temperature control adj The selection probability of each selection strategy in , and based on the selection probability, select the adjacent selection strategy set T adj Choose a new selection strategy Among them, the selection probability Calculated by the following formula: Among them, A ij is the i-th design factor F i The jth option f ij Advantages, is the current option for the i-th design factor in the current selection strategy advantage, t is the current search step number, M ij It is option f ij The number of occurrences in the entire search process, λ, τ are hyperparameters.

4. The prompt word optimization method of the text evaluator according to claim 3 is characterized in that: The execution steps of the global perturbation method include: Calculate the sum of advantages of each design factor for all selection strategies that have not been searched in the current search process; Select the selection strategy with the largest sum of advantages as the new selection strategy; Among them, the new selection strategy T max Determined by the following formula: Among them, {F1,F2,…,F n } represents the design factor, Indicates the design factor F i The advantage calculated during the current search.

5. The prompt word optimization method of the text evaluator according to claim 4, characterized in that: Also includes: Calculate the advantage of each option for each design factor: Among them, the i-th design factor F i The jth option f ij Advantages of A ij Defined as follows: in, Indicates prompt words Evaluation performance on the validation set, m i F i The number of options; To remove F from all design factors i Expected calculation of other factors; The advantage is used to adjust the probability of selecting a strategy during the search process so that options with higher advantage values ​​are more likely to be selected in subsequent iterations than options with lower advantage values.

6. The prompt word optimization method of the text evaluator according to claim 5, characterized in that: After using the local perturbation method to generate a new selection strategy, the method further includes: Calculate the design factor F i The jth option f ij Evaluation performance gain: To evaluate the performance of the new selection strategy, Select strategy for the current Evaluation performance; is the current option for the i-th design factor in the current selection strategy Advantages: Update design factor F i The jth option f ij Advantages: Normalization is performed so that the same design factor F i The sum of the advantages of all options in is 0, Among them, N ij f ij The number of times it was explored during the search.

7. The prompt word optimization method of a text evaluator according to any one of claims 1 to 6, characterized in that: The selecting and updating the selection strategy in the selection strategy cluster from the current selection strategy cluster and the new selection strategy based on the correlation coefficient includes: After the number of search steps reaches the maximum number of search steps, the iterative search terminates and the final selection strategy cluster is output.

8. The prompt word optimization method of a text evaluator according to any one of claims 1 to 6, characterized in that: The design factors of the evaluation prompt words include: scoring range, evaluation examples, evaluation criteria, reference answers, output format, evaluation steps, reference questions, and placement order.

9. The prompt word optimization method of a text evaluator according to any one of claims 1 to 6, characterized in that: The selecting and updating the selection strategy in the selection strategy cluster from the current selection strategy cluster and the new selection strategy based on the correlation coefficient includes: For the dataset and evaluation dimension a, the large language model of evaluation is M, and the evaluation prompt words used are In the i-th sample d of the data set i The evaluation results on The manual evaluation results of the corresponding samples are By searching for evaluation prompt words To maximize the correlation between the large language model evaluation results and the manual evaluation results: Among them, the evaluation prompt words Including n design factors {F1, F2, ..., F n }, the i-th design factor F i There are i Options corr(·) is the correlation coefficient, Representative evaluation prompt words Evaluation performance on the evaluation dimension α of dataset D.

10. A prompt word optimization device for a text evaluator, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the prompt word optimization method of the text evaluator according to any one of claims 1 to 9 is implemented.

11. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the prompt word optimization method of the text evaluator according to any one of claims 1-9 is implemented.

Citation Information

Patent Citations

  • Prompt word updating method, device and system based on multi-dimensional large language model

    CN118132716A

  • Low-complexity safety method for enhancing robustness of large model

    CN118885404A

  • Self-reflection cue word optimization method and system based on large language model

    CN118966208A

  • Dynamic cue word example recall method and system based on reinforcement learning

    CN118966209A

  • Private data protection method and device, storage medium and computer program product

    CN119323051A