Large language model evaluation method based on multi-identity play for power field
By employing multiple role-playing and dynamic weight adjustment methods, the issues of professionalism and consistency in the evaluation of large language models in the power sector were resolved, improving the accuracy of the evaluation and its consistency with human preferences, thus achieving more accurate and fair evaluation results.
Patent Information
- Application Number
- CN202411143740.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2026-03-03
AI Technical Summary
Existing large language model evaluation methods in the power sector lack professionalism, traditional evaluation methods are difficult to cover the needs of the power industry, and multi-domain evaluation is costly and prone to leakage risks, and model evaluation results are inconsistent with human judgment.
By employing a multi-identity role-playing and dynamic weight adjustment approach, a multi-identity evaluation model is constructed by setting identities based on downstream tasks, rewriting the original dataset, and adjusting model weights, thereby optimizing the evaluation results to approximate human preferences.
This improves the accuracy and consistency with human preferences in the evaluation of large language models in the power sector, reduces bias in the model evaluation process, and achieves more accurate and fair evaluation results.
Smart Images

Figure CN121599085A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the problem of evaluating natural language generation tasks for large language models in the power industry, and to a method for evaluating large language models based on multiple role-playing in the power industry. Background Technology
[0002] With the rapid development of generative AI, the capabilities of Large Language Models (LLMs) have been gradually enhanced, and they are increasingly being used in a range of Natural Language Generation (NLG) tasks, such as text summarization, text translation, and dialogue generation. In the power industry, LLMs can also solve system management problems, overcoming the limitations of traditional applications such as system operation and fault detection. With the emergence of various LLMs, there is an increasing need to evaluate their capabilities in handling downstream tasks.
[0003] Traditional NLG evaluation methods use matching metrics such as BLEU or ROUGE to calculate the similarity between the language model's output and the standard answer. While these methods can measure the quality of the generated text to some extent, they overlook some text that has low literal relevance to the standard answer but is semantically relevant.
[0004] Some large language model evaluation benchmarks that emerged after the launch of ChatGP, such as BigBench and LongBench, cover multiple NLG domains and test the performance of language models by evaluating their generative capabilities in these domains. While these benchmarks are multi-domain, they still struggle to cover a wider range of downstream tasks. Furthermore, although they avoid ignoring semantic similarity, building such benchmarks is costly, and there is a risk of benchmark leakage over time.
[0005] For specialized fields, such as the power industry, there is a greater lack of specialized evaluation methods. Most existing evaluations are limited to general NLP capabilities, such as text generation and text summarization, and cannot cover the needs of the power industry. Using manually constructed evaluation methods to assess the capabilities of large language models specific to a particular professional domain requires significantly higher costs, including guidance from domain experts and the construction of domain-specific test questions. Therefore, it is necessary to explore alternative evaluation methods besides manually constructed evaluation sets.
[0006] Recent researchers have begun exploring the use of more advanced large language models to evaluate text generated by language models. Because large language models possess strong generalization capabilities, they are better able to assess semantic accuracy compared to traditional evaluation methods. Specifically, when using large language models for NLG evaluation, a good evaluation prompt is first set for the model. This prompt can specify details for model evaluation or employ in-context learning to provide the model with relevant information during evaluation.
[0007] Using large language models for evaluation places high demands on the capabilities of the models themselves. If the model has a small number of parameters, low evaluation accuracy may occur when evaluating a single model. One solution is to use multiple models for joint evaluation, but the relationships between the models and the prompt settings need to be explored. On the other hand, the evaluation of large language models reveals that the models themselves may be biased towards specific text content, thus the evaluation results may not accurately align with human judgment. Summary of the Invention
[0008] To overcome the shortcomings of the existing technologies, this invention proposes a method for constructing a multi-identity evaluation model using two approaches: multi-identity role-playing and dynamic weight adjustment. The aim is to eliminate potential biases in the model's treatment of the evaluation text using multi-identity role-playing, while simultaneously using preference data to fine-tune the weights to optimize the overall evaluation results, making them closer to human judgment. To this end, this invention adopts the following technical solution:
[0009] A method for evaluating large language models based on multiple identity roles in the power industry is characterized by including identity retrieval based on downstream tasks;
[0010] Rewriting the original dataset based on a large language model;
[0011] Adjust the weights of different identity models based on the preference dataset;
[0012] Evaluation of large language models based on linear weighting of multi-identity models.
[0013] Furthermore, each large language model used for evaluation (such as GPT-4, DeepSeek-V2, Llama3-70B, Tigerbot-70B, etc.) is assigned a different identity, and the identity retrieval includes:
[0014] Based on the text description of the downstream task, a large language model is used to generate professional identities that are related to but different from the downstream task, each with a certain background in the power industry. The best-matching identity is then assigned to each evaluation model for role-playing.
[0015] Furthermore, the original dataset is rewritten using the large language model to obtain preference data. The original dataset refers to an expert knowledge question-and-answer dataset in the power industry. Each sample in the dataset contains a question and a reference answer (positive sample) and an incorrect answer (negative sample) in natural language form. The rewritten original dataset based on the large language model includes:
[0016] Rewrite the samples included in the original dataset;
[0017] Positive samples are rewritten using semantically similar but content-diverse methods to construct positive samples under human cognition, but the model may produce rewritten data with different evaluation results; keywords in positive samples are reversed to construct negative samples under human cognition, but the model may produce rewritten data with different evaluation results.
[0018] Negative samples are rewritten using semantically similar but content-diverse methods to construct negative samples under human cognition, but the model may produce rewritten data with different evaluation results; keywords in negative samples are reversed to construct positive samples under human cognition, but the model may produce rewritten data with different evaluation results.
[0019] Combining positive and negative samples yields a training dataset that can represent human cognitive preferences.
[0020] Furthermore, the method of adjusting the weights of different identity models based on the preference dataset, using positive and negative samples to evaluate different identity models, and adjusting the relative weights, includes:
[0021] Initialize the same weights for each evaluation model;
[0022] Each training data point is evaluated using an evaluation model with different identities. The model scores the data based on its assigned identity, thereby offsetting any bias that the model itself may contain.
[0023] The model's score is multiplied by its weight and then compared with the score of the dataset itself. The weight of the model whose score is closer to the ground truth is increased, and the weight of the model whose score deviates from the ground truth is decreased.
[0024] After optimization using multiple training datasets, the model identity and corresponding weights are obtained by dynamically adjusting them based on different tasks.
[0025] Furthermore, the convenience includes the following steps:
[0026] Step 1: Assign different identities based on the given downstream task. In Step 1, using a large language model, the following instructions are input for the given downstream task to generate several role identity information related to but different from that downstream task:
[0027] Text Generation Task: Based on the given task, generate {} different information on power industry expert reviewers. Reviewer identities should be concise and focused on specific evaluation angles. Strive for creativity and diversity in the generated identity information, and avoid omissions. The obtained power industry expert identity information includes:
[0028] Environmental Impact Assessment Engineer / Supervision Engineer / First-Class Cost Engineer / Intermediate Registered Safety Engineer / Registered Urban and Rural Planner / Registered Electrical Engineer / Registered Public Utility Equipment Engineer / Registered Architect / Registered Structural Engineer / Registered Equipment Supervisor / Registered Civil Engineer (Port and Waterway) / Registered Civil Engineer (Hydraulic and Hydropower Engineering) / Registered Civil Engineer (Geotechnical) / Registered Fire Protection Engineer / Registered First-Class Construction Engineer / Registered Consulting Engineer.
[0029] Based on the above identities, the embedding model is used to retrieve the n identities most relevant to the power assessment task, and these identities are assigned to the n models in sequence.
[0030] Step 2: Construct preference data based on the original dataset. Rewrite the original dataset using the evol-instruct method. The rewriting strategy of the evol-instruct method is (1) In-depth Evolving: making the text more complex and difficult by adding constraints, deepening, concretizing, adding reasoning steps, and complicating the input; (2) In-breadth Evolving: generating a completely new instruction based on the given instruction to increase diversity. The above two methods generate positive samples. On the basis of the evol-instruct method, the function of generating counterfactual data is added to generate negative samples.
[0031] Step 3: Adjust the scoring weights for different identity models. First, for the n different identity review models m1, m2, ..., mn obtained in Step 1... n Initialize the scoring weights w1, w2, ..., w for each model in sequence. n =1. Based on the positive and negative samples x1, x2, ..., x obtained in step 2 l Let the true scores of positive and negative samples be The positive samples should receive full marks, that is, if sample x j If it is a positive sample, then its true score is If sample x j As a negative sample, its true score is Each sample is scored by models with different identities: p1, p2, ..., p n The scores are multiplied by their weights, and then summed to obtain the multi-identity model's score for sample x. j Total score In each scoring process, the weight w i Dynamic fine-tuning to make the weighted score approximate the sample score
[0032] The final optimization goal is:
[0033]
[0034] Step 4: After the scoring weights of the entire multi-identity model are optimized, for a given input x, the overall score of the entire multi-identity evaluation model is...
[0035] This invention addresses the problems of model bias and insufficient alignment in current NLG assessments by utilizing two methods: multi-identity role-playing and dynamic weight adjustment, to improve the accuracy of NLG assessments in the power sector and their consistency with human preferences.
[0036] Compared with existing technologies, the innovation of this invention lies in two aspects. Firstly, it assigns identities based on downstream tasks in the power sector to the evaluation model, optimizing the model's assessment capabilities from a prompt perspective. Secondly, by constructing positive and negative samples to fine-tune the weights in model scoring, the overall model evaluation can be fitted to human preferences, thereby improving consistency with human assessments. Attached Figure Description
[0037] Figure 1 This is a schematic diagram of the large language model evaluation process used in this invention. Detailed Implementation
[0038] The following detailed description of the proposed solution using specific examples further illustrates the present invention. The technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are merely some embodiments of the present invention, used only as examples, and should not be construed as limiting this patent. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of the present invention.
[0039] The specific embodiments of the present invention can be further described in detail below, taking into account specific examples:
[0040] Identity retrieval based on downstream tasks: When evaluating NLG tasks, firstly, based on the specific needs of downstream tasks in the power field, multiple professional identities related to the task but with different roles are generated using a large language model. For example, if the downstream task involves fault detection in the power system, the identities that may be generated include power system engineer, data analyst, power equipment maintenance technician, etc. In this case, the downstream task is knowledge question answering in the power field. The embedding model (BGE-m3 in this case) was used to retrieve the three most relevant identities, and the corresponding identity prompts are: (1) You are an electrician, and you are very concerned about the clarity of the text. The following is the answer of the large language model to a question. From the perspective of expressiveness, score this answer from 0 to 5, and write the score in 【】. (2) You are an electrical engineer, and you are very concerned about the accuracy of the text. The following is the answer of the large language model to a question. From the perspective of factuality, score this answer from 0 to 5, and write the score in 【】. (3) You are a consulting engineer, and you are very concerned about the accuracy of the logical reasoning in the text. Below is the large language model's answer to a question. Rate this answer from 0 to 5 points based on its logical reasoning, and write your score in the brackets [ ]. These different professional identities will be used in subsequent evaluation processes to simulate the evaluation of the language model's generated content from different professional perspectives.
[0041] The original dataset was rewritten using a large language model to generate training data that is comprehensive and conforms to human cognitive preferences. Specifically, the original "question + correct answer" pairs in the power industry knowledge-answering dataset were used as the original positive samples, and the "question + incorrect answer" pairs were used as the original negative samples. The original positive samples were rewritten using semantically similar but content-diverse methods to construct positive samples based on human cognition, although the model may produce different evaluation results. Keywords in the original positive samples were reversed to construct negative samples based on human cognition, again with potentially different evaluation results. Similarly, the original negative samples were rewritten using semantically similar but content-diverse methods to construct negative samples based on human cognition, again with potentially different evaluation results. Keywords in the original negative samples were also reversed to construct positive samples based on human cognition, again with potentially different evaluation results. The positive and negative samples were then combined to obtain a training dataset that represents human cognitive preferences.
[0042] Adjusting the weights of different identity models based on the preference dataset: The weights of evaluation models for different identities are adjusted using a constructed preference dataset. Specifically, each model evaluates the training data according to the identity it represents, and then the weights are adjusted using linear regression based on how closely the model evaluation results match the actual preferences. For the n=3 different identity review models m1, m2, m3 in this case, the scoring weights w1, w2, w3 are initialized to 1 for each model sequentially. Based on the positive and negative samples x1, x2, ..., x3 obtained in the previous step... l Let the true scores of positive and negative samples be The positive samples should receive full marks, that is, if sample x j If it is a positive sample, then its true score is If sample x j As a negative sample, its true score is Each sample is scored by models with different identities (p1, p2, p3). The scores are multiplied by their weights and then summed to obtain the multi-identity model's score for sample x. j Total score In each scoring process, the weight w i Dynamic fine-tuning to make the weighted score approximate the sample score
[0043] Its optimization objective is:
[0044]
[0045] Training the model using gradient descent can optimize the weights of the multi-identity evaluation model, ensuring that the evaluation results better reflect human preferences and evaluation criteria.
[0046] Large language model evaluation based on linear weighting of multiple identity models: Finally, the evaluation results of different identity models are linearly weighted to form the final evaluation score. Different identity models participate in the calculation of the final score according to the weights adjusted in the previous step, ensuring that the evaluation process can fully consider the evaluation results from different professional perspectives, thereby improving the comprehensiveness and accuracy of the evaluation.
[0047] By employing this multi-identity role-playing and dynamic weight adjustment method, this invention improves the performance of large language models in NLG task evaluation in the power sector, reduces bias in the model evaluation process, enhances consistency with human evaluation, and ultimately achieves more accurate and fair evaluation results.
[0048] In this case study experiment, the prediction results showed that, based on the Tigerbot-70B model, a multi-identity evaluation model was constructed. Compared with the evaluation using a single Tigerbot-70B model, the similarity to human evaluation improved from 0.08 to 0.68 in the summarization task, from 0.39 to 0.70 in the rewriting task, and from 0.06 to 0.22 in the question answering task.
[0049] Analysis: The original Tigerbot-70B model itself may have inherent limitations, and in evaluations, especially in specialized fields, it may not align well with human preferences. However, by adopting a multi-identity, multi-model joint evaluation method, the bias of the model itself towards specific evaluation content can be effectively offset, making it closer to human preferences and improving the comprehensiveness and accuracy of the large language model's ability assessment in the power field.
[0050] Obviously, the above description is only a partial embodiment of the present invention. It should be noted that, for those skilled in the art, several improvements and adjustments can be made according to actual needs without departing from the principle of the present invention, and these improvements and adjustments should also be considered within the scope of protection of the present invention.
Claims
1. A method for evaluating large language models based on multi-identity role-playing in the power sector, characterized in that, The method includes: Identity retrieval based on downstream tasks; Rewriting the original dataset based on a large language model; Adjust the weights of different identity models based on the preference dataset; Evaluation of large language models based on linear weighting of multi-identity models.
2. The method for evaluating large language models based on multi-identity role-playing in the power sector according to claim 1, characterized in that, Assigning a different identity to each large language model used for evaluation, the identity retrieval method based on downstream tasks includes: Based on the textual description of the downstream task, a large language model is used to generate professional identities that are related to but different from the downstream task, each with a background in the power industry. The most matching identity is assigned to each evaluation model for role-playing.
3. The method for evaluating large language models based on multi-identity role-playing in the power sector according to claim 2, characterized in that, The original dataset, rewritten using the large language model, is used to obtain preference data. The original dataset refers to an expert knowledge question-and-answer dataset in the power industry. Each sample in the dataset contains a question and a reference answer and an incorrect answer in natural language form. The reference answer is a positive sample, and the incorrect answer is a negative sample. The rewritten original dataset based on the large language model includes: Rewrite the samples included in the original dataset; Positive samples are rewritten using semantically similar but content-diverse methods to construct positive samples based on human cognition; keywords in positive samples are reversed to construct negative samples based on human cognition. Negative samples are rewritten using semantically similar but content-diverse methods to construct negative samples based on human cognition; keywords in negative samples are reversed to construct positive samples based on human cognition. Combining positive and negative samples yields a training dataset that can represent human cognitive preferences.
4. The method for evaluating large language models based on multi-identity role-playing in the power sector according to claim 3, characterized in that, The aforementioned adjustment of the weights of different identity models based on the preference dataset involves using positive and negative samples to evaluate different identity models and adjusting their relative weights. This adjustment includes: Initialize the same weights for each evaluation model; Each training data point is evaluated using an evaluation model with different identities. The model scores the data based on its assigned identity, thereby offsetting any bias that the model itself may contain. The model's score is multiplied by its weight and then compared with the score of the dataset itself. The weight of the model whose score is closer to the ground truth is increased, and the weight of the model whose score deviates from the ground truth is decreased. After optimization using multiple training datasets, the model identity and corresponding weights are obtained by dynamically adjusting them based on different tasks.