Method and system for evaluating effect of large language model in power field

By constructing an evaluation question bank covering all aspects of knowledge in the power field, and combining multi-model testing and manual evaluation, the subjectivity and applicability issues of large language model evaluation in the power field have been resolved, achieving efficient and accurate effect evaluation.

CN118093371BActive Publication Date: 2025-12-26BEIJING GUODIANTONG NETWORK TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410083297.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2025-12-26
Estimated Expiration
2044-01-19

AI Technical Summary

Technical Problem

In existing technologies, the evaluation methods for large language models in the power field mainly rely on manual evaluation, which is time-consuming, labor-intensive, and highly subjective. Furthermore, existing text matching models are not very applicable to the evaluation of long text generation.

Method used

We will construct a large language model evaluation question bank for the power industry. Through multi-model testing and manual testing, we will form an evaluation question bank covering all aspects of knowledge in the power industry. The model performance will be evaluated based on the accuracy of the answers.

Benefits of technology

It achieves objective and highly applicable assessment of large language models in the power sector, reduces the subjectivity of manual evaluation, and improves the efficiency and accuracy of evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118093371B_ABST
    Figure CN118093371B_ABST
Patent Text Reader

Abstract

The application provides an evaluation method and system for the effect of a large language model in the field of electric power, comprising: inputting a pre-constructed evaluation question bank of the large language model in the field of electric power into the large language model in the field of electric power to obtain an answer result; calculating an answer accuracy based on the answer result, and evaluating the effect of the large language model in the field of electric power based on the answer accuracy; wherein the evaluation question bank of the large language model in the field of electric power is constructed through investigation of various application scenarios in the field of electric power, and multi-model testing and manual testing. The application constructs the evaluation question bank of the large language model in the field of electric power through investigation of various application scenarios in the field of electric power, and multi-model testing and manual testing. The question bank covers knowledge in all aspects of the field of electric power, can objectively evaluate the effect of the large language model in the field of electric power, and has high applicability.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of electric power and the field of large language models, and in particular to a method and system for evaluating the effect of a large language model in the field of electric power. BACKGROUND

[0002] (1) In the current use scenarios of large language models in vertical fields, most evaluation methods are manual evaluation. Specifically, the results output by the large language model are fed back to the researchers, who then manually judge whether the results are reasonable. Manual design of evaluation indicators is time-consuming and labor-intensive, and may have certain subjectivity and limitations.

[0003] (2) In some applications of large language models in vertical fields, a text matching model (such as QA matching and phrase matching) is used for evaluation. Specifically, the answer generated by the large language model is matched with the labeled answer in terms of similarity. This evaluation is more suitable for matching of words or phrases, and has certain difficulties in evaluating long text generation, and has low applicability.

[0004] (3) There is an evaluation method in existing technologies that uses a well-performing open-source large language model (such as GPT4) to fine-tune a transfer model for evaluation, and then uses the transfer model to evaluate the fine-tuned model in the vertical field. This method is limited by the performance of the transfer model and has certain limitations. SUMMARY

[0005] To solve the problem of manual design of evaluation indicators in existing technologies, which is time-consuming and labor-intensive, and may have certain subjectivity and limitations, and the problem that the evaluation of large language models using text matching models is more suitable for matching of words or phrases, and has certain difficulties in evaluating long text generation, and has low applicability, the present application proposes a method for evaluating the effect of a large language model in the field of electric power, comprising:

[0006] Substitute the pre-constructed large language model evaluation question bank in the field of electric power into the large language model in the field of electric power to obtain the answer results;

[0007] Calculate the correct answer rate based on the answer results, and evaluate the effect of the large language model in the field of electric power based on the correct answer rate; wherein the large language model evaluation question bank in the field of electric power is constructed through research on various application scenarios in the field of electric power, and through multiple model testing and manual testing.

[0008] Optionally, the construction of the large language model evaluation question bank in the field of electric power comprises:

[0009] Construct a preliminary large language model evaluation question bank in the field of electric power by researching various application scenarios in the field of electric power;

[0010] The preliminary power field large language model evaluation question bank is subjected to multi-model testing and artificial evaluation;

[0011] The preliminary power field large language model evaluation question bank is modified based on the multi-model testing results and the artificial evaluation results to obtain a modified power field large language model evaluation question bank;

[0012] The modified power field large language model evaluation question bank is subjected to multi-model testing and artificial evaluation until the power field large language model answer accuracy reaches a set value, and no modification suggestion is provided by artificial evaluation, and the modified power field large language model evaluation question bank is used as the power field large language model evaluation question bank.

[0013] Optionally, the constructing of the preliminary power field large language model evaluation question bank through the investigation of various application scenarios in the power field comprises:

[0014] Various questions and answers with a frequency reaching a set frequency in the power field professional knowledge question bank are collected by investigating and using the power field professional knowledge scenarios to form a power field professional knowledge question bank;

[0015] A part of questions and answers are extracted from the power field common sense question bank according to the commonality and importance of common sense to form a power field common sense question bank;

[0016] Questions and answers with a frequency reaching a set threshold are collected from various safety scenarios in the power field, and knowledge is extracted from typical safety accidents to form a power field safety knowledge question bank;

[0017] The proportion of the power field professional knowledge question bank, the power field common sense question bank and the power field safety knowledge question bank in the business scenarios is obtained by investigating and inquiring the employees in the power industry and by making the employees test the dialogue model;

[0018] The preliminary power field large language model evaluation question bank is formed based on the power field professional knowledge question bank, the power field common sense question bank and the power field safety knowledge question bank and the proportion of the power field professional knowledge question bank, the power field common sense question bank and the power field safety knowledge question bank.

[0019] Optionally, the multi-model testing and artificial evaluation of the preliminary power field large language model evaluation question bank comprises:

[0020] The preliminary power field large language model evaluation question bank is substituted into each open source large language model without vertical field fine-tuning to obtain a testing result without fine-tuning;

[0021] The preliminary power field large language model evaluation question bank is substituted into each open source large language model fine-tuned in the power field to obtain fine-tuned test results;

[0022] The preliminary power field large language model evaluation question bank is substituted into a commercial large language model to obtain commercial test results.

[0023] The preliminary power field large language model evaluation question bank is artificially evaluated.

[0024] Optionally, the preliminary power field large language model evaluation question bank is modified based on the multi-model test results and the artificial evaluation results to obtain a modified power field large language model evaluation question bank, which includes:

[0025] The un-tuned test results are analyzed to record questions that none of the large language models can output correct answers.

[0026] For power field professional knowledge questions, collect questions that none of the power field large language models can output correct answers to form a mistake question set.

[0027] For common sense and safety field question banks, check the record, compare the fine-tuned test results, collect questions that cannot be answered accurately and appear in the record, and add them to the mistake question set.

[0028] The commercial test results are analyzed to collect questions that cannot output correct answers, compared with the mistake question set, the questions appearing in the mistake question set are marked, and the collected questions that cannot output correct answers are added to the mistake question set to improve the mistake question set.

[0029] Calculate the error rate of the questions answered incorrectly in the artificial evaluation results, compare the questions with the error rate reaching the set threshold with the mistake question set, and mark the questions with the error rate reaching the set threshold in the mistake question set.

[0030] The preliminary power field large language model evaluation question bank is modified based on the mistake question set, the opinions and suggestions of the artificial evaluation results, and the opinions and suggestions of the answerers.

[0031] The multi-model test results include un-tuned test results, fine-tuned test results, and commercial test results.

[0032] Optionally, the preliminary power field large language model evaluation question bank is modified based on the mistake question set, the opinions and suggestions of the artificial evaluation results, and the opinions and suggestions of the answerers, which includes:

[0033] According to the opinions and suggestions of the answerers, the preliminary power field large language model evaluation question bank is modified.

[0034] Analyzing the marked questions in the wrong question set, checking whether the question expression has ambiguity, modifying the expression method for the questions with ambiguity, deleting the questions with correct expression, and replacing the questions by re-screening the questions in each application scenario in the power field through investigation.

[0035] Optionally, the manual evaluation on the preliminary power field large language model evaluation question bank comprises:

[0036] Manually answering the questions in the power field large language model evaluation question bank.

[0037] After the answering is completed, collecting and sorting the opinions and suggestions of the participants on the preliminary power field large language model evaluation question bank.

[0038] Optionally, the effect of the power field large language model is evaluated based on the correct rate of the answers, which comprises:

[0039] If the correct rate of the answers reaches a correct rate threshold, the effect of the power field large language model is good, otherwise, the effect of the power field large language model is poor.

[0040] In still another aspect, the application further provides an evaluation system for the effect of a power field large language model, which comprises:

[0041] An answering module, configured to input a pre-constructed power field large language model evaluation question bank into a power field large language model to obtain an answer result.

[0042] An evaluation module, configured to calculate a correct rate of answers based on the answer result, and evaluate the effect of the power field large language model based on the correct rate of answers.

[0043] The power field large language model evaluation question bank is constructed through investigation of each application scenario in the power field, and multi-model testing and manual testing.

[0044] Optionally, it further comprises a question bank construction module, configured to construct a power field large language model evaluation question bank.

[0045] Optionally, the question bank construction module comprises:

[0046] A preliminary construction submodule, configured to construct a preliminary power field large language model evaluation question bank through investigation of each application scenario in the power field.

[0047] A test and evaluation submodule, configured to perform multi-model testing and manual evaluation on the preliminary power field large language model evaluation question bank.

[0048] The modifying submodule is configured to modify the preliminary power field large language model evaluation question bank based on the multi-model test results and the artificial evaluation results, to obtain a modified power field large language model evaluation question bank.

[0049] The question bank determining submodule is configured to perform multi-model testing and artificial evaluation on the modified power field large language model evaluation question bank until the correct answer rate of the large language model after fine-tuning in the power field reaches a set value, and no modification suggestion is provided by artificial evaluation, and the modified power field large language model evaluation question bank is used as the power field large language model evaluation question bank.

[0050] Optionally, the preliminary construction submodule is specifically configured to:

[0051] The frequency of each question and answer in the power field professional knowledge question bank is set to a certain frequency through investigation and collection of professional knowledge scenarios in the power field.

[0052] The frequency of each question and answer in the power field common sense question bank is set to a certain threshold through investigation and collection of common sense knowledge in the power field.

[0053] The frequency of each question and answer in the power field safety knowledge question bank is set to a certain threshold through investigation and collection of safety scenarios in the power field.

[0054] The proportion of each question and answer in the power field professional knowledge question bank, the power field common sense question bank and the power field safety knowledge question bank is obtained through investigation and inquiry of employees in the power industry and dialogue model testing of employees.

[0055] The preliminary power field large language model evaluation question bank is formed based on the power field professional knowledge question bank, the power field common sense question bank and the power field safety knowledge question bank, and the proportion of the power field professional knowledge question bank, the power field common sense question bank and the power field safety knowledge question bank.

[0056] Optionally, the test and evaluation submodule includes:

[0057] The un-debugged test submodule is configured to input the preliminary power field large language model evaluation question bank into each open source large language model that has not been fine-tuned in the vertical field, to obtain un-tuned test results.

[0058] The power test submodule is configured to input the preliminary power field large language model evaluation question bank into each open source large language model that has been fine-tuned in the power field, to obtain fine-tuned test results.

[0059] A commercial test submodule is configured to input the preliminary power field large language model evaluation question bank into a commercial large language model to obtain a commercial test result.

[0060] An artificial evaluation submodule is configured to artificially evaluate the preliminary power field large language model evaluation question bank.

[0061] Optionally, the modifying submodule is specifically configured to:

[0062] analyze the un-tuned test result, and record questions that cannot be answered correctly by each large language model;

[0063] collect questions that cannot be answered correctly by each large language model to form a wrong question set for a power field professional knowledge question;

[0064] for a common sense and safety field question bank, check the record, compare the tuned test result, collect questions that cannot be answered correctly and appear in the record, and add the questions to the wrong question set;

[0065] analyze the commercial test result, collect questions that cannot be answered correctly, compare the questions with the wrong question set, mark questions in the wrong question set, and add the collected questions that cannot be answered correctly to the wrong question set to improve the wrong question bank;

[0066] calculate the error rate of questions answered incorrectly in the artificial evaluation result, compare the questions with the wrong question set, and mark questions with an error rate reaching a set threshold in the wrong question set;

[0067] modify the preliminary power field large language model evaluation question bank based on the wrong question set, opinions and suggestions of the artificial evaluation result;

[0068] The multiple model test result includes an un-tuned test result, a tuned test result and a commercial test result.

[0069] Optionally, the implementation steps of the modifying submodule for modifying the preliminary power field large language model evaluation question bank based on the wrong question set and opinions and suggestions of the artificial evaluation result include:

[0070] modify the preliminary power field large language model evaluation question bank according to the opinions and suggestions of the answerers;

[0071] analyze the questions marked in the wrong question set, check whether the question expression has ambiguity, modify the expression method for the questions with ambiguity, delete the questions with correct expression, and replace the questions by re-screening the questions through research on various application scenarios in the power field.

[0072] Optionally, the artificial evaluation sub-module is specifically used for:

[0073] artificially answering the power field large language model evaluation question bank;

[0074] After the answering is completed, the opinions and suggestions of the participants on the preliminary power field large language model evaluation question bank are collected and arranged.

[0075] Optionally, the evaluation module is specifically used for:

[0076] If the correct answer rate reaches the correct answer rate threshold, the effect of the power field large speech model is good, otherwise, the effect of the power field large speech model is poor.

[0077] In another aspect, the present application also provides a computing device, comprising: one or more processors;

[0078] a processor for executing one or more programs;

[0079] When the one or more programs are executed by the one or more processors, a method for evaluating the effect of a large language model in the power field is realized.

[0080] In another aspect, the present application also provides a computer readable storage medium having a computer program stored thereon, which, when executed, realizes a method for evaluating the effect of a large language model in the power field.

[0081] Compared with the prior art, the present application has the following advantages:

[0082] The present application provides a method for evaluating the effect of a large language model in the power field, comprising substituting a pre-constructed power field large language model evaluation question bank into a power field large language model for answering to obtain an answer result; calculating an answer correct rate based on the answer result, and evaluating the effect of the power field large language model based on the answer correct rate; wherein the power field large language model evaluation question bank is constructed through investigating various application scenarios in the power field, and through multi-model testing and artificial testing. The present application constructs a power field large language model evaluation question bank through investigating various application scenarios in the power field, and through multi-model testing and artificial testing. The question bank covers knowledge in all aspects of the power field, and can objectively evaluate the effect of the power field large language model, and has high applicability. BRIEF DESCRIPTION OF DRAWINGS

[0083] Figure 1 A flow chart of a method for evaluating the effect of a large language model in the power field of the present application;

[0084] Figure 2 A process chart for constructing a power field large language model evaluation question bank of the present application. Detailed Implementation

[0085] Since the public release of ChatGPT 3.0, large language models have sparked a research boom. Subsequent iterations and updates have led to the development of several foundational large language models, including ChatGPT 3.5, ChatGPT 4.0, Llama, ChatGLM, MOSS, and Wenxin Yiyan. However, limitations imposed by hardware capabilities such as server computing power, and the relatively low open-source nature of these foundational models and the difficulty in acquiring and integrating pre-training datasets, have limited most applications across various fields and industries to fine-tuning large language models within specific vertical domains. While some fine-tuning applications of large language models exist in the power industry, there is a lack of metrics for evaluating these fine-tuned models. This invention proposes a method for evaluating the performance of large language models in the power industry. It presents a general evaluation metric library for large language models in the power industry, covering specialized knowledge, general encyclopedic knowledge, and safety knowledge within the power sector, which can be used to evaluate large language models in this field. This invention also proposes a design concept and method for an evaluation metric library for large language models in vertical domains, providing guidance for the design and evaluation of other large language models in other vertical domains.

[0086] Example 1:

[0087] A method for evaluating the performance of large language models in the power sector, such as... Figure 1 As shown, it includes:

[0088] Step S1: Substitute the pre-built large language model evaluation question bank in the power field into the large language model in the power field to answer questions and obtain the answer results;

[0089] Step S2: Calculate the answer accuracy rate based on the answer results, and evaluate the effectiveness of the power field large language model based on the answer accuracy rate;

[0090] The large language model evaluation question bank for the power sector was constructed through research on various application scenarios in the power sector, and after multi-model testing and manual testing.

[0091] Before step S1, the process also includes constructing a large language model evaluation question bank for the power industry. The following section will combine... Figure 2 The construction process of this large language model evaluation question bank in the power field is introduced:

[0092] A preliminary evaluation question bank for a large language model in the power field was constructed by investigating various application scenarios in the power sector.

[0093] The preliminary large-scale language model evaluation question bank in the power field was subjected to multi-model testing and manual evaluation.

[0094] modify the preliminary power field large language model evaluation question bank based on the multi-model test results and artificial evaluation results to obtain a modified power field large language model evaluation question bank;

[0095] Perform multi-model testing and artificial evaluation on the modified power field large language model evaluation question bank until the correct answer rate of the large language model after fine-tuning in the power field reaches a set value, and there is no modification suggestion from artificial evaluation. The modified power field large language model evaluation question bank is used as the power field large language model evaluation question bank.

[0096] (1) The construction process of the preliminary power field large language model evaluation question bank specifically includes:

[0097] To design a more professional and practical power field large language model evaluation question bank, we investigated various application scenarios in the power field, consulted with professors and experts in the industry, and summarized the opinions of all parties. Finally, we selected the professional knowledge scenarios in the power field, various safety scenarios in the power field, and common encyclopedic knowledge in the power field as the basis for constructing the question bank.

[0098] First, we investigated the use of professional knowledge scenarios in the power field, including college power professional test questions, industry recruitment test questions, and power field professional title examination test questions. We collected high-frequency questions and answers from various test questions to form a power field professional knowledge question bank.

[0099] Second, we investigated and viewed common encyclopedic knowledge in the power field, including power encyclopedias and power common sense. Based on the commonality and importance of common sense, we extracted some questions and answers to form a power field common sense question bank.

[0100] Third, we investigated various safety scenarios in the power field, including common power field safety knowledge, recent power field safety accidents, and safety test questions in the Xitexuesheng (Think Extreme School). We collected high-frequency questions and answers, as well as related knowledge from typical safety accidents to form a power field safety knowledge question bank. Finally, we summarized and formed a power field large language model evaluation question bank.

[0101] The power industry employees are investigated and inquired, and the employees are tested by the dialogue model. It is learned that the use of professional knowledge in the power field is the highest in the business scenario, and the employees raise the most questions during the test. Therefore, the scene topic is designed to account for the first set ratio of the entire power field large language model evaluation question bank, and the embodiment takes 50%; the employees are relatively concerned about the safety knowledge in the power field, so the scene topic is set to account for the second set ratio of the entire power field large language model evaluation question bank, and the embodiment takes 30%; the employees have relatively low attention to the power common sense knowledge, so the scene topic is set to account for the third set ratio of the entire power field large language model evaluation question bank, and the embodiment takes 20%. The sum of the first set ratio, the second set ratio and the third set ratio is 1.

[0102] The preliminary power field large language model evaluation question bank is constructed by 50% of the power field professional knowledge question bank, 30% of the power field common sense question bank and 20% of the power field safety knowledge question bank.

[0103] (2) The preliminary power field large language model evaluation question bank is tested by multiple models.

[0104] After forming the preliminary power field large model performance evaluation question bank, the open source large language model supporting Chinese and the commercial large language model such as ChatGLM, MOSS, Llama-2 and Wenxin Yanyan are deployed to verify the question bank.

[0105] First, the open source large language model is inputted into the question bank without vertical field fine-tuning, and the correctness and accuracy of the model answering the questions are tested.

[0106] Secondly, after fine-tuning the open source large language model in the power field, the question bank is inputted again, and the correctness and accuracy of the model answering the questions are tested. Here, the open source large language model is fine-tuned in the power field, which is specifically to construct question and answer pairs by using professional knowledge and business knowledge related to the power field, and provide the question and answer pairs to the open source large language model for training, to obtain the open source large language model fine-tuned in the power field.

[0107] Thirdly, the commercial large language model is used to test the correctness and accuracy of the model answering the questions.

[0108] After the above tests, the results obtained by testing the un-tuned large language models are analyzed, mainly to view the common sense and safety field question bank, and to record the questions that cannot output correct answers in each large language model.

[0109] After analyzing the results of the test on the fine-tuned large language models, for the professional knowledge questions in the power field, collect the questions that none of the large language models can output the correct answer to form a "mistake set". For the common sense and safety field question bank, check the previous records and compare the current test results. If there are still questions that cannot be answered accurately in the current test and appear in the records, collect such questions and add them to the "mistake set".

[0110] Finally, analyze the results of the test on the commercial large language models, collect the questions that cannot output the correct answer, and compare them with the previously formed mistake set. For the questions that appear in this test and in the "mistake set", mark them as important. For questions that do not appear in the "mistake set", add them to the mistake set.

[0111] (3) Manually evaluate the preliminary power field large language model evaluation question bank.

[0112] After forming the power field large model performance evaluation question bank, invite professors, well-known experts, business backbone and people outside the industry in the industry to manually answer the questions. After completing the answers, collect the questions that appear frequently in the answer sheets and compare them with the "mistake set" formed previously in the model test. Mark the questions that exist in both. In addition, after completing the answers, ask and collect the opinions and suggestions of the professors, experts, backbone and people outside the industry for the question bank for modification.

[0113] (4) Modify the preliminary power field large language model evaluation question bank.

[0114] After model testing and manual answering, a "mistake set" is formed. For the questions marked in the "mistake set", modification is needed. First, modify the power field large language model evaluation question bank according to the opinions and suggestions provided by professors, experts and people outside the industry. Second, analyze the questions marked in the "mistake set", check if there is ambiguity in the question statement, modify the expression method for questions with ambiguity, delete questions with correct expression, and replace them with new questions selected according to the method in (1).

[0115] (5) Repeat the cross-validation of the modified preliminary power field large language model.

[0116] After modifying the question bank based on the "mistake set" and expert opinions, repeat steps (2), (3) and (4) to re-test and modify the modified question bank until the large language model fine-tuned in the power field can answer most of the questions in the question bank correctly and experts have no further opinions. Form a power field large language model evaluation question bank and a power field large language model rating system methodology.

[0117] The application provides a power field large language model evaluation question bank, which can be used for performance evaluation of a fine-tuned power field large language model, and provides technical index support for iterative optimization of the model.

[0118] The application provides a method for evaluating the effect of a power field large language model, which is used to support subsequent self-defined construction of a power field large language model evaluation question bank required by a user.

[0119] The power field large language model evaluation method provided by the application can be extended to other fields, and provides a feasible method support for designing a large language model performance evaluation.

[0120] Embodiment 2

[0121] The application based on the same inventive concept also provides an evaluation system for the effect of a power field large language model, which comprises:

[0122] An answering module is configured to input a pre-constructed power field large language model evaluation question bank into the power field large language model to obtain an answer result.

[0123] An evaluation module is configured to calculate an answer accuracy based on the answer result, and evaluate the effect of the power field large language model based on the answer accuracy.

[0124] The power field large language model evaluation question bank is constructed through investigation of various application scenarios in the power field, and multiple model testing and manual testing.

[0125] Optionally, the system further comprises a question bank construction module configured to construct the power field large language model evaluation question bank.

[0126] Optionally, the question bank construction module comprises:

[0127] A preliminary construction submodule is configured to construct a preliminary power field large language model evaluation question bank through investigation of various application scenarios in the power field.

[0128] A test and evaluation submodule is configured to perform multiple model testing and manual evaluation on the preliminary power field large language model evaluation question bank.

[0129] A modification submodule is configured to modify the preliminary power field large language model evaluation question bank based on the multiple model testing results and the manual evaluation results, to obtain a modified power field large language model evaluation question bank.

[0130] The question bank determining sub-module is configured to perform multi-model testing and manual evaluation on the modified power field large language model evaluation question bank until the correct answer rate of the large language model after power field fine-tuning reaches a set value and no modification suggestion is provided by manual evaluation, and the modified power field large language model evaluation question bank is used as the power field large language model evaluation question bank.

[0131] Optionally, the preliminary construction sub-module is specifically configured to:

[0132] collect questions and answers with a frequency reaching a set frequency in various test questions by investigating professional knowledge scenarios in the power field to form a power field professional knowledge question bank;

[0133] extract part of the questions and answers according to commonality and importance of common sense to form a power field common sense question bank by investigating and viewing common sense knowledge in the power field;

[0134] collect questions and answers with a frequency reaching a set threshold in various test questions by investigating various safety scenarios in the power field, and extract knowledge from typical safety accidents to form a power field safety knowledge question bank;

[0135] obtain the proportion of the power field professional knowledge question bank, the power field common sense question bank and the power field safety knowledge question bank in the business scenarios by investigating and inquiring employees in the power industry and making employees test the dialogue model;

[0136] form a preliminary power field large language model evaluation question bank based on the power field professional knowledge question bank, the power field common sense question bank and the power field safety knowledge question bank, and the proportion of the power field professional knowledge question bank, the power field common sense question bank and the power field safety knowledge question bank.

[0137] Optionally, the test and evaluation sub-module includes:

[0138] The un-debugged test sub-module is configured to input the preliminary power field large language model evaluation question bank into each open source large language model without vertical field fine-tuning to obtain un-fine-tuned test results;

[0139] The power test sub-module is configured to input the preliminary power field large language model evaluation question bank into each open source large language model fine-tuned in the power field to obtain fine-tuned test results;

[0140] The commercial test sub-module is configured to input the preliminary power field large language model evaluation question bank into a commercial large language model to obtain commercial test results;

[0141] The manual evaluation sub-module is configured to perform manual evaluation on the preliminary power field large language model evaluation question bank.

[0142] Optionally, the modifying submodule is specifically used for:

[0143] analyzing the un-tuned test results, and recording the questions for which each large language model fails to output a correct answer;

[0144] for the professional knowledge questions in the power field, collecting the questions for which the large language models in the power field all fail to output a correct answer, to form a question set;

[0145] for the common sense and safety field question bank, checking the record, comparing the tuned test results, collecting questions that fail to answer accurately and appear in the record, and adding them to the question set;

[0146] analyzing the commercial test results, collecting questions that fail to output a correct answer, comparing them with the question set, highlighting questions that appear in the question set, and adding the collected questions that fail to output a correct answer to the question set to improve the question bank;

[0147] calculating the error rate of the questions answered incorrectly in the artificial evaluation results, comparing the questions with an error rate reaching a set threshold with the question set, and marking the questions with an error rate reaching a set threshold in the question set;

[0148] modifying the preliminary large language model evaluation question bank in the power field based on the question set, the opinions and suggestions of the question answering personnel in the artificial evaluation results;

[0149] The multiple model test results include: un-tuned test results, tuned test results, and commercial test results.

[0150] Optionally, the implementation steps of modifying the preliminary large language model evaluation question bank in the power field based on the question set, the opinions and suggestions of the question answering personnel in the artificial evaluation results in the modifying submodule include:

[0151] modifying the preliminary large language model evaluation question bank in the power field according to the opinions and suggestions of the question answering personnel;

[0152] analyzing the questions marked in the question set, checking whether the question expression has ambiguity, modifying the expression method for questions with ambiguity, deleting questions with correct expression, and replacing them with new questions selected by researching various application scenarios in the power field.

[0153] Optionally, the artificial evaluation submodule is specifically used for:

[0154] artificially answering the large language model evaluation question bank in the power field;

[0155] After the answering is completed, opinions and suggestions of the people participating in the answering on the preliminary power field large language model evaluation question bank are collected and arranged.

[0156] Optionally, the evaluation module is specifically configured to:

[0157] If the correct answer rate reaches the correct rate threshold, the effect of the power field large speech model is good, otherwise, the effect of the power field large speech model is poor.

[0158] Embodiment 3:

[0159] Based on the same invention concept, the application further provides a computer device, which comprises a processor and a memory, the memory is used for storing a computer program, the computer program comprises program instructions, and the processor is used for executing the program instructions stored in the computer storage medium. The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions in the computer storage medium to realize the corresponding method process or corresponding function, so as to realize the steps of the above-mentioned embodiment of the method for evaluating the effect of the power field large language model.

[0160] Embodiment 4:

[0161] Based on the same inventive concept, the present application also provides a storage medium, specifically a computer readable storage medium (Memory), which is a memory device in a computer device, used for storing programs and data. It can be understood that the computer readable storage medium here can include the built-in storage medium in the computer device, and of course can also include the expansion storage medium supported by the computer device. The computer readable storage medium provides a storage space, which stores the operating system of the terminal. Moreover, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and the instructions can be one or more computer programs (including program codes). It should be noted that the computer readable storage medium here can be a high-speed RAM memory, or a non-volatile memory such as at least one disk memory. One or more instructions stored in the computer readable storage medium can be loaded and executed by the processor to implement the steps of the above-mentioned embodiment of the method for evaluating the effect of the large language model in the power field.

[0162] Those skilled in the art will appreciate that embodiments of the application can be supplied as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.

[0163] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocks Figure 1 The means for implementing the functions specified in one or more flows and / or blocks.

[0164] These computer program instructions can also be stored in a computer readable memory capable of guiding the computer or other programmable data processing devices to work in a specific way, so that the instructions stored in the computer readable memory produce a product including instruction means, which implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocksFigure 1 the function specified in the one or more blocks.

[0165] These computer program instructions can also be loaded into computer or other programmable data processing devices, so that a series of operation steps are performed on the computer or other programmable data processing devices to generate computer-implemented processes, thus the instructions executed on the computer or other programmable data processing devices provide processes for implementing the flows Figure 1 the flows or the flows and / or blocks Figure 1 the steps of the function specified in the one or more blocks.

[0166] The above merely illustrates the embodiments of the present application, and is not intended to limit the present application, and any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are included in the scope of the claims of the present application.

Claims

1. An evaluation method for the effect of a large language model in the field of electric power, characterized by, The application comprises the following steps: The pre-constructed power field large language model evaluation question bank is input into the power field large language model to obtain an answer result; The effect of the power field large language model is evaluated based on the answer correctness rate. The power field large language model evaluation question bank is constructed by investigating various application scenarios in the power field, and is constructed through multi-model testing and manual testing. The construction of the power field large language model evaluation question bank comprises the following steps: A preliminary power field large language model evaluation question bank is constructed by investigating various application scenarios in the power field; The preliminary power field large language model evaluation question bank is subjected to multi-model testing and manual evaluation; The preliminary power field large language model evaluation question bank is modified based on the multi-model testing results and the manual evaluation results to obtain a modified power field large language model evaluation question bank; The modified power field large language model evaluation question bank is subjected to multi-model testing and manual evaluation until the answer correctness rate of the power field large language model reaches a set value, and no modification suggestion is provided by manual evaluation, and the modified power field large language model evaluation question bank is used as the power field large language model evaluation question bank. The modification of the preliminary power field large language model evaluation question bank based on the multi-model testing results and the manual evaluation results comprises the following steps: The test results of the un-tuned models are analyzed, and the questions that cannot be answered correctly by all the large language models are recorded; For the professional knowledge questions in the power field, the questions that cannot be answered correctly by all the large language models are collected to form a mistake question set; For the common sense and safety field question bank, the record is checked, the test results of the tuned models are compared, the questions that cannot be answered correctly and appear in the record are collected, and the questions are added to the mistake question set; The commercial test results are analyzed, the questions that cannot be answered correctly are collected, and the questions are compared with the mistake question set, the questions appearing in the mistake question set are marked, and the questions that cannot be answered correctly are added to the mistake question set to improve the mistake question set; The error rate of the questions answered incorrectly in the manual evaluation results is calculated, the questions with the error rate reaching a set threshold are compared with the mistake question set, and the questions with the error rate reaching the set threshold are marked in the mistake question set; The preliminary power field large language model evaluation question bank is modified based on the mistake question set, the opinions and suggestions of the manual evaluation results; The multi-model testing results comprise the un-tuned test results, the tuned test results and the commercial test results.

2. The method of claim 1, wherein, The preliminary power field large language model evaluation question bank is constructed by investigating various application scenarios in the power field, and is constructed through multi-model testing and manual testing. The professional knowledge question bank in the power field is formed by collecting the questions and answers with a frequency reaching a set frequency in various test questions in the professional knowledge scenarios in the power field; The common sense question bank in the power field is formed by extracting part of the questions and answers according to the commonality and importance of common sense. A safety knowledge question bank in the power field is formed by investigating various safety scenarios in the power field, collecting questions and answers with a frequency reaching a set threshold, and extracting knowledge from typical safety accidents; The proportion of the professional knowledge question bank in the power field, the common sense question bank in the power field, and the safety knowledge question bank in the power field in the business scenarios is obtained by investigating employees in the power industry and having the employees test the dialogue model; A preliminary evaluation question bank of the large language model in the power field is formed based on the professional knowledge question bank in the power field, the common sense question bank in the power field, and the safety knowledge question bank in the power field, and the proportion of the professional knowledge question bank in the power field, the common sense question bank in the power field, and the safety knowledge question bank in the power field.

3. The method of claim 1, wherein, The preliminary evaluation question bank of the large language model in the power field is tested by multiple models and artificially evaluated, including: The preliminary evaluation question bank of the large language model in the power field is substituted into each open source large language model that has not been fine-tuned in the vertical field to obtain an un-tuned test result; The preliminary evaluation question bank of the large language model in the power field is substituted into each open source large language model that has been fine-tuned in the power field to obtain a fine-tuned test result; The preliminary evaluation question bank of the large language model in the power field is substituted into a commercial large language model to obtain a commercial test result; The preliminary evaluation question bank of the large language model in the power field is artificially evaluated.

4. The method of claim 1, wherein, The preliminary evaluation question bank of the large language model in the power field is modified based on the mistake question set, the opinions and suggestions of the answerers in the artificial evaluation results, including: The preliminary evaluation question bank of the large language model in the power field is modified according to the opinions and suggestions of the answerers; The questions marked in the mistake question set are analyzed to check whether the question expressions have ambiguities. For the questions with ambiguities, the expression methods are modified. For the questions with correct expressions, the questions are deleted and replaced by re-screening the questions through investigating various application scenarios in the power field.

5. The method of claim 3, wherein, The preliminary evaluation question bank of the large language model in the power field is artificially evaluated, including: The preliminary evaluation question bank of the large language model in the power field is artificially answered; After the answering is completed, the opinions and suggestions of the participants on the preliminary evaluation question bank of the large language model in the power field are collected and sorted.

6. The method of claim 1, wherein, The effect of the large language model in the power field is evaluated based on the answering accuracy, including: If the answering accuracy reaches a correct rate threshold, the large language model in the power field is effective, otherwise, the large language model in the power field is ineffective.

7. A system for implementing the method for evaluating the effect of a large language model in the field of electric power according to any one of claims 1-6, characterized in that, Including: An answering module configured to substitute a pre-constructed evaluation question bank of a large language model in the power field into the large language model in the power field to obtain an answering result; An evaluation module configured to calculate an answering accuracy based on the answering result and evaluate the effect of the large language model in the power field based on the answering accuracy; The evaluation question bank of the large language model in the power field is constructed through investigating various application scenarios in the power field and testing by multiple models and artificially.

8. A computer device, comprising: Including: One or more processors; The processor is configured to store one or more programs; When the one or more programs are executed by the one or more processors, a method for evaluating the effect of a large language model in the power field according to any one of claims 1-6 is implemented.

9. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed, a method for evaluating the effect of a large language model in the power field according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Assessment method and device of large language model, storage medium and computer equipment

    CN117291184A