Comparison, evaluation, selection and measurement method for large language model of electric power system

By dividing the large language model into first-level and second-level capabilities and adopting quantitative evaluation methods, the problem of lack of objectivity and accuracy of traditional evaluation methods is solved, and a comprehensive and scientific evaluation of the performance of the large language model is achieved.

CN119962527APending Publication Date: 2025-05-09CHINA POWER ENG CONSULTING GRP CORP EAST CHINA ELECTRIC POWER DESIGN INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311475189.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-07
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

It is difficult for the prior art to choose the large language model that is most suitable for the application of power system knowledge service through objective and accurate methods. The traditional method is based on human subjective feelings and it is difficult to ensure the objectivity and accuracy of the results.

Method used

By dividing the large language model into first-level and second-level capabilities, and matching the weight coefficients and adjustment coefficients for each ability, the scores of each first-level and overall ability are calculated using quantitative and universal evaluation methods.

Benefits of technology

It realizes a comprehensive, scientific and quantifiable evaluation of the performance of large language models, provides more objective and accurate evaluation results, and can be applied in different fields and scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962527A_ABST
    Figure CN119962527A_ABST
Patent Text Reader

Abstract

The invention relates to the field of electric power systems, and discloses a comparison, evaluation, selection and measurement method for a large language model of an electric power system, which can find the large language model most suitable for knowledge service application of the electric power system through a quantitative and universal evaluation method. The method comprises the following steps: dividing a large language model of the power system into a plurality of first-level capabilities, and matching a weight coefficient for each first-level capability; and based on the plurality of first-level capabilities, dividing each first-level capability into second-level capabilities of multiple dimensions, and matching a corresponding calculation method and an adjustment coefficient for each second-level capability. And calculating the total score of the first-level capability, wherein the score of the first-level capability is equal to sigma score of the second-level capability * the adjustment coefficient. And calculating the total score of the overall capability of the large language model, wherein the total score of the overall capability of the large language model is equal to the score of the sigma first-level capability * the weight coefficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of power systems, and in particular to a comparison, evaluation and selection method for a large language model of a power system. Background Art

[0002] This section is intended to provide a background or context to the embodiments of the present application as recited in the claims. The description herein is not admitted to be prior art as disclosed by virtue of its inclusion in this section.

[0003] In recent years, with the continuous development and popularization of artificial intelligence technology, large models have been used more and more widely in various fields. Among them, the field of natural language processing is the most successful field for the application of large models, and it is also one of the most concerned fields at present. In the field of natural language processing, the development of large models is mainly reflected in two aspects: first, the scale of the model is getting larger and larger, and second, the complexity of the model is getting higher and higher. These models are usually composed of billions or even tens of billions of parameters, and require large-scale computing resources for training. At the same time, these models also need to solve more complex problems, such as semantic understanding and sentiment analysis. The development of large models has brought many benefits. First, they can better handle the complex structure and semantic information in natural language, thereby improving the accuracy and efficiency of natural language processing tasks. At the same time, large models can discover some new laws and knowledge by learning massive data, bringing more inspiration and innovation to mankind. Large models can also be applied to various practical scenarios, such as intelligent customer service, machine translation, speech recognition, etc.

[0004] In the field of power system knowledge services, large language models are an important tool and application method. Large language models can automatically generate and understand natural language, thereby providing intelligent knowledge services for power systems. However, due to the wide variety of large language models and the differences in performance of different models, selecting the most suitable large language model for power system knowledge service applications has become one of the important technical issues facing the industry. Traditional evaluation methods are usually based on the overall subjective feelings of human use, which makes it difficult to guarantee the objectivity and accuracy of the evaluation results. In addition, due to the differences in power system knowledge and application scenarios in different fields, a single evaluation method and indicator system often cannot meet the needs of different applications and scenarios.

[0005] There are already some large language model selection and evaluation methods. For example, the large model is scored from various dimensions, and each dimension corresponds to an evaluation data set, which contains several questions. Each question is given 1 to 5 points based on the quality of the large model's response. The scores of all questions in the evaluation set are accumulated and normalized to 100 points, which is the final score. There is also a method of evaluating the ability of large models purely through subjective feelings during human use. This method usually requires inviting experts or experienced personnel to evaluate and improve the model. Summary of the invention

[0006] The purpose of this application is to provide a comparative evaluation method for large language models of power systems, which can find the large language model that is most suitable for power system knowledge service applications through a quantitative and universal evaluation method.

[0007] The present application discloses a method for comparing and evaluating a large language model of an electric power system, comprising:

[0008] Divide the large language model of the power system into a plurality of first-level capabilities, and match a weight coefficient for each of the first-level capabilities;

[0009] Based on the plurality of first-level capabilities, each first-level capability is divided into second-level capabilities of multiple dimensions, a corresponding calculation method is matched for each second-level capability to obtain a score for each second-level capability, and a corresponding adjustment coefficient is matched for each second-level capability;

[0010] The score of the first-level ability is calculated, the score of the first-level ability=∑the score of the second-level ability×the adjustment coefficient;

[0011] The total score of the overall ability of the large language model is calculated, where the total score of the overall ability of the large language model=∑the score of the first-level ability×the weight coefficient.

[0012] In a preferred example, the first-level capabilities include basic capabilities, summary capabilities, translation capabilities, reasoning capabilities, imitation capabilities, learning capabilities, real-time capabilities, and multimodal capabilities;

[0013] The weight coefficient of the basic ability is 0.05, the weight coefficient of the summary ability is 0.1, the weight coefficient of the translation ability is 0.04, the weight coefficient of the reasoning ability is 0.2, the weight coefficient of the imitation ability is 0.02, the weight coefficient of the learning ability is 0.04, the weight coefficient of the real-time is 0.2, and the weight coefficient of the multimodal ability is 0.1.

[0014] In a preferred example, the secondary capabilities of the basic capabilities include: model scale, model data, and model efficiency;

[0015] The adjustment coefficient of the model scale is 0.6, the adjustment coefficient of the model data is 0.2, and the adjustment coefficient of the model efficiency is 0.2;

[0016] The model scale score is calculated as follows: 50 points for a scale of less than 100 billion; 60 points for a scale of 100 billion to 200 billion; 70 points for a scale of 200 billion to 300 billion; 80 points for a scale of 300 billion to 400 billion; 90 points for a scale of 400 billion to 500 billion; 100 points for a scale of 500 billion or more;

[0017] The calculation method of the model data score is as follows: the data set used by the model covers training data of multiple fields and multiple tasks, and the score is 70 points. In subsequent iterations, the score increases by 5 points for each additional field of training data, and the full score is 100 points.

[0018] The model efficiency score is calculated as follows: a single training time of 3 months, 4 weeks per month, is 70 points. For every week of training time reduced, the score increases by 5 points. The full score is 100 points. For every week of training time increased, the score decreases by 5 points.

[0019] In a preferred embodiment, the secondary capabilities of the summary capabilities include: relevance of the summary, consistency of the summary, fluency of the summary, and coherence of the summary;

[0020] The adjustment factor for the relevance of the abstract is 0.4, the adjustment factor for the consistency of the abstract is 0.3, the adjustment factor for the fluency of the abstract is 0.1, and the adjustment factor for the coherence of the abstract is 0.2;

[0021] The calculation method of the relevance score of the summary is: if the summary output by the large language model contains a% of important keywords in the text, the score is a;

[0022] The calculation method of the consistency score of the summary is: the summary output by the large language model reflects b% of the purpose in the text, and the score is b points;

[0023] The calculation method of the fluency score of the abstract is as follows: the number and position of the subject, predicate, object and other components in the sentence, the correct rate of the noun, verb, adjective and other parts of speech in the sentence is c%, and the score is c points;

[0024] The calculation method of the coherence score of the summary is: the logical coherence rate between each sentence in the summary output by the large language model is d%, and the score is d points.

[0025] In a preferred embodiment, the secondary capabilities of the translation capability classification include: translation accuracy, translation fluency, translation naturalness, and translation diversity;

[0026] The adjustment coefficient for the accuracy of the translation is 0.4, the adjustment coefficient for the fluency of the translation is 0.3, the adjustment coefficient for the naturalness of the translation is 0.1, and the adjustment coefficient for the diversity of the translation is 0.2;

[0027] The calculation method of the accuracy score of the translation is: the accuracy of each word in the translation result generated by the large language model is e%, and the score is e points;

[0028] The calculation method of the translation fluency score is as follows: the grammatical accuracy of each sentence in the translation result generated by the large language model is f%, and the score is f points;

[0029] The naturalness score of the translation is calculated by scoring whether the translation result generated by the large language model takes cultural differences into consideration, with the score ranging from 0 to 100.

[0030] The calculation method of the translation diversity score is: 60 points for the ability to translate into both Chinese and English, 5 points for each additional language, and 0 points if the ability to translate into both Chinese and English is not available.

[0031] In a preferred embodiment, the secondary abilities of the reasoning ability classification include: rationality of reasoning, validity of reasoning, persuasiveness of reasoning, and diversity of reasoning;

[0032] The adjustment coefficient for the rationality of the reasoning is 0.5, the adjustment coefficient for the validity of the reasoning is 0.2, the adjustment coefficient for the persuasiveness of the reasoning is 0.1, and the adjustment coefficient for the diversity of the reasoning is 0.2;

[0033] The rationality score of the reasoning is calculated as follows: if the reasoning is based on clear facts and evidence in the text and is in line with logic and common sense, it is 100 points; if the reasoning is not based on facts and evidence in the text and is completely inconsistent with logic and common sense, it is 0 points, and the score is 0-100 points;

[0034] The calculation method of the validity score of the reasoning is: if the reasoning can achieve the expected purpose and can strengthen or weaken the argument of the text according to logic and common sense, it will be 100 points; if the reasoning cannot achieve the expected purpose at all and cannot strengthen or weaken the argument of the text according to logic and common sense, it will be 0 points, and the score range is 0-100 points;

[0035] The persuasiveness score of the reasoning is calculated as follows: if the reasoning uses appropriate language and rhetoric and can influence and persuade the reader, 100 points; if the reasoning cannot influence and persuade the reader at all, 0 points, and the score is 0-100 points;

[0036] The calculation method of the reasoning diversity score is: the large language model can perform reasoning tasks in multiple fields and multiple tasks, with a score of 70 points. The score increases by 10 points for each additional field of reasoning ability, and increases by 5 points for each additional task of reasoning ability. The full score is 100 points. If the large language model is completely unable to handle reasoning in different fields and tasks, it will be scored 0 points.

[0037] In a preferred embodiment, the secondary capabilities of the imitation capability classification include: similarity of imitation, creativity of imitation, adaptability of imitation and diversity of imitation;

[0038] The adjustment coefficient of the similarity of the imitation is 0.5, the adjustment coefficient of the creativity of the imitation is 0.2, the adjustment coefficient of the adaptability of the imitation is 0.1, and the adjustment coefficient of the diversity of the imitation is 0.2;

[0039] The similarity score of the imitation is calculated as follows: if the imitation completely maintains the characteristics and features of the text and is close to the style and tone of the text, it is scored as 100 points; if it is completely unrelated to the imitated content, it is scored as 0 points;

[0040] The calculation method of the creativity score of the imitation is as follows: 100 points are given if the imitation completely maintains the characteristics and features of the text and combines the reasoning function to output accurate content similar to the style and tone of the text; 0 points are given if it is completely unrelated to the imitated content;

[0041] The calculation method of the adaptability score of the imitation is as follows: the large language model adjusts its output according to different inputs and prompts to adapt to different fields and tasks, and the score is 70 points. For each additional field of imitation ability, the score increases by 10 points, and for each additional task of imitation ability, the score increases by 5 points. The full score is 100 points; if it is completely unable to handle the imitation of different fields and tasks, it will be 0 points;

[0042] The calculation method of the imitation diversity score is as follows: the large language model processes different text types and sources and simulates different authors or characters, with a score of 70 points. For each additional text type and source of imitation ability, the score increases by 10 points, with a full score of 100 points; if it is completely unable to process different text types and sources, it will be scored 0 points.

[0043] In a preferred embodiment, the secondary capabilities of the learning capability classification include: learning depth, learning breadth, learning persistence and learning adaptability;

[0044] The adjustment coefficient of the depth of learning is 0.3, the adjustment coefficient of the breadth of learning is 0.3, the adjustment coefficient of the persistence of learning is 0.3, and the adjustment coefficient of the adaptability of learning is 0.1;

[0045] The depth score of the learning is calculated as follows: 10 points are awarded for each detail and implicit information captured in the text, with the highest score being 100 points;

[0046] The calculation method of the breadth of learning score is: starting from 100 points, 10 points will be deducted for each missing detail and implicit information in the text, and the lowest score is 0 points;

[0047] The calculation method of the learning persistence score is as follows: if the knowledge and information in the text can still be retained and recalled after 5 conversations, it will be 70 points; for each additional conversation, the score will increase by 5 points, up to a maximum of 100 points; if the knowledge and information in the text are ignored in this conversation, it will be 0 points;

[0048] The adaptive score of the learning is calculated as follows: 100 points are given if the learning system can fully adjust its output according to different inputs and prompts and adapt to different fields and tasks; 0 points are given if the learning system cannot adjust its output according to different inputs and prompts and cannot adapt to different fields and tasks.

[0049] In a preferred example, the second-level capabilities of the real-time classification include: real-time accuracy and real-time relevance;

[0050] The adjustment coefficient of the real-time accuracy is 0.5, and the adjustment coefficient of the real-time correlation is 0.6;

[0051] The real-time accuracy score is calculated as follows: the current date and time is asked, including the five elements of year, month, day, hour, and minute, and each correct answer to one element is counted as 20 points;

[0052] The real-time relevance score is calculated as follows: the text of the question contains date and time information, and on this basis, the large language model is used to perform 5 date calculations, including the three elements of year, month, and day, and each correct answer is counted as 20 points.

[0053] In a preferred example, the secondary capabilities of the multimodal capability division include: multimodal consistency, multimodal richness, multimodal creativity, and multimodal diversity;

[0054] The adjustment coefficient of the consistency of the multimodality is 0.3, the adjustment coefficient of the richness of the multimodality is 0.3, the adjustment coefficient of the creativity of the multimodality is 0.1, and the adjustment coefficient of the diversity of the multimodality is 0.3;

[0055] The multimodal consistency score is calculated as follows: if the extracted content from other types of content other than the text content is completely consistent with the purpose of the text or file, it is scored 100 points, and if the extracted content is completely unrelated to the text or file, it is scored 0 points;

[0056] The calculation method of the multimodal richness score is as follows: 70 points are given for the ability to interact with text within the interface, and 10 points are added for each additional type of file, with the maximum score being 100 points;

[0057] The calculation method of the multimodal creativity score is as follows: 70 points are given for the ability to generate images based on text content, and 10 points are added for each additional type of file, with the maximum score being 100 points;

[0058] The calculation method of the multimodal diversity score is: 100 points for being fully able to handle different fields and tasks and being able to adapt to different languages ​​and regions; 0 points for being completely unable to handle different fields and tasks and being unable to adapt to different languages ​​and regions.

[0059] In the implementation mode of the present application, by dividing the large language model into first-level capabilities, the project setting of the first-level capabilities covers the 8 important capability requirements of the large language model in the application of power system knowledge services, and each first-level capability is divided into corresponding second-level capabilities of multiple dimensions. The evaluation items of the second-level capabilities are formulated according to the scenarios of the large language model in the application of power system knowledge services, and the corresponding weight coefficient is matched for each first-level capability. The weight coefficient is set according to the capability focus requirements of the large language model in the application of power system knowledge services, and the corresponding adjustment coefficient is matched for each second-level capability. The adjustment coefficient is set according to the requirements of the large language model in the application of power system knowledge services, so that the first-level capability score of the large language model and the total score of the overall capability of the large language model can be calculated. Compared with the traditional evaluation method based entirely on human usage experience, this method of constructing a comprehensive, scientific, and quantifiable multi-dimensional evaluation index system can comprehensively evaluate the performance of different large language models in different aspects, provide more objective and accurate evaluation results, thereby achieving a comprehensive understanding and comparison of the performance of the large language model, and finally finding the most suitable large language model for power system knowledge service applications;

[0060] Furthermore, the evaluation system of the present invention is universal and can be applied to different fields and scenarios;

[0061] Furthermore, the evaluation system of the present invention is targeted and professional, and can be applied to power system knowledge services.

[0062] Each technical feature disclosed in the above invention content, each technical feature disclosed in each implementation mode and example below, and each technical feature disclosed in the accompanying drawings can be freely combined with each other to form various new technical solutions (these technical solutions should be deemed to have been recorded in this specification), unless such combination of technical features is technically infeasible. For example, in one example, feature A+B+C is disclosed, and in another example, feature A+B+D+E is disclosed, and features C and D are equivalent technical means that play the same role. Technically, only one of them can be used, and it is impossible to use them at the same time. Feature E can be combined with feature C technically. Then, the solution of A+B+C+D should not be deemed to have been recorded because it is technically infeasible, while the solution of A+B+C+E should be deemed to have been recorded. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 It is a flow chart of a comparative evaluation and selection method of a large language model of an electric power system according to an embodiment of the present application. DETAILED DESCRIPTION

[0064] In the following description, many technical details are provided to help readers better understand the present application. However, those skilled in the art can understand that the technical solution claimed in the present application can be implemented even without these technical details and various changes and modifications based on the following embodiments.

[0065] Description of some concepts:

[0066] Big Language Models: Big language models are AI models designed to understand and generate human language. They are trained on large amounts of text data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, and more.

[0067] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below in conjunction with the accompanying drawings.

[0068] The first embodiment of the present application relates to a method for comparing and evaluating a large language model of a power system, and the process is as follows: Figure 1 As shown, the method comprises the following steps:

[0069] In step S1, the large language model of the power system is divided into a plurality of first-level capabilities, and a weight coefficient is matched for each first-level capability.

[0070] In step S2, based on multiple first-level capabilities, each first-level capability is divided into second-level capabilities of multiple dimensions, a corresponding calculation method is matched for each second-level capability to obtain the score of each second-level capability, and a corresponding adjustment coefficient is matched for each second-level capability.

[0071] In step S3, the score of the first-level ability is calculated, and the score of the first-level ability=∑the score of the second-level ability×the adjustment coefficient.

[0072] In step S4, the total score of the overall ability of the large language model is calculated, and the total score of the overall ability of the large language model = ∑ first-level ability score × weight coefficient.

[0073] In an optional embodiment, as shown in Table 1, the first-level capabilities include basic capabilities, summary capabilities, translation capabilities, reasoning capabilities, imitation capabilities, learning capabilities, real-time capabilities and multimodal capabilities. The weight coefficient of the basic capabilities is 0.05, the weight coefficient of the summary capabilities is 0.1, the weight coefficient of the translation capabilities is 0.04, the weight coefficient of the reasoning capabilities is 0.2, the weight coefficient of the imitation capabilities is 0.02, the weight coefficient of the learning capabilities is 0.04, the weight coefficient of the real-time capabilities is 0.2, and the weight coefficient of the multimodal capabilities is 0.1.

[0074]

Table 1

[0075] First level capability Weight coefficient Basic Abilities 0.05 Summary ability 0.1 Translation skills 0.04 reasoning ability 0.2 Imitation ability 0.02 Learning ability 0.04 Real-time 0.2 Multimodal capabilities 0.1

[0076] In an optional embodiment, as shown in Table 2, the secondary capabilities of the basic capability division include: model scale, model data, and model efficiency; the adjustment coefficient of the model scale is 0.6, the adjustment coefficient of the model data is 0.2, and the adjustment coefficient of the model efficiency is 0.2.

[0077] The model scale score is calculated as follows: 50 points for a scale of less than 100 billion; 60 points for a scale of 100 billion to 200 billion; 70 points for a scale of 200 billion to 300 billion; 80 points for a scale of 300 billion to 400 billion; 90 points for a scale of 400 billion to 500 billion; and 100 points for a scale of 500 billion and above. The model data score is calculated as follows: the model uses a dataset that covers training data from multiple fields and multiple tasks, with a score of 70 points. For each additional field of training data in subsequent iterations, the score increases by 5 points, with a full score of 100 points. The model efficiency score is calculated as follows: 70 points for a single training time of 3 months and 4 weeks per month. For each week of training time reduction, the score increases by 5 points, with a full score of 100 points. For each week of training time increase, the score decreases by 5 points.

[0078]

Table 2

[0079]

[0080] In an optional embodiment, as shown in Table 3, the secondary capabilities of the summary capability division include: relevance of the summary, consistency of the summary, fluency of the summary and coherence of the summary, the adjustment coefficient of the relevance of the summary is 0.4, the adjustment coefficient of the consistency of the summary is 0.3, the adjustment coefficient of the fluency of the summary is 0.1, and the adjustment coefficient of the coherence of the summary is 0.2.

[0081] The relevance score of the summary is calculated as follows: the summary output by the large language model contains a% of the important keywords in the text, and the score is a. The consistency score of the summary is calculated as follows: the summary output by the large language model reflects b% of the purpose of the text, and the score is b. The fluency score of the summary is calculated as follows: the number and position of the subject, predicate, object and other components in the sentence, and the accuracy of the nouns, verbs, adjectives and other parts of speech in the sentence are c%, and the score is c. The coherence score of the summary is calculated as follows: the logical coherence rate between each sentence in the summary output by the large language model is d%, and the score is d.

[0082]

Table 3

[0083]

[0084] In an optional embodiment, as shown in Table 4, the secondary capabilities of the translation capability division include: translation accuracy, translation fluency, translation naturalness, and translation diversity. The adjustment coefficient of the translation accuracy is 0.4, the adjustment coefficient of the translation fluency is 0.3, the adjustment coefficient of the translation naturalness is 0.1, and the adjustment coefficient of the translation diversity is 0.2.

[0085] The calculation method of the accuracy score of the translation is: the accuracy rate of each word in the translation result generated by the large language model is e%, and the score is e. The calculation method of the fluency score of the translation is: the grammatical accuracy rate of each sentence in the translation result generated by the large language model is f%, and the score is f. The calculation method of the naturalness score of the translation is: whether the translation result generated by the large language model takes cultural differences into consideration for scoring, and the score is 0-100. The calculation method of the diversity score of the translation is: 60 points for the ability to translate into both Chinese and English, and 5 points for each additional language. If the ability to translate into both Chinese and English is not possible, the score is 0.

[0086]

Table 4

[0087]

[0088]

[0089] In an optional embodiment, as shown in Table 5, the secondary capabilities of the reasoning ability classification include: rationality of reasoning, validity of reasoning, persuasiveness of reasoning, and diversity of reasoning. The adjustment coefficient of the rationality of reasoning is 0.5, the adjustment coefficient of the validity of reasoning is 0.2, the adjustment coefficient of the persuasiveness of reasoning is 0.1, and the adjustment coefficient of the diversity of reasoning is 0.2.

[0090] The calculation method of the reasonableness score of the reasoning is: 100 points if the reasoning is based on clear facts and evidence in the text and is in line with logic and common sense; 0 points if the reasoning is not based on facts and evidence in the text and is completely inconsistent with logic and common sense, and the score is 0-100. The calculation method of the effectiveness score of the reasoning is: 100 points if the reasoning can achieve the expected purpose and can strengthen or weaken the argument of the text according to logic and common sense; 0 points if the reasoning cannot achieve the expected purpose at all and cannot strengthen or weaken the argument of the text according to logic and common sense, and the score is 0-100. The calculation method of the persuasiveness score of the reasoning is: 100 points if the reasoning uses appropriate language and rhetoric and can influence and persuade readers; 0 points if the reasoning cannot influence and persuade readers at all, and the score is 0-100. The calculation method for the reasoning diversity score is as follows: if the large language model can perform reasoning tasks in multiple fields and multiple tasks, the score is 70 points. For each additional field of reasoning ability, the score increases by 10 points. For each additional task of reasoning ability, the score increases by 5 points. The full score is 100 points. If it is completely unable to handle reasoning in different fields and tasks, the score is 0 points.

[0091]

Table 5

[0092]

[0093]

[0094] In an optional embodiment, as shown in Table 6, the secondary capabilities of the imitation capability classification include: similarity of imitation, creativity of imitation, adaptability of imitation, and diversity of imitation. The adjustment coefficient of the similarity of imitation is 0.5, the adjustment coefficient of the creativity of imitation is 0.2, the adjustment coefficient of the adaptability of imitation is 0.1, and the adjustment coefficient of the diversity of imitation is 0.2.

[0095] The calculation method of the similarity score of the imitation is: 100 points if the imitation completely maintains the characteristics and features of the text and is close to the style and tone of the text; 0 points if it is completely unrelated to the imitated content. The calculation method of the creativity score of the imitation is: 100 points if the imitation fully maintains the characteristics and features of the text and combines the reasoning function to output accurate content similar to the style and tone of the text; 0 points if it is completely unrelated to the imitated content. The calculation method of the adaptability score of the imitation is: the large language model adjusts its output according to different inputs and prompts to adapt to different fields and tasks, and the score is 70 points. For each additional field of imitation ability, the score increases by 10 points, and for each additional task of imitation ability, the score increases by 5 points. The full score is 100 points; 0 points are completely unable to handle the imitation of different fields and tasks. The calculation method of the imitation diversity score is as follows: the large language model handles different text types and sources and simulates different authors or characters, with a score of 70 points. For each additional text type and source of imitation ability, the score increases by 10 points, with a full score of 100 points; if it is completely unable to handle different text types and sources, it will receive 0 points.

[0096]

Table 6

[0097]

[0098]

[0099] In an optional embodiment, as shown in Table 7, the secondary capabilities of the learning capability classification include: learning depth, learning breadth, learning persistence, and learning adaptability. The adjustment coefficient of learning depth is 0.3, the adjustment coefficient of learning breadth is 0.3, the adjustment coefficient of learning persistence is 0.3, and the adjustment coefficient of learning adaptability is 0.1.

[0100] The depth score of learning is calculated as follows: 10 points are awarded for each detail and implicit information captured in the text, with a maximum score of 100 points. The breadth score of learning is calculated as follows: starting from 100 points, 10 points are deducted for each detail and implicit information in the text that is missed, and the minimum score is 0 points. The persistence score of learning is calculated as follows: 70 points are awarded for being able to retain and recall the knowledge and information in the text after 5 conversations; 5 points are added for each additional conversation, with a maximum score of 100 points; 0 points are awarded for ignoring the knowledge and information in the text during this conversation. The adaptability score of learning is calculated as follows: 100 points are awarded for being able to fully adjust their output according to different inputs and prompts and adapt to different fields and tasks; 0 points are awarded for being completely unable to adjust their output according to different inputs and prompts and unable to adapt to different fields and tasks.

[0101]

Table 7

[0102]

[0103]

[0104] In an optional embodiment, as shown in Table 8, the second-level capabilities of real-time classification include: real-time accuracy and real-time relevance. The adjustment coefficient of real-time accuracy is 0.5, and the adjustment coefficient of real-time relevance is 0.6.

[0105] The real-time accuracy score is calculated as follows: the current date and time of the question is asked, including the five elements of year, month, day, hour, and minute, and each correct answer counts for 20 points. The real-time relevance score is calculated as follows: the text of the question contains date and time information, and on this basis, the large language model is used to perform five date calculations, including the three elements of year, month, and day, and each correct answer counts for 20 points.

[0106]

Table 8

[0107]

[0108]

[0109] In an optional embodiment, as shown in Table 9, the secondary capabilities of the multimodal capability division include: multimodal consistency, multimodal richness, multimodal creativity, and multimodal diversity. The adjustment coefficient of multimodal consistency is 0.3, the adjustment coefficient of multimodal richness is 0.3, the adjustment coefficient of multimodal creativity is 0.1, and the adjustment coefficient of multimodal diversity is 0.3.

[0110] The calculation method for the consistency score of multimodality is: if the content extracted from other types of content except text content is completely consistent with the purpose of the text or file, it will be 100 points; if the extracted content is completely unrelated to the text or file, it will be 0 points. The calculation method for the richness score of multimodality is: 70 points for the ability to interact with text within the interface, and 10 points for each additional type of file, with a maximum score of 100 points. The calculation method for the creativity score of multimodality is: 70 points for the ability to generate images based on text content, and 10 points for each additional type of file, with a maximum score of 100 points. The calculation method for the diversity score of multimodality is: 100 points for being able to fully handle different fields and tasks and adapt to different languages ​​and regions; 0 points for being completely unable to handle different fields and tasks and unable to adapt to different languages ​​and regions.

[0111]

Table 9

[0112]

[0113]

[0114] Accordingly, the embodiments of the present application also provide a computer-readable storage medium, in which computer executable instructions are stored, and when the computer executable instructions are executed by the processor, the various method embodiments of the present application are implemented. Computer-readable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be a computer-readable instruction, a data structure, a module of a program, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable storage media does not include temporary computer-readable media (transitory media), such as modulated data signals and carriers.

[0115] It should be noted that, in the present application, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the term "include", "comprise" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or equipment. In the absence of more restrictions, the elements defined by the sentence "include one" do not exclude the existence of other identical elements in the process, method, article or equipment including the elements. In the present application, if it is mentioned that a certain action is performed according to a certain element, it means at least the meaning of performing the action according to the element, which includes two situations: performing the action only according to the element and performing the action according to the element and other elements. Multiple, multiple, multiple, etc. expressions include 2, 2 times, 2 kinds and more than 2, more than 2 times, more than 2 kinds.

[0116] The serial numbers used in describing the steps of the method do not themselves constitute any limitation on the order of these steps. For example, the step with a larger serial number does not necessarily have to be executed after the step with a smaller serial number. The step with a larger serial number may be executed first and then the step with a smaller serial number. They may also be executed in parallel, as long as this execution order is reasonable for those skilled in the art. For another example, multiple steps with consecutive serial numbers (such as step S1, step S2, step S3, etc.) do not limit other steps that can be executed in between. For example, there may be other steps between step S1 and step S2.

[0117] This specification includes combinations of the various embodiments described herein. Individual references to embodiments (e.g., "one embodiment" or "some embodiments" or "preferred embodiments"); however, these embodiments are not mutually exclusive unless indicated as mutually exclusive or it is clear to a person skilled in the art that they are mutually exclusive. It should be noted that the word "or" is used in this specification in a non-exclusive sense unless the context clearly indicates or requires otherwise.

[0118] All documents mentioned in this specification are considered to be included in the disclosure of this application as a whole, so that they can be used as a basis for modification when necessary. In addition, it should be understood that the above is only a preferred embodiment of this specification and is not intended to limit the scope of protection of this specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of this specification should be included in the scope of protection of one or more embodiments of this specification.

[0119] In some cases, the actions or steps described in the claims may be performed in a different order than in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A method for comparing and evaluating a large language model of a power system, characterized in that: include: Divide the large language model of the power system into a plurality of first-level capabilities, and match a weight coefficient for each of the first-level capabilities; Based on the plurality of first-level capabilities, each first-level capability is divided into second-level capabilities of multiple dimensions, a corresponding calculation method is matched for each second-level capability to obtain a score for each second-level capability, and a corresponding adjustment coefficient is matched for each second-level capability; The score of the first-level ability is calculated, the score of the first-level ability=∑the score of the second-level ability×the adjustment coefficient; The total score of the overall ability of the large language model is calculated, where the total score of the overall ability of the large language model=∑the score of the first-level ability×the weight coefficient.

2. The method for comparing and evaluating a large language model of a power system according to claim 1, characterized in that: The first-level capabilities include basic capabilities, summary capabilities, translation capabilities, reasoning capabilities, imitation capabilities, learning capabilities, real-time capabilities, and multimodal capabilities; The weight coefficient of the basic ability is 0.05, the weight coefficient of the summary ability is 0.1, the weight coefficient of the translation ability is 0.04, the weight coefficient of the reasoning ability is 0.2, the weight coefficient of the imitation ability is 0.02, the weight coefficient of the learning ability is 0.04, the weight coefficient of the real-time is 0.2, and the weight coefficient of the multimodal ability is 0.

1.

3. The method for comparing and evaluating a large language model of a power system according to claim 2, characterized in that: The secondary capabilities of the basic capabilities include: model scale, model data, and model efficiency; The adjustment coefficient of the model scale is 0.6, the adjustment coefficient of the model data is 0.2, and the adjustment coefficient of the model efficiency is 0.2; The model scale score is calculated as follows: 50 points for a scale of less than 100 billion; 60 points for a scale of 100 billion to 200 billion; 70 points for a scale of 200 billion to 300 billion; 80 points for a scale of 300 billion to 400 billion; 90 points for a scale of 400 billion to 500 billion; 100 points for a scale of 500 billion or more; The calculation method of the model data score is as follows: the data set used by the model covers training data of multiple fields and multiple tasks, and the score is 70 points. In subsequent iterations, the score increases by 5 points for each additional field of training data, and the full score is 100 points. The model efficiency score is calculated as follows: a single training time of 3 months, 4 weeks per month, is 70 points. For every week of training time reduced, the score increases by 5 points. The full score is 100 points. For every week of training time increased, the score decreases by 5 points.

4. The method for comparing and evaluating a large language model of a power system according to claim 2, characterized in that: The secondary abilities of the summary ability classification include: relevance of the summary, consistency of the summary, fluency of the summary, and coherence of the summary; The adjustment factor for the relevance of the abstract is 0.4, the adjustment factor for the consistency of the abstract is 0.3, the adjustment factor for the fluency of the abstract is 0.1, and the adjustment factor for the coherence of the abstract is 0.2; The calculation method of the relevance score of the summary is: if the summary output by the large language model contains a% of important keywords in the text, the score is a; The calculation method of the consistency score of the summary is: the summary output by the large language model reflects b% of the purpose in the text, and the score is b points; The calculation method of the fluency score of the abstract is as follows: the number and position of the subject, predicate, object and other components in the sentence, the correct rate of the noun, verb, adjective and other parts of speech in the sentence is c%, and the score is c points; The calculation method of the coherence score of the summary is: the logical coherence rate between each sentence in the summary output by the large language model is d%, and the score is d points.

5. The method for comparing and evaluating a large language model of a power system according to claim 2, characterized in that: The secondary abilities of the translation ability classification include: translation accuracy, translation fluency, translation naturalness and translation diversity; The adjustment coefficient for the accuracy of the translation is 0.4, the adjustment coefficient for the fluency of the translation is 0.3, the adjustment coefficient for the naturalness of the translation is 0.1, and the adjustment coefficient for the diversity of the translation is 0.2; The calculation method of the accuracy score of the translation is: the accuracy of each word in the translation result generated by the large language model is e%, and the score is e points; The calculation method of the translation fluency score is as follows: the grammatical accuracy of each sentence in the translation result generated by the large language model is f%, and the score is f points; The naturalness score of the translation is calculated by scoring whether the translation result generated by the large language model takes cultural differences into consideration, with the score ranging from 0 to 100. The calculation method of the translation diversity score is: 60 points for the ability to translate into both Chinese and English, 5 points for each additional language, and 0 points if the ability to translate into both Chinese and English is not available.

6. The method for comparing and evaluating a large language model of a power system according to claim 2, characterized in that: The secondary abilities of the reasoning ability classification include: rationality of reasoning, validity of reasoning, persuasiveness of reasoning and diversity of reasoning; The adjustment coefficient for the rationality of the reasoning is 0.5, the adjustment coefficient for the validity of the reasoning is 0.2, the adjustment coefficient for the persuasiveness of the reasoning is 0.1, and the adjustment coefficient for the diversity of the reasoning is 0.2; The rationality score of the reasoning is calculated as follows: if the reasoning is based on clear facts and evidence in the text and is in line with logic and common sense, it is 100 points; if the reasoning is not based on facts and evidence in the text and is completely inconsistent with logic and common sense, it is 0 points, and the score is 0-100 points; The calculation method of the validity score of the reasoning is: if the reasoning can achieve the expected purpose and can strengthen or weaken the argument of the text according to logic and common sense, it will be 100 points; if the reasoning cannot achieve the expected purpose at all and cannot strengthen or weaken the argument of the text according to logic and common sense, it will be 0 points, and the score range is 0-100 points; The persuasiveness score of the reasoning is calculated as follows: if the reasoning uses appropriate language and rhetoric and can influence and persuade the reader, 100 points; if the reasoning cannot influence and persuade the reader at all, 0 points, and the score is 0-100 points; The calculation method of the reasoning diversity score is: the large language model can perform reasoning tasks in multiple fields and multiple tasks, with a score of 70 points. The score increases by 10 points for each additional field of reasoning ability, and increases by 5 points for each additional task of reasoning ability. The full score is 100 points. If the large language model is completely unable to handle reasoning in different fields and tasks, it will be scored 0 points.

7. The method for comparing and evaluating a large language model of a power system according to claim 2, characterized in that: The second-level abilities of the imitation ability classification include: similarity of imitation, creativity of imitation, adaptability of imitation and diversity of imitation; The adjustment coefficient of the similarity of the imitation is 0.5, the adjustment coefficient of the creativity of the imitation is 0.2, the adjustment coefficient of the adaptability of the imitation is 0.1, and the adjustment coefficient of the diversity of the imitation is 0.2; The similarity score of the imitation is calculated as follows: if the imitation completely maintains the characteristics and features of the text and is close to the style and tone of the text, it is scored as 100 points; if it is completely unrelated to the imitated content, it is scored as 0 points; The calculation method of the creativity score of the imitation is as follows: 100 points are given if the imitation completely maintains the characteristics and features of the text and combines the reasoning function to output accurate content similar to the style and tone of the text; 0 points are given if it is completely unrelated to the imitated content; The calculation method of the adaptability score of the imitation is as follows: the large language model adjusts its output according to different inputs and prompts to adapt to different fields and tasks, and the score is 70 points. For each additional field of imitation ability, the score increases by 10 points, and for each additional task of imitation ability, the score increases by 5 points. The full score is 100 points; if it is completely unable to handle the imitation of different fields and tasks, it will be 0 points; The calculation method of the imitation diversity score is as follows: the large language model processes different text types and sources and simulates different authors or characters, with a score of 70 points. For each additional text type and source of imitation ability, the score increases by 10 points, with a full score of 100 points; if it is completely unable to process different text types and sources, it will be scored 0 points.

8. The method for comparing and evaluating a large language model of a power system according to claim 2, characterized in that: The secondary abilities of the learning ability classification include: learning depth, learning breadth, learning persistence and learning adaptability; The adjustment coefficient of the depth of learning is 0.3, the adjustment coefficient of the breadth of learning is 0.3, the adjustment coefficient of the persistence of learning is 0.3, and the adjustment coefficient of the adaptability of learning is 0.1; The depth score of the learning is calculated as follows: 10 points are awarded for each detail and implicit information captured in the text, with the highest score being 100 points; The calculation method of the breadth of learning score is: starting from 100 points, 10 points will be deducted for each missing detail and implicit information in the text, and the lowest score is 0 points; The calculation method of the learning persistence score is as follows: if the knowledge and information in the text can still be retained and recalled after 5 conversations, it will be 70 points; for each additional conversation, the score will increase by 5 points, up to a maximum of 100 points; if the knowledge and information in the text are ignored in this conversation, it will be 0 points; The adaptive score of the learning is calculated as follows: 100 points are given if the learning system can fully adjust its output according to different inputs and prompts and adapt to different fields and tasks; 0 points are given if the learning system cannot adjust its output according to different inputs and prompts and cannot adapt to different fields and tasks.

9. The method for comparing and evaluating a large language model of a power system according to claim 2, characterized in that: The second-level capabilities of the real-time division include: real-time accuracy and real-time relevance; The adjustment coefficient of the real-time accuracy is 0.5, and the adjustment coefficient of the real-time correlation is 0.6; The real-time accuracy score is calculated as follows: the current date and time is asked, including the five elements of year, month, day, hour, and minute, and each correct answer to one element is counted as 20 points; The real-time relevance score is calculated as follows: the text of the question contains date and time information, and on this basis, the large language model is used to perform 5 date calculations, including the three elements of year, month, and day, and each correct answer is counted as 20 points.

10. The method for comparing and evaluating a large language model of a power system according to claim 2, characterized in that: The secondary capabilities of the multimodal capability division include: multimodal consistency, multimodal richness, multimodal creativity and multimodal diversity; The adjustment coefficient of the consistency of the multimodality is 0.3, the adjustment coefficient of the richness of the multimodality is 0.3, the adjustment coefficient of the creativity of the multimodality is 0.1, and the adjustment coefficient of the diversity of the multimodality is 0.3; The multimodal consistency score is calculated as follows: if the extracted content from other types of content except the text content is completely consistent with the purpose of the text or file, it will be scored 100 points, and if the extracted content is completely unrelated to the text or file, it will be scored 0 points; The calculation method of the multimodal richness score is as follows: 70 points are given for the ability to interact with text within the interface, and 10 points are added for each additional type of file, with the maximum score being 100 points; The calculation method of the multimodal creativity score is as follows: 70 points are given for the ability to generate images based on text content, and 10 points are added for each additional type of file, with the maximum score being 100 points; The calculation method of the multimodal diversity score is: 100 points for being fully able to handle different fields and tasks and being able to adapt to different languages ​​and regions; 0 points for being completely unable to handle different fields and tasks and being unable to adapt to different languages ​​and regions.