A method and apparatus for evaluating the performance of model inference in power systems.
By constructing multiple evaluation indicators and test questions, and combining them with practical application scenarios to evaluate the reasoning performance of large language models, the problem of model evaluation in power systems has been solved, achieving effective evaluation of logic and interpretability, and improving the reliability and efficiency of the power grid.
Patent Information
- Application Number
- CN202411861366.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-12-17
AI Technical Summary
How to effectively evaluate the reasoning performance of large language models in power systems, especially in terms of logic and interpretability, taking into account the complexity and security requirements of power systems.
Multiple evaluation indicators are constructed to form a target evaluation indicator set. Evaluation dimensions are created in combination with actual application scenarios. Test questions are constructed and the model's reasoning effect is evaluated using preset correct answers. Evaluation is carried out through logical reasoning tests.
It achieves effective evaluation in terms of logic and interpretability, provides a comprehensive and objective assessment of the model's reasoning performance, and improves the reliability and efficiency of the power grid.
Smart Images

Figure CN119903311B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power system technology, and in particular to a method and apparatus for evaluating the performance of model reasoning in power systems. Background Technology
[0002] Large Language Models (LLMs) are an important component of the field of artificial intelligence, playing a key role in natural language processing. In the field of power system technology, LLMs can process and analyze massive amounts of power data, providing decision support for power grid operation, optimizing energy allocation, predicting and preventing potential faults, and thus improving the reliability and efficiency of the power grid.
[0003] However, the complexity of power systems and the extremely high requirements for safety present unique challenges to LLM. How to measure the inference performance of LLM in practical applications is an urgent problem to be solved. Summary of the Invention
[0004] In view of this, this application provides a method and apparatus for UAV flight path planning. The main purpose is to use logical reasoning tests implemented by constructing test questions to effectively evaluate the reasoning ability of the model in terms of logic and interpretability, in combination with the actual application scenario of the model, thereby providing a solution for effectively evaluating the reasoning effect of the model.
[0005] To achieve the above objectives, this application mainly provides the following technical solutions:
[0006] The first aspect of this application provides a method for evaluating the performance of model inference on a power system, the method comprising:
[0007] Multiple evaluation indicators are constructed for the model, and each evaluation indicator corresponds to a first evaluation dimension. There is no information overlap between the first evaluation dimensions corresponding to different evaluation indicators. The model is a large language model that performs inference operations on power data in a power system.
[0008] At least one evaluation indicator is selected from the plurality of evaluation indicators to form a target evaluation indicator set. The target evaluation indicator set corresponds to at least one second evaluation dimension, which is an evaluation dimension created in combination with the actual application scenario.
[0009] Test questions are constructed for the set of target evaluation indicators. The test questions are used to test the reasoning performance of the model on the second evaluation dimension. Each test question corresponds to a preset correct answer.
[0010] Obtain the target power data required for the test questions;
[0011] The target power data is processed using the model to output inference result data;
[0012] The reasoning results data are evaluated using the preset correct answers corresponding to the test questions, so as to output an evaluation result of the reasoning effect of the model.
[0013] A second aspect of this application provides an evaluation apparatus for model inference performance on a power system, the apparatus comprising:
[0014] The first construction unit is used to construct multiple evaluation indicators for the model. Each evaluation indicator corresponds to a first evaluation dimension. There is no information overlap between the first evaluation dimensions corresponding to different evaluation indicators. The model is a large language model that performs inference operations on power data in a power system.
[0015] The selection unit is used to select at least one evaluation indicator from the plurality of evaluation indicators to form a target evaluation indicator set. The target evaluation indicator set corresponds to at least one second evaluation dimension, which is an evaluation dimension created in combination with the actual application scenario.
[0016] The second construction unit is used to construct test questions for the target evaluation index set. The test questions are used to test the reasoning effect of the model on the second evaluation dimension. Each test question corresponds to a preset correct answer.
[0017] The first acquisition unit is used to acquire the target power data required for the test question;
[0018] The processing unit is used to process the target power data using the model to output inference result data;
[0019] An evaluation unit is used to evaluate the reasoning result data by using the preset correct answers corresponding to the test questions, so as to output an evaluation result of the reasoning effect of the model.
[0020] A third aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described above for evaluating the performance of model inference on a power system.
[0021] A fourth aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the method described above for evaluating the performance of model inference on a power system.
[0022] By employing the above-described technical solution, the technical solution provided in this application has at least the following advantages:
[0023] This application provides a method and apparatus for evaluating the performance of model inference in power systems. For large language models that perform inference operations on power data in power systems, multiple evaluation indicators are first constructed for the model. Since evaluating the model using a single evaluation indicator or simply by superimposing evaluation indicators is not comprehensive or objective, this application combines the evaluation indicators to obtain a target evaluation indicator set. Then, evaluation dimensions are created for this target evaluation indicator set based on the actual application scenario. Based on these evaluation dimensions, test questions are constructed for the target evaluation indicator set, and each test question has a pre-set correct answer. The target power data required for the test questions is input into the model for processing, and the inference result data given by the model is output. By comparing the inference result data with the pre-set correct answer of each test question and performing evaluation, the evaluation result of the model's inference performance is obtained.
[0024] In summary, this application utilizes the logical reasoning test implemented by constructing test questions. This test can effectively evaluate the model's reasoning ability in terms of logic and explanatory power by combining the model with the actual application scenario, thus providing a solution for effectively evaluating the model's reasoning performance.
[0025] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description
[0026] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present application. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0027] Figure 1 A flowchart illustrating a method for evaluating the performance of model inference in a power system, provided as an embodiment of this application;
[0028] Figure 2 A block diagram illustrating the composition of an evaluation device for model inference performance in a power system, provided in an embodiment of this application;
[0029] Figure 3 A block diagram of another device for evaluating the performance of model inference on a power system, provided as an embodiment of this application. Detailed Implementation
[0030] The following describes exemplary embodiments of the present application in more detail with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0031] This application provides a method for evaluating the performance of model inference in power systems, such as... Figure 1 As shown, the following specific steps are provided in this embodiment of the application:
[0032] 101. Construct multiple evaluation indicators for the model, each of which corresponds to a first evaluation dimension. There is no information overlap between the first evaluation dimensions corresponding to different evaluation indicators. The model is a large language model that performs inference operations on power data in the power system.
[0033] In this embodiment, a Large Language Model (LLM) is used to process power data. Power data refers to various data related to power supply, use, and management generated, collected, and processed through power facilities. It mainly includes power generation data, power consumption data, quality data, equipment data, and transaction data. This data is collected and processed through acquisition and transmission systems, providing a foundation for power grid monitoring, control, and optimization.
[0034] For example, electricity data may typically have, but is not limited to, the following characteristics:
[0035] (1) Real-time: Real-time updates of power data are crucial for the safe and stable operation of the power grid.
[0036] (2) Diversity: Power data comes from various stages such as power generation, transmission, transformation, distribution, consumption and dispatch, and the data types are diverse.
[0037] (3) High value density: Power data contains a wealth of information, which is of great value for the optimization and management of power systems.
[0038] (4) Accuracy: The accuracy of power data is directly related to the operating efficiency and security of the power grid.
[0039] Furthermore, the application of large language models in power processing can include, but is not limited to, the following:
[0040] (1) Data cleaning and integration: LLM has powerful text understanding and reasoning capabilities, which can automatically identify and correct errors and outliers in power data. Through contextual learning, LLM can understand and integrate information from different data sources to form a unified and accurate dataset.
[0041] (2) Load Forecasting and Power Generation Forecasting: LLM can perform more accurate load and power generation forecasting based on multi-dimensional information such as historical electricity data, meteorological information, and socio-economic activities. In terms of load forecasting, LLM's small-sample learning capability helps solve the data scarcity problem in new user and low-frequency electricity consumption scenarios. For medium- and long-term load forecasting, LLM can solve the problems of feature selection and non-electricity consumption data generation. In terms of power generation forecasting, LLM can consider multiple factors (such as weather changes and equipment status) to improve forecast accuracy.
[0042] (3) Power System Planning: In terms of uncertainty simulation, LLM can generate intuitive simulation results based on knowledge, saving the sampling calculation process and achieving approximate results, while realizing power system risk assessment based on uncertainty simulation results. In terms of planning scenario generation, LLM can automatically generate typical planning scenarios through natural language descriptions and adjust and optimize specific scenarios.
[0043] In terms of planning scheme optimization, LLM can be applied to the automatic interpretation of optimization schemes, generating natural language descriptions of the optimal solutions. Simultaneously, LLM can quickly generate corresponding planning models and constraints, reducing the workload of power system planning.
[0044] (4) Power System Dispatch: LLM can assist power system dispatch in many aspects, such as dispatcher experience extraction, dispatch modeling, dispatch decision-making, operation execution, and power system situation awareness. The generalization and natural language processing capabilities of LLM can help dispatchers quickly learn dispatch knowledge, improve dispatch efficiency and accuracy, thereby achieving economical, safe and low-carbon operation of the power system.
[0045] (5) Fault Diagnosis and Recovery: LLM can quickly identify fault types and locations based on fault data and a professional knowledge base, and generate corresponding fault handling solutions. In the fault diagnosis of power transmission and distribution systems, the versatility of LLM is expected to enable simultaneous diagnosis of multiple types of faults. In terms of system recovery, LLM can conduct multi-dimensional safety comprehensive assessments of recovery plans and quickly generate recovery plans.
[0046] (6) Electricity Market Research: In electricity market modeling, LLM can construct more accurate models of trading behavior even with few or no samples. In electricity market decision-making, LLM has the ability to autonomously generate strategies, enabling it to cope with rapidly changing market environments. In electricity market mechanism design, LLM can be combined with mechanism design theory to simulate the behavior of market participants, thereby better ensuring the rationality and practicality of electricity market mechanism design.
[0047] In the field of power system technology, LLM can process and analyze massive amounts of power data, providing decision support for power grid operation, optimizing energy allocation, predicting and preventing potential faults, and thus improving the reliability and efficiency of the power grid.
[0048] In response to the characteristics of power data and the application of LLM in power processing, and considering how to evaluate the inference effect of the model, this application embodiment initially constructs multiple evaluation indicators. For example, a preset input template can be used for construction. Each evaluation indicator corresponds to an evaluation dimension, and there is no information overlap between each evaluation dimension to avoid information overlap between evaluation indicators and the occurrence of evaluation redundancy.
[0049] It should be noted that the evaluation indicators described here are only preliminary ones, and the preliminary evaluation dimensions obtained using these preliminary indicators are preliminary ones. In order to distinguish them from the evaluation dimensions that will be further developed in the future, the embodiments of this application use the terms "first" and "second" to refer to them.
[0050] For example, embodiments of this application may, but are not limited to, construct the following evaluation indicators as preliminary evaluation indicators, which can constitute a comprehensive evaluation indicator system.
[0051] Accuracy refers to the correctness of a model in a reasoning task, that is, the degree to which the model's answer or reasoning result matches the standard answer or reasoning result for a given reasoning question. Specific metrics include comprehension accuracy and error rate. For fixed-answer questions, a pre-labeled test dataset is used, and accuracy is evaluated by comparing the model's output with the standard answer. For open-ended questions, an accuracy score is derived through expert evaluation of the model's output.
[0052] Logic: This refers to the logical rigor of the model's reasoning process, i.e., whether the reasoning process conforms to logical rules and whether it has a reasonable chain of reasoning and logical structure. Specific indicators include logical consistency score and completeness of reasoning chain. Logic analysis tools can be used to examine the reasoning chain and logical structure output by the model, and the logic can be judged by experts reviewing the model's reasoning process.
[0053] Consistency refers to the stability and consistency of a model's reasoning across different contexts; that is, whether the model's reasoning results are consistent for similar reasoning problems. Specific metrics include consistency scores and coefficients of variation. Multiple tests are conducted on the same or similar problems at different times and in different environments to compare the consistency of the reasoning results. The stability of the model's reasoning results is observed by changing the problem background or details.
[0054] Coverage refers to the breadth of different domains, topics, and scenarios that a model can cover in inference tasks, i.e., the model's performance in various inference situations. Specific metrics include domain coverage and scenario adaptability scores. The model is tested and evaluated in different professional domains and topics, and its inference capabilities are tested by simulating different application scenarios.
[0055] Interpretability refers to whether a model can clearly explain its reasoning process and logical reasoning, enabling users to understand and accept the results. Specific metrics include clarity of explanation score and user satisfaction. Model interpretation tools (such as LIME and SHAP) are used to analyze the interpretability of the model output, and users are rated based on their understanding and acceptance of the model's explanation process.
[0056] Transferability refers to a model's ability to be transferred between different inference tasks or domains; that is, whether the model's performance on one inference task can be transferred and applied to other inference tasks. Specific metrics include transfer success rate and adaptability score. Transfer tests are conducted on different tasks and domains to evaluate its performance. Through fine-tuning or with a small amount of training data, the degree of performance improvement on new tasks is observed.
[0057] Efficiency refers to the speed and resource utilization efficiency of the model when performing inference tasks, meaning that the inference result can be obtained within a reasonable time, and the computing resources consumed in the processing are moderate. Specific indicators include response time and resource utilization rate. By actually running the model, its processing speed and resource consumption are measured to analyze the model's computing resource usage during the inference process.
[0058] 102. Select at least one evaluation indicator from multiple evaluation indicators to form a target evaluation indicator set. The target evaluation indicator set corresponds to at least one second evaluation dimension, which is an evaluation dimension created in combination with the actual application scenario.
[0059] The multiple evaluation indicators constructed in 101 above each have different purposes and can be evaluated using different methods, as shown in Table 1 below, which lists the "evaluation indicators", their corresponding descriptions, and evaluation methods.
[0060] Table 1:
[0061]
[0062]
[0063] In the above, each evaluation indicator evaluates the model from different evaluation dimensions, and the sum of these evaluation results equals the sum of the evaluation results on each evaluation indicator. Based on the different divisions of evaluation indicators, this is equivalent to evaluating the model at a certain fine granularity, which has a good breadth of evaluation, but the depth of evaluation is still insufficient. Therefore, the embodiment of this application is to create new evaluation dimensions based on one or more evaluation indicators and actual application scenarios, to obtain multiple new evaluation dimensions corresponding to a target set of evaluation indicators, and then evaluate the model on the new evaluation dimensions.
[0064] For example, taking the evaluation indicators "logic" and "consistency" as examples, "logic" describes whether the reasoning process conforms to logical rules, which can be measured by the rigor of logical rules; "consistency" describes whether the reasoning results are consistent in similar situations. This application's embodiment integrates these two evaluation indicators and selects an application scenario to construct an evaluation dimension. For example, combining the actual application scenario of "electricity market," it tests the reasoning performance of "in a power company's electricity market, if electricity prices increase by 10%, what is the relationship between electricity demand and electricity supply?" while considering both "logic" and "consistency."
[0065] Therefore, the new evaluation dimension is limited to: "real-world application scenario" + "logic" + "consistency". Furthermore, the evaluation focuses on how well the model performs in reasoning within this new evaluation dimension, i.e., assessing the model's reasoning effectiveness.
[0066] 103. Construct test questions for the target evaluation index set. The test questions are used to test the inference effect of the model on the second evaluation dimension. Each test question corresponds to a pre-set correct answer.
[0067] As illustrated in example 102, test questions are constructed based on the new evaluation dimensions corresponding to the set of target evaluation indicators.
[0068] For example: Construct a test question; the question stem is "In the electricity market of a power company, if the electricity price increases by 10%, which of the following is most likely to happen?"; and four options A, B, C, and D, along with the pre-defined correct answer in the question stem.
[0069] Option A: Electricity demand decreases by 5%, but electricity supply increases by 3%;
[0070] Option B: Electricity demand increases by 8%, but electricity supply decreases by 6%;
[0071] Option C: Electricity demand increases by 12%, but electricity supply increases by 10%;
[0072] Option D: Electricity demand decreases by 3%, but electricity supply decreases by 5%;
[0073] The presupposed correct answer to this question is "B".
[0074] The test questions constructed above are those with logical reasoning ability. They are input into the model to test the model's reasoning effectiveness by checking whether the model's output reasoning results are consistent with the preset correct answers. In this way, the model's reasoning ability can be evaluated in terms of logic and interpretability through logical reasoning tests.
[0075] 104. Obtain the target power data required for the test questions.
[0076] 105. Use the model to process the target power data to output inference result data.
[0077] 106. The reasoning results data are evaluated by using the preset correct answers corresponding to the test questions, so as to output the evaluation results of the model's reasoning performance.
[0078] As shown in 104-106, based on a test question exemplified in 103, this application embodiment obtains the relevant power data corresponding to the question stem and four options and inputs them into the model so that the model processes this data information and outputs the reasoning result. When the reasoning result is consistent with the preset correct answer of the test question, it indicates that the model's reasoning effect on this question is very good.
[0079] The test questions exemplified in section 103 are merely exemplary examples of embodiments of this application. To better test the inference performance of the model, embodiments of this application can create multiple new evaluation dimensions for a set of target evaluation indicators for a given set, combined with different real-world scenarios, and create multiple test questions for each new evaluation dimension. That is, a set of target evaluation indicators corresponds to multiple new evaluation dimensions, and multiple test questions are constructed for each new evaluation dimension. In fact, these test questions are all applicable to the set of target evaluation indicators.
[0080] Furthermore, multiple sets of target evaluation indicators can be constructed based on the 101 evaluation indicators, resulting in diverse new evaluation dimensions. Therefore, compared to the 101 evaluation indicators, this application embodiment is equivalent to creating more numerous, diverse, and in-depth new evaluation dimensions, enabling a more comprehensive and in-depth evaluation of the model's reasoning performance.
[0081] The embodiments of this application, as described above (101-106), provide a method for evaluating the performance of model inference in power systems. For large language models that perform inference operations on power data in power systems, multiple evaluation indicators are first constructed for the model. Since evaluating the model using a single evaluation indicator or simply by superimposing evaluation indicators is not comprehensive or objective, this application combines the evaluation indicators to obtain a target evaluation indicator set. Then, evaluation dimensions are created for this target evaluation indicator set based on the actual application scenario. This allows for the construction of test questions for the target evaluation indicator set based on these evaluation dimensions, with each test question having a pre-set correct answer. The target power data required for the test questions is input into the model for processing, and the inference result data given by the model is output. By comparing the inference result data with the pre-set correct answer of each test question and performing an evaluation, the evaluation result of the model's inference performance is obtained.
[0082] Therefore, the logical reasoning test implemented by constructing test questions in this application embodiment can effectively evaluate the model's reasoning ability in terms of logic and interpretability by combining the model's actual application scenario, thus providing a solution for effectively evaluating the model's reasoning performance.
[0083] In some modified embodiments, the set of target evaluation indicators in 102 corresponds to at least one second evaluation dimension. This application embodiment provides a detailed explanation of the construction of this second evaluation dimension, and provides the following detailed implementation steps:
[0084] A1 determines the target evaluation indicators included in the target evaluation indicator set and the first evaluation dimension corresponding to each target evaluation indicator;
[0085] A2 expands the first semantic information on each first evaluation dimension;
[0086] A3 combines the first semantic information obtained from different first evaluation dimensions with different pre-set application scenarios to construct multiple second semantic information.
[0087] A4 constructs a corresponding second evaluation dimension based on each second semantic information;
[0088] A5 uses the second evaluation dimension as the evaluation dimension for the set of target evaluation indicators;
[0089] In this embodiment of the application, the constructed second evaluation dimension is actually a new evaluation dimension that integrates the data information of one or more evaluation indicators in the first evaluation dimension and the actual application scenario.
[0090] Specifically, in this application embodiment, natural language processing can be used to expand semantic information on the first evaluation dimension corresponding to each evaluation indicator. For example, for the evaluation indicator "logic", explanatory information, conditional constraint information, and usage scenario requirement information related to "logic" can be obtained by constructing related words and other related methods, so as to obtain extended semantic information related to "logic".
[0091] When the target evaluation index set includes one target evaluation index, it is sufficient to obtain the extended semantic information related to that target evaluation index. However, when the target evaluation index set includes multiple target evaluation indexes, this application embodiment needs to integrate the extended semantic information corresponding to multiple evaluation indexes and then combine it with different preset application scenarios as the extended semantic information corresponding to each target index set. Based on such semantic information, a second evaluation dimension is created.
[0092] It should be noted that a set of target evaluation indicators can correspond to the creation of one or more second evaluation dimensions, thereby enabling in-depth evaluation of the model based on different real-world scenarios.
[0093] In some modified embodiments, the construction of test questions for the target evaluation index set in step 103 is explained in detail. The embodiments of this application provide the following detailed implementation steps:
[0094] B1 assigns different scores to each of the at least one second evaluation dimension in the target evaluation index set, and the scores are used to measure the importance of the second evaluation dimension.
[0095] B2 allocates the proportion of questions to the second evaluation dimension corresponding to different scores according to the preset weight allocation method;
[0096] B3 constructs test questions based on the proportion of questions in different second evaluation dimensions.
[0097] As in B1-B3, the embodiments of this application are based on the different difficulty of the inference operations performed by the model on different second evaluation dimensions, and different scoring values are configured for each second evaluation dimension. The higher the scoring value, the more important the evaluation is on the corresponding second evaluation dimension.
[0098] Furthermore, after assigning scores, embodiments of this application can also assign different weights to different scores. These weights are used to indicate the number of test questions set, thereby constraining the second evaluation dimension by assigning high or low scores and the number of questions, thereby achieving a better evaluation of the model reasoning in depth.
[0099] In some modified embodiments, the evaluation of the reasoning result data by using the preset correct answers corresponding to the test questions in 106 is explained in detail. The embodiments of this application provide the following detailed implementation steps:
[0100] For each test question, C1 retrieves the inference result corresponding to the test question from the inference result data;
[0101] C2 compares the reasoning result with the preset correct answer corresponding to the test question. If they match, the score for the test question is obtained based on the score value of the second evaluation dimension corresponding to the test question; if they do not match, the test question receives no score.
[0102] C3 compares each test question in the reasoning result data one by one and obtains the cumulative score corresponding to multiple test questions;
[0103] C4 determines the comprehensive evaluation result corresponding to the model by comparing the cumulative score with the preset comprehensive evaluation score range.
[0104] The embodiments of this application are explained in conjunction with the test questions shown in Table 2 below, for C1-C4 above.
[0105] Table 2
[0106]
[0107]
[0108]
[0109] Furthermore, the constructed test questions can also be multiple-choice questions, as shown in Table 3 below.
[0110] Table 3
[0111]
[0112]
[0113]
[0114] For Tables 2 and 3, the questions are scientifically scored according to the test question settings. For example, questions related to policy interpretation and reasonable extrapolation are scored based on the model's reasoning ability based on text comprehension, with 2 points awarded; questions on the analysis and accurate judgment of scenarios such as trading and dispatching in the electricity market are scored based on the model's scenario understanding and reasoning ability, with 3 points awarded; and questions involving power system node flow calculations related to scenario changes are scored based on the model's comprehensive reasoning and calculation abilities, with 4 points awarded. Furthermore, but not limited to, additional operations can be added. Since multiple-choice questions are slightly more difficult than single-choice questions, the weighted score for correctly answered multiple-choice questions can be increased compared to correctly answering single-choice questions. After forming a complete set of test questions, the questions are input into the model for testing, thereby obtaining a more scientific and accurate assessment of the large-scale model's reasoning ability.
[0115] The above 101-106 are actually evaluations of the model's inference performance based on the "model output results". In some modified embodiments, the evaluation of the model can also take into account the three aspects of the data source processed by the model, the model quality, and the application scenario. These can be used as reference factors, and their weights can be increased to comprehensively update the evaluation results of the model's inference performance in the output of 101-106, thereby achieving a more comprehensive evaluation of the model.
[0116] Data sources refer to the data used to build the model and perform inference, including factors such as the quantity, quality, and diversity of the data. The quality of the data sources directly affects the model's inference performance. Model quality evaluates the model's accuracy, stability, and generalization ability, including the modeling methods and training techniques. The quality of the model determines the reliability of the inference results. Output results refer to the results obtained by the model's inference, including the accuracy, completeness, and consistency of the information. The quality of the output results determines the model's practical utility. Application scenarios refer to the model's performance in real-world applications, including inference speed, adaptability, and stability. The quality of the application scenarios affects the model's effectiveness in real-world scenarios. A comprehensive evaluation system needs to consider all four factors, evaluating data quality, model quality, output results, and application scenarios to comprehensively assess the inference performance of large language models and provide a reference for model improvement and optimization.
[0117] For example, in terms of operation, for the historical data sources processed by the model, a first evaluation result of the historical sources is obtained; for the design architecture of the model itself, a second evaluation result of the design architecture is obtained; for the historical application scenarios corresponding to the model, a third evaluation result of the historical application scenarios is obtained; and the first evaluation result, the second evaluation result, and the third evaluation result are combined as reference evaluation factors for the model's reasoning ability on power data.
[0118] Furthermore, as a response to the above Figure 1The implementation of the method shown in this application provides an evaluation device for model inference performance in power systems. This device embodiment corresponds to the foregoing method embodiment. For ease of reading, this device embodiment will not repeat the details of the foregoing method embodiment, but it should be understood that the device in this embodiment can implement all the contents of the foregoing method embodiment. This device is used to achieve a more in-depth evaluation of model inference performance, specifically as follows... Figure 2 As shown, the device includes:
[0119] The first construction unit 21 is used to construct multiple evaluation indicators for the model. Each evaluation indicator corresponds to a first evaluation dimension. There is no information overlap between the first evaluation dimensions corresponding to different evaluation indicators. The model is a large language model that performs inference operations on power data in a power system.
[0120] Selection unit 22 is used to select at least one evaluation indicator from the plurality of evaluation indicators to form a target evaluation indicator set, wherein the target evaluation indicator set corresponds to at least one second evaluation dimension, and the second evaluation dimension is an evaluation dimension created in combination with the actual application scenario.
[0121] The second construction unit 23 is used to construct test questions for the target evaluation index set. The test questions are used to test the reasoning effect of the model on the second evaluation dimension. Each test question corresponds to a preset correct answer.
[0122] The first acquisition unit 24 is used to acquire the target power data required for the test question;
[0123] Processing unit 25 is used to process the target power data using the model to output inference result data;
[0124] Evaluation unit 26 is used to evaluate the reasoning result data by using the preset correct answers corresponding to the test questions, so as to output an evaluation result of the reasoning effect of the model.
[0125] Furthermore, such as Figure 3 As shown, the second building unit 23 includes:
[0126] Configuration module 231 is used to configure different scoring values for each of the second evaluation dimensions in the target evaluation index set corresponding to at least one second evaluation dimension, wherein the scoring value is used to measure the importance of the second evaluation dimension.
[0127] Allocation module 232 is used to allocate the proportion of the number of questions to the second evaluation dimension corresponding to different scores according to a preset weight allocation method;
[0128] Module 233 is used to construct test questions on different second evaluation dimensions according to the proportion of the number of questions.
[0129] Furthermore, such as Figure 3 As shown, the evaluation unit 26 includes:
[0130] The first acquisition module 261 is used to acquire the reasoning result corresponding to each test question from the reasoning result data;
[0131] The comparison module 262 is used to compare the reasoning result with the preset correct answer corresponding to the test question. If they match, the score corresponding to the test question is obtained according to the score value of the second evaluation dimension corresponding to the test question. If they do not match, the test question is not scored.
[0132] The second acquisition module 263 is used to obtain the cumulative score corresponding to multiple test questions after comparing each test question in the reasoning result data one by one;
[0133] The determination module 264 is used to determine the comprehensive evaluation result corresponding to the model by comparing the cumulative score with the preset comprehensive evaluation score range.
[0134] Furthermore, such as Figure 3 As shown, after the set of target evaluation indicators is formed, the device further includes: a third construction unit 27, specifically used for:
[0135] In the set of target evaluation indicators, the target evaluation indicators included and the first evaluation dimension corresponding to each target evaluation indicator are determined;
[0136] Expand the first semantic information on each of the first evaluation dimensions;
[0137] Based on the first semantic information obtained from different first evaluation dimensions, it is combined with different preset application scenarios to construct multiple second semantic information.
[0138] Based on each piece of the second semantic information, a corresponding second evaluation dimension is constructed;
[0139] The second evaluation dimension is used as the evaluation dimension for the target evaluation index set.
[0140] Furthermore, such as Figure 3 As shown, the device further includes:
[0141] The second acquisition unit 28 is used to acquire a first evaluation result of the historical data source processed by the model for the historical data source;
[0142] The second acquisition unit 28 is used to acquire a second evaluation result of the design architecture of the model itself.
[0143] The second acquisition unit 28 is used to acquire a third evaluation result for the historical application scenario corresponding to the model.
[0144] The determining unit 29 is used to integrate the first evaluation result, the second evaluation result, and the third evaluation result as a reference evaluation factor for the model's reasoning ability on power data.
[0145] In summary, the device for evaluating the model inference effect on the power system includes a processor and a memory. The first construction unit, selection unit, second construction unit, first acquisition unit, processing unit and evaluation unit are all stored in the memory as program units, and the processor executes the program units stored in the memory to realize the corresponding functions.
[0146] The processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured. By adjusting kernel parameters, logical reasoning tests implemented using constructed test questions can be conducted. This allows for effective evaluation of the model's reasoning ability in terms of logic and interpretability, taking into account the model's actual application scenarios, thus providing a solution for effectively evaluating the model's reasoning performance.
[0147] This application provides a storage medium storing a program that, when executed by a processor, implements a method for evaluating the performance of model inference on a power system.
[0148] This application provides a processor for running a program, wherein the program executes an evaluation method for model inference performance on a power system.
[0149] This application also provides a computer program product that, when executed on a data processing device, is adapted to perform the steps of an evaluation method for the effectiveness of model reasoning on a power system.
[0150] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0151] In a typical configuration, the device includes one or more processors (CPUs), memory, and a bus. The device may also include input / output interfaces, network interfaces, etc.
[0152] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and memory includes at least one memory chip. Memory is an example of computer-readable media.
[0153] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0154] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0155] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0156] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for evaluating the performance of model inference in power systems, characterized in that, The method includes: Multiple evaluation indicators are constructed for the model, and each evaluation indicator corresponds to a first evaluation dimension. There is no information overlap between the first evaluation dimensions corresponding to different evaluation indicators. The model is a large language model that performs inference operations on power data in a power system. Select at least one evaluation indicator from the plurality of evaluation indicators to form a target evaluation indicator set; In the set of target evaluation indicators, the target evaluation indicators included and the first evaluation dimension corresponding to each target evaluation indicator are determined; Expand the first semantic information on each of the first evaluation dimensions; Based on the first semantic information obtained from different first evaluation dimensions, it is combined with different preset application scenarios to construct multiple second semantic information. Based on each piece of the second semantic information, a corresponding second evaluation dimension is constructed; The second evaluation dimension is used as the evaluation dimension for the target evaluation index set, and the target evaluation index set corresponds to at least one second evaluation dimension. The second evaluation dimension is an evaluation dimension created in combination with the actual application scenario. Test questions are constructed for the set of target evaluation indicators. The test questions are used to test the reasoning performance of the model on the second evaluation dimension. Each test question corresponds to a preset correct answer. Obtain the target power data required for the test questions; The target power data is processed using the model to output inference result data; The reasoning result data is evaluated by using the preset correct answers corresponding to the test questions, so as to output an evaluation result of the reasoning effect of the model; For the historical data sources processed by the model, obtain a first evaluation result for the historical data sources; For the design architecture of the model itself, obtain a second evaluation result of the design architecture; For the historical application scenarios corresponding to the model, obtain the third evaluation result for the historical application scenarios; The combined results of the first, second, and third evaluations are used as reference evaluation factors for the model's reasoning ability on power data.
2. The method according to claim 1, characterized in that, The construction of test questions for the set of target evaluation indicators includes: In the target evaluation index set corresponding to at least one second evaluation dimension, different scoring values are configured for each second evaluation dimension, and the higher or lower scoring values are used to measure the importance of the second evaluation dimension. According to the preset weight allocation method, the proportion of the number of questions is allocated to the second evaluation dimension corresponding to different scores; Based on the proportion of the number of questions, test questions are constructed on different second evaluation dimensions.
3. The method according to claim 2, characterized in that, The step of evaluating the reasoning result data using the preset correct answers corresponding to the test questions to output an evaluation result of the model's reasoning performance includes: For each test question, the reasoning result corresponding to the test question is obtained from the reasoning result data; The reasoning result is compared with the preset correct answer corresponding to the test question. If they match, the score corresponding to the test question is obtained according to the score value of the second evaluation dimension corresponding to the test question. If they do not match, the test question is not scored. After comparing each test question one by one in the reasoning result data, the cumulative score corresponding to multiple test questions is obtained; The comprehensive evaluation result corresponding to the model is determined by comparing the cumulative score with the preset comprehensive evaluation score range.
4. An evaluation device for model inference performance in power systems, characterized in that, The device includes: The first construction unit is used to construct multiple evaluation indicators for the model. Each evaluation indicator corresponds to a first evaluation dimension. There is no information overlap between the first evaluation dimensions corresponding to different evaluation indicators. The model is a large language model that performs inference operations on power data in a power system. The selection unit is used to select at least one evaluation indicator from the plurality of evaluation indicators to form a target evaluation indicator set. The third construction unit is configured to: determine the target evaluation indicators included in the target evaluation indicator set and a first evaluation dimension corresponding to each target evaluation indicator; extend first semantic information on each first evaluation dimension; combine the first semantic information obtained on different first evaluation dimensions with different preset application scenarios to construct multiple second semantic information; construct a corresponding second evaluation dimension according to each second semantic information; and use the second evaluation dimension as the evaluation dimension for the target evaluation indicator set; wherein the target evaluation indicator set corresponds to at least one second evaluation dimension, and the second evaluation dimension is an evaluation dimension created in combination with actual application scenarios; The second construction unit is used to construct test questions for the target evaluation index set. The test questions are used to test the reasoning effect of the model on the second evaluation dimension. Each test question corresponds to a preset correct answer. The first acquisition unit is used to acquire the target power data required for the test question; The processing unit is used to process the target power data using the model to output inference result data; An evaluation unit is used to evaluate the reasoning result data by using the preset correct answers corresponding to the test questions, so as to output an evaluation result of the reasoning effect of the model; The second acquisition unit is used to acquire a first evaluation result of the historical data source processed by the model for the historical data source; The second acquisition unit is used to acquire a second evaluation result of the design architecture of the model itself. The second acquisition unit is used to acquire a third evaluation result for the historical application scenario corresponding to the model. The determining unit is used to integrate the first evaluation result, the second evaluation result, and the third evaluation result as a reference evaluation factor for the model's reasoning ability on power data.
5. The apparatus according to claim 4, characterized in that, The second building unit includes: The configuration module is used to configure different scoring values for each of the second evaluation dimensions in the target evaluation index set corresponding to at least one second evaluation dimension, wherein the scoring value is used to measure the importance of the second evaluation dimension. The allocation module is used to allocate the proportion of the number of questions to the second evaluation dimension corresponding to different scores according to a preset weight allocation method. The module is used to construct test questions on different second evaluation dimensions based on the proportion of the number of questions.
6. The apparatus according to claim 5, characterized in that, The evaluation unit includes: The first acquisition module is used to acquire the reasoning result corresponding to each test question from the reasoning result data; The comparison module is used to compare the reasoning result with the preset correct answer corresponding to the test question. If they match, the score corresponding to the test question is obtained according to the score value of the second evaluation dimension corresponding to the test question; if they do not match, the test question is not scored. The second acquisition module is used to obtain the cumulative score corresponding to multiple test questions after comparing each test question in the reasoning result data one by one; The determination module is used to determine the comprehensive evaluation result corresponding to the model by comparing the cumulative score with a preset comprehensive evaluation score range.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method for evaluating the effectiveness of model reasoning on a power system as described in any one of claims 1-3.
8. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the method for evaluating the effectiveness of model reasoning on a power system as described in any one of claims 1-3.
Citation Information
Patent Citations
Evaluation method and device of natural language processing model and electronic equipment
CN117291169A
Method and system for evaluating effect of large language model in power field
CN118093371A