Comprehensive Assessment Method and Related Devices for Large Language Models

By using dynamic difficulty adjustment and multi-dimensional evaluation methods, the problem of insufficient identification of ability ranges in the evaluation of large language models is solved, and accurate and objective evaluation of the capabilities of large language models is achieved, which is applicable to the full life cycle management of large language models.

CN121581225BActive Publication Date: 2026-06-30CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD
Filing Date
2025-11-28
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing large language model evaluation methods cannot accurately identify subtle changes within the ability range, ignore the intrinsic connections and synergistic effects between various ability dimensions, and are difficult to accurately identify ability collapse phenomena, resulting in evaluations that are not objective and reliable enough.

Method used

We employ dynamic difficulty adjustment and multi-dimensional evaluation methods. By acquiring evaluation datasets for each capability dimension of a large language model, we calculate critical difficulty, balance, and stability to construct a comprehensive capability evaluation result, including the calculation of evaluation text difficulty for dimensions such as mathematical reasoning, code generation, and logical reasoning.

Benefits of technology

It achieves accuracy and objectivity in the assessment of large language model capabilities, can identify weaknesses and imbalances in capabilities, improves assessment efficiency, is applicable to the full lifecycle management of large language models, and provides capability comparison standards and optimization guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121581225B_ABST
    Figure CN121581225B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of artificial intelligence technology and discloses a method and related apparatus for comprehensive capability evaluation of large language models. This method acquires multi-difficulty evaluation datasets across various capability dimensions, dynamically evaluates and extracts three core indicators: critical difficulty, capability balance, and stability, and then generates a comprehensive capability evaluation result. It abandons the traditional single accuracy indicator, using critical difficulty to pinpoint weaknesses, balance to prevent "uneven development," and stability to assess performance fluctuations under the same difficulty level, achieving three-dimensional accurate evaluation. By combining dynamic difficulty with capability boundary mapping, it significantly improves evaluation efficiency and is applicable to the R&D optimization and selection comparison of large language models. It can provide capability comparison standards during the selection stage, guide data allocation and architecture design during the R&D stage, accurately identify shortcomings during the evaluation stage, and warn of capability decay during the application stage, thus providing core technical support for the full lifecycle management of large language models in various industries.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology and relates to a method and related apparatus for evaluating the comprehensive capabilities of a large language model. Background Technology

[0002] Large language models refer to language models with massive parameter scales trained on massive datasets using deep learning techniques. Large language models have driven the rapid development of artificial intelligence technology, greatly expanding the boundaries of capabilities and finding wide application across various business domains. How to measure the comprehensive capabilities of large language models has become a crucial issue affecting their effective deployment. Currently, existing large language model evaluation methods mainly adopt two technical approaches: benchmark dataset-based evaluation and adversarial testing.

[0003] Benchmark dataset evaluation is the most common method. It assesses the performance of large language models in areas such as language understanding, logical reasoning, question answering, and code generation by constructing standardized datasets covering multiple domains and tasks. These methods typically rely on pre-defined answers or human annotations and measure the performance of large language models by calculating metrics such as accuracy, F1 score, and Rouge score. Adversarial testing, on the other hand, aims to discover the vulnerabilities or unexpected behaviors of large language models under specific inputs. For example, it can induce large language models to produce incorrect, unsafe, or biased outputs by generating offensive samples, thereby evaluating the capabilities of large language models. These two methods, each with its own focus, together constitute the main framework for current large language model capability evaluation.

[0004] It is evident that existing ability assessment methods generally calculate and present the ability indicators of each dimension of a large language model independently, neglecting the inherent connections and synergistic effects between these dimensions. Furthermore, they often employ discrete, tiered assessment methods, failing to accurately identify subtle changes within ability ranges, and particularly struggling to capture ability collapse phenomena in key score intervals. These shortcomings make it difficult for existing assessment methods to accurately evaluate the true capabilities of large language models. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and related apparatus for evaluating the comprehensive capabilities of a large language model.

[0006] To achieve the above objectives, the present invention employs the following technical solution:

[0007] In a first aspect, this invention provides a method for evaluating the comprehensive capabilities of a large language model, comprising: acquiring an evaluation dataset for each capability dimension of the large language model; wherein the evaluation dataset includes several evaluation texts of different difficulties; evaluating the large language model based on the evaluation datasets for each capability dimension to obtain the critical difficulty, capability dimension balance, and capability dimension stability of each capability dimension of the large language model; wherein the critical difficulty for each capability dimension is the difficulty of the evaluation text corresponding to the accuracy of each capability dimension of the large language model dropping to a preset accuracy threshold, the capability dimension balance is determined based on the difference in critical difficulty of each capability dimension, and the capability dimension stability is determined based on the variance of the accuracy of the large language model on evaluation texts within a preset difficulty range; and obtaining the comprehensive capability evaluation result of the large language model based on the critical difficulty, capability dimension balance, and capability dimension stability of each capability dimension.

[0008] Optionally, the difficulty of the evaluation text is obtained by the following formula:

[0009]

[0010] in, To assess the difficulty of the text, As the first weight, To assess the number of key knowledge points in a text, To evaluate the word count of a text, As the second weight, To evaluate the number of reasoning steps in a text, As the third weight, This is used to evaluate the number of irrelevant information words in a text.

[0011] Optionally, when evaluating the large language model based on the evaluation datasets for each capability dimension, a dynamic difficulty adjustment strategy is employed. This dynamic difficulty adjustment strategy includes: for each capability dimension of the large language model, repeating the adjustment steps until the construction interval of the current capability dimension includes a preset accuracy interval; where the construction interval is the interval between the maximum and minimum evaluation accuracy among all evaluation accuracies for the current capability dimension; the adjustment steps include: when the current evaluation accuracy of the current capability dimension of the large language model is greater than a first preset accuracy, increasing the current difficulty of the current capability dimension and repeating the evaluation steps; when the current evaluation accuracy of the current capability dimension of the large language model is less than a second preset accuracy, decreasing the current difficulty of the current capability dimension and repeating the evaluation steps; where the evaluation steps include: based on the current difficulty of the current capability dimension of the large language model, extracting the current evaluation text for the current capability dimension from the evaluation dataset, and obtaining the accuracy of the current evaluation text based on the large language model to obtain the current evaluation accuracy of the current capability dimension of the large language model.

[0012] Optionally, the step of evaluating the large language model based on the evaluation datasets for each ability dimension to obtain the critical difficulty, ability dimension balance, and ability dimension stability of the large language model includes: for each ability dimension of the large language model, obtaining several evaluation texts of different difficulties for the current ability dimension of the large language model and the accuracy of each evaluation text based on the large language model; and fitting a relational function based on the several evaluation texts of different difficulties for the current ability dimension of the large language model and the accuracy of each evaluation text based on the large language model.

[0013]

[0014] in, For difficulty The evaluation text is based on the expected evaluation accuracy of the large language model. The slope of the curve. For the large language model The critical difficulty of each capability dimension; based on the relation function, obtain the first capability dimension of the large language model. The critical difficulty of each capability dimension.

[0015] Optionally, the evaluation of the large language model based on the evaluation datasets for each ability dimension, to obtain the critical difficulty, ability dimension balance, and ability dimension stability of the large language model, includes obtaining the ability dimension balance of the large language model through the following formula. :

[0016]

[0017] in, This is a vector composed of the critical difficulty of each ability dimension of the large language model. for standard deviation for The average value.

[0018] Optionally, the evaluation of the large language model based on the evaluation datasets for each ability dimension, and the resulting determination of the critical difficulty, ability dimension balance, and ability dimension stability of the large language model, includes obtaining the ability dimension stability of the large language model using the following formula. :

[0019]

[0020] in, For all The average value; For the large language model The stability difficulty range of the first capability dimension is included in the large language model. The range of critical difficulty for each ability dimension; For the large language model Difficulty in each ability dimension The variance of the evaluation accuracy of several evaluation texts based on a large language model.

[0021] Optionally, obtaining the comprehensive ability assessment result of the large language model based on the critical difficulty, balance, and stability of each ability dimension includes: obtaining the comprehensive ability assessment result of the large language model through the following formula. :

[0022]

[0023] in, The number of capability dimensions for a large language model. For the first The basic weights of each capability dimension for The normalized value, For the large language model The critical difficulty of each ability dimension. To ensure the balance of capabilities across the large language model, To ensure the stability of the capability dimension of a large language model.

[0024] In a second aspect, the present invention provides a comprehensive ability evaluation system for a large language model, comprising: a data acquisition module for acquiring evaluation datasets for each ability dimension of the large language model; wherein the evaluation datasets include several evaluation texts of different difficulties; a model evaluation module for evaluating the large language model based on the evaluation datasets for each ability dimension, and obtaining the critical difficulty, ability dimension balance, and ability dimension stability of each ability dimension of the large language model; wherein the critical difficulty for each ability dimension is the difficulty of the evaluation text corresponding to the accuracy of each ability dimension of the large language model dropping to a preset accuracy threshold, the ability dimension balance is determined based on the difference in critical difficulty of each ability dimension, and the ability dimension stability is determined based on the variance of the accuracy of the large language model on evaluation texts within a preset difficulty range; and an ability evaluation module for obtaining the comprehensive ability evaluation result of the large language model based on the critical difficulty, ability dimension balance, and ability dimension stability of each ability dimension.

[0025] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described large language model comprehensive ability evaluation method.

[0026] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method for evaluating the comprehensive capabilities of large language models.

[0027] Compared with the prior art, the present invention has the following beneficial effects:

[0028] This invention presents a comprehensive ability assessment method for large language models. It evaluates the overall capabilities of large language models across various ability dimensions, using assessment texts of varying difficulty as the basis. The method obtains the critical difficulty, balance, and stability of each ability dimension, and then derives the comprehensive ability assessment result. This method employs a dynamic mapping between difficulty and ability boundaries, upgrading the traditional single-dimensional accuracy assessment to a three-dimensional evaluation of ability, balance, and stability. The critical difficulty of each ability dimension reveals weaknesses in the large language model; the balance of ability dimensions reveals gaps between different dimensions, preventing the model from being biased towards any particular area; and the stability of ability dimensions reveals fluctuations in accuracy at the same difficulty level. This three-dimensional assessment accurately identifies bottlenecks in the large language model's capabilities. Compared to traditional static assessments, this method significantly improves assessment efficiency, reduces errors, and strongly guarantees the objectivity and reliability of the large language model assessment. It is applicable to scenarios such as R&D optimization and selection comparison of large language models, and ultimately achieves accurate and objective evaluation of the capabilities of large language models. It can be widely used in the full life cycle management of large language models in various industries, such as providing comparison standards of different model capabilities in the selection stage, guiding data matching and architecture design in the R&D stage, accurately locating capability shortcomings in the evaluation stage, and warning of the degradation of large language model capabilities in the application stage, thus ensuring business stability. Attached Figure Description

[0029] Figure 1 This is a flowchart of the method for evaluating the comprehensive capabilities of a large language model according to an embodiment of the present invention.

[0030] Figure 2 This is a block diagram of the large language model comprehensive ability evaluation system according to an embodiment of the present invention. Detailed Implementation

[0031] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0033] The present invention will now be described in further detail with reference to the accompanying drawings:

[0034] See Figure 1 In one embodiment of the present invention, a method for evaluating the comprehensive capabilities of a large language model is provided, which enables accurate evaluation of the comprehensive capabilities of a large language model and supports the research, development and application deployment of large language models.

[0035] Specifically, the comprehensive ability evaluation method for large language models of this invention includes the following steps:

[0036] S1: Obtain the evaluation dataset for each capability dimension of the large language model; the evaluation dataset includes several evaluation texts of different difficulty levels.

[0037] S2: Evaluate the large language model based on the evaluation datasets for each capability dimension to obtain the critical difficulty, capability balance, and capability stability of each capability dimension of the large language model. Among them, the critical difficulty of each capability dimension is the difficulty of the evaluation text when the accuracy of each capability dimension of the large language model drops to a preset accuracy threshold. The capability balance is determined based on the difference in the critical difficulty of each capability dimension, and the capability stability is determined based on the variance of the accuracy of the large language model on the evaluation text within the preset difficulty range.

[0038] S3: Based on the critical difficulty, balance, and stability of each ability dimension of the large language model, the comprehensive ability assessment result of the large language model is obtained.

[0039] This invention presents a comprehensive ability assessment method for large language models. It evaluates the overall capabilities of large language models across various ability dimensions, using assessment texts of varying difficulty as the basis. The method obtains the critical difficulty, balance, and stability of each ability dimension, and then derives the comprehensive ability assessment result. This method employs a dynamic mapping between difficulty and ability boundaries, upgrading the traditional single-dimensional accuracy assessment to a three-dimensional evaluation of ability, balance, and stability. The critical difficulty of each ability dimension reveals weaknesses in the large language model; the balance of ability dimensions reveals gaps between different dimensions, preventing the model from being biased towards any particular area; and the stability of ability dimensions reveals fluctuations in accuracy at the same difficulty level. This three-dimensional assessment accurately identifies bottlenecks in the large language model's capabilities. Compared to traditional static assessments, this method significantly improves assessment efficiency, reduces errors, and strongly guarantees the objectivity and reliability of the large language model assessment. It is applicable to scenarios such as R&D optimization and selection comparison of large language models, and ultimately achieves accurate and objective evaluation of the capabilities of large language models. It can be widely used in the full life cycle management of large language models in various industries, such as providing comparison standards of different model capabilities in the selection stage, guiding data matching and architecture design in the R&D stage, accurately locating capability shortcomings in the evaluation stage, and warning of the degradation of large language model capabilities in the application stage, thus ensuring business stability.

[0040] In one possible implementation, the capabilities of the large language model include: mathematical reasoning, code generation, text summarization, and logical reasoning. For each capability dimension of the large language model, several evaluation texts of varying difficulty are constructed; for example, for the mathematical reasoning capability dimension, several mathematical reasoning problems of varying difficulty are constructed.

[0041] In one possible implementation, the difficulty of the evaluation text is obtained by the following formula:

[0042]

[0043] in, To assess the difficulty of the text, As the first weight, To assess the number of key knowledge points in a text, To evaluate the word count of a text, As the second weight, To evaluate the number of reasoning steps in a text, As the third weight, This is used to evaluate the number of irrelevant information words in a text.

[0044] Explanatoryly, the difficulty of the evaluation text is determined by three factors: knowledge density, reasoning depth, and interference noise. Among them, knowledge density refers to the number of key knowledge points in the evaluation text. The depth of reasoning refers to the number of reasoning steps in the evaluated text. Interference noise refers to the number of irrelevant information words in the evaluation text. .

[0045] Among them, the number of key knowledge points Number of reasoning steps and the number of irrelevant information words The corresponding weight, i.e., the first weight Second weight and third weight It can be set to average weight, meaning that the three dimensions of knowledge density, reasoning depth, and interference noise are equally important, or the weight can be dynamically calculated using the entropy weight method according to actual needs.

[0046] For example, the number of words and key knowledge points in the evaluation text. Number of reasoning steps and the number of irrelevant information words It can be calculated using a pre-set large language model, and can also further calculate the difficulty value of each evaluation text. It can also be combined with manual verification to correct the above parameter values ​​and difficulty values.

[0047] In one possible implementation, when evaluating the large language model based on the evaluation datasets for each capability dimension, a dynamic difficulty adjustment strategy is employed. This strategy includes: for each capability dimension of the large language model, repeating the adjustment steps until the construction interval of the current capability dimension includes a preset accuracy interval; where the construction interval is the interval between the maximum and minimum evaluation accuracy among all evaluation accuracies for the current capability dimension; the adjustment steps include: when the current evaluation accuracy of the current capability dimension of the large language model is greater than a first preset accuracy, increasing the current difficulty of the current capability dimension and repeating the evaluation steps; when the current evaluation accuracy of the current capability dimension of the large language model is less than a second preset accuracy, decreasing the current difficulty of the current capability dimension and repeating the evaluation steps; and the evaluation steps include: based on the current difficulty of the current capability dimension of the large language model, extracting the current evaluation text for the current capability dimension from the evaluation dataset, and obtaining the accuracy of the current evaluation text based on the large language model to obtain the current evaluation accuracy of the current capability dimension of the large language model.

[0048] The explanatory approach ensures the accuracy of the evaluation by dynamically adjusting the difficulty level, guaranteeing that the evaluation of large language models involves evaluation texts of varying difficulty, thereby accurately identifying the bottlenecks in the large language model's capabilities. Furthermore, evaluating the large language model using evaluation texts of different difficulties allows for the differentiation of the strengths and weaknesses in different capability dimensions of the large language model, enabling targeted optimization of the weaker capability dimensions.

[0049] Explanatory, the dynamic adjustment is guided by a first preset accuracy rate and a second preset accuracy rate, the first preset accuracy rate being less than the second preset accuracy rate, and it is recommended that the interval between the first preset accuracy rate and the second preset accuracy rate cover as large a range as possible.

[0050] In one possible implementation, the step of evaluating the large language model based on the evaluation datasets for each capability dimension to obtain the critical difficulty, capability dimension balance, and capability dimension stability of the large language model includes: for each capability dimension of the large language model, obtaining several evaluation texts of different difficulties for the current capability dimension of the large language model and the accuracy of each evaluation text based on the large language model; and fitting a relational function based on the several evaluation texts of different difficulties for the current capability dimension of the large language model and the accuracy of each evaluation text based on the large language model.

[0051]

[0052] in, For difficulty The evaluation text is based on the expected evaluation accuracy of the large language model. The slope of the curve. For the large language model The critical difficulty of each capability dimension.

[0053] Based on the relational function, obtain the first... The critical difficulty of each capability dimension.

[0054] Explanatory methods are used to assess the critical difficulty of each ability dimension of a large language model. This involves extracting multiple sets of evaluation texts of varying difficulty for each ability dimension, with a general recommendation of at least 100 questions. By recording the accuracy of the large language model and the difficulty of the evaluation texts, the accuracy is fitted to the variation with difficulty based on the Sigmoid function.

[0055] For example, the least squares method can be used, and the solution can be obtained through optimization algorithms (such as gradient descent). and The goal is to minimize the mean square error between the fitted curve and the measured data, thus completing the fitting process.

[0056] Explanatoryly, in this implementation, the critical difficulty of the ability dimension is defined as the difficulty of the evaluation text when the accuracy of the large language model first drops to 50%, representing the upper limit of the difficulty of this ability dimension.

[0057] In one possible implementation, the step of evaluating the large language model based on the evaluation datasets for each capability dimension to obtain the critical difficulty, capability dimension balance, and capability dimension stability of the large language model includes: obtaining the capability dimension balance of the large language model through the following formula. :

[0058]

[0059] in, This is a vector composed of the critical difficulty of each ability dimension of the large language model. for standard deviation for The average value.

[0060] Explanatory, balanced ability dimensions Based on the differences in the critical difficulty of each ability dimension, this implementation method calculates the critical difficulty vector of each ability dimension according to the standard deviation and average value, so as to reflect the ability balance of each ability dimension of the large language model.

[0061] In one possible implementation, the step of evaluating the large language model based on the evaluation datasets for each capability dimension to obtain the critical difficulty, capability dimension balance, and capability dimension stability of the large language model includes: obtaining the capability dimension stability of the large language model through the following formula. :

[0062]

[0063] in, For all The average value; For the large language model The stability difficulty range of the first capability dimension is included in the large language model. The range of critical difficulty for each ability dimension; For the large language model Difficulty in each ability dimension The variance of the evaluation accuracy of several evaluation texts based on a large language model.

[0064] Explanatory, capability dimension stability Based on the variance of the accuracy of the large language model on the evaluation texts within a preset difficulty range, this embodiment characterizes the stability of the large language model's ability dimensions by setting a stable difficulty range for each ability dimension, then obtaining the variance of the evaluation accuracy of several evaluation texts within the range based on the large language model, and taking the mean. This is used to reflect the stability of the capabilities of a large language model.

[0065] Among them, the large language model The stability difficulty range of the first ability dimension is based on the large language model. The critical difficulty of each ability dimension is generally the first one in a large language model. The design uses a narrow neighborhood of the critical difficulty range for each ability dimension. This approach effectively represents the fluctuations of the large language model, reflecting its true performance near the ability boundary to the greatest extent possible and avoiding misjudgments of stability due to difficulty selection bias. For example, if low-difficulty or high-difficulty evaluation texts are selected, the accuracy of the large language model will approach 100% or 0%, resulting in a variance close to 0, making it impossible to distinguish the differences between the large language models.

[0066] In one possible implementation, obtaining the comprehensive ability assessment result of the large language model based on the critical difficulty, balance, and stability of each ability dimension includes: obtaining the comprehensive ability assessment result of the large language model using the following formula. :

[0067]

[0068] in, The number of capability dimensions for a large language model. For the first The basic weights of each capability dimension for The normalized value, For the large language model The critical difficulty of each ability dimension. To ensure the balance of capabilities across the large language model, To ensure the stability of the capability dimension of a large language model.

[0069] Explanatory, the first The basic weights for each capability dimension are average weights, meaning that each capability dimension is considered equally important. Alternatively, the weights can be dynamically calculated using the entropy weight method based on actual capability requirements.

[0070] For example, it can be calculated using the following formula. :

[0071]

[0072] in, for The maximum value in, for The minimum value in.

[0073] Explanatory The closer the value is to 1, the stronger the comprehensive capability of the large language model. It can be used for iterative optimization of large language models and comparison of the comprehensive capabilities of multiple large language models. A higher value indicates a broader and stronger comprehensive boundary for the large language model under multiple capabilities and high difficulty levels. This can be achieved through... Directly compare the comprehensive capabilities of different large language models (such as large language model A). =2.8, Large Language Model B =1.5, then the comprehensive ability of the large language model A is stronger), which can also be achieved through analysis. The shortcomings of the large language model are identified (e.g., the critical difficulty of the logical reasoning ability dimension of large language model B is 1.2, which is much lower than other ability dimensions, so the logical reasoning ability of large language model B needs to be optimized in a targeted manner).

[0074] In one possible implementation, the comprehensive ability evaluation method of the large language model of the present invention is illustrated by taking a large language model (denoted as the target large language model) as an example.

[0075] First, the target large language model has four capability dimensions: mathematical reasoning, code generation, text summarization, and logical reasoning. The critical difficulty of each capability dimension is obtained, as shown in Table 1.

[0076] Table 1

[0077]

[0078] Calculate the basic score of the target large language model based on Table 1. :

[0079] =0.25×0.74+0.25×0+0.25×1.00+0.25×0.42=0.54

[0080] The balance of the capability dimensions of the target large language model is calculated based on Table 1. :

[0081] =1 2.67 / ((7.2 + 2.1 + 9.0 + 5.0) / 4) ≈ 0.54

[0082] Stability in the ability dimension of computing target large language models Before, determine the first step of the large language model. The stability difficulty ranges for each ability dimension are as follows: mathematical reasoning has a stability difficulty range of [7.0, 7.5], code generation has a stability difficulty range of [2.0, 2.2], text summarization has a stability difficulty range of [8.8, 9.2], and logical reasoning has a stability difficulty range of [4.8, 5.2]. Specific evaluation information is shown in Table 2.

[0083] Table 2

[0084]

[0085] The stability of the capability dimension of the large language model was calculated based on Table 2. :

[0086] =1 - (0.0004 + 0.01 + 0.00003 + 0.00029) / 4 ≈ 0.997

[0087] but: =0.54*0.54*0.997≈0.291.

[0088] The comprehensive capability assessment result of the target large language model is 0.291, which can be used to quantitatively assess the comprehensive capability of the target large language model for iterative optimization and capability comparison.

[0089] In one possible implementation, the method for evaluating the comprehensive capabilities of a large language model in the power sector is illustrated using this invention as an example.

[0090] First, in terms of constructing the evaluation text, questions of varying difficulty were designed for each power-related capability dimension. In the mathematical reasoning dimension, low-difficulty questions might involve calculating simple electricity consumption, while high-difficulty questions involve predicting the transient stability of the power grid based on fluctuating load curves. In the code generation dimension, the difficulty ranged from writing electricity bill calculation functions (low difficulty) to generating scripts capable of parsing complex fault recording data (high difficulty). In the text summarization dimension, the difficulty ranged from summarizing a single power outage notice (low difficulty) to extracting a comprehensive accident analysis that includes technical parameters, event details, and reports from multiple parties (high difficulty). In the logical reasoning dimension, the difficulty ranged from judging whether a single operation complies with safety regulations (low difficulty) to analyzing the logical sequence and causal relationship of multiple protection actions in a cascading fault (high difficulty).

[0091] This evaluation method directly addresses two major pain points in the deployment of power big data language models: first, accurately identifying deficiencies. Poor stability in the logical reasoning ability dimension indicates unreliable output when processing causal chains of power grid faults, necessitating targeted enhancements in related training; second, guiding optimization directions. Poor balance in the power big data language model requires supplementing data in weaker capability dimensions, while low overall critical difficulty necessitates a comprehensive improvement in knowledge depth and reasoning ability. Ultimately, the evaluation results provide a scientific basis for the safe and efficient deployment of power big data language models in high-value scenarios such as dispatch instruction generation and intelligent fault diagnosis.

[0092] The following are embodiments of the apparatus of the present invention, which can be used to execute embodiments of the method of the present invention. For details not disclosed in the apparatus embodiments, please refer to the embodiments of the method of the present invention.

[0093] See Figure 2 In another embodiment of the present invention, a comprehensive evaluation system for large language models is provided, which can be used to implement the above-mentioned comprehensive evaluation method for large language models. Specifically, the comprehensive evaluation system for large language models includes a data acquisition module, a model evaluation module, and an ability evaluation module.

[0094] The system comprises the following modules: a data acquisition module for acquiring evaluation datasets for each capability dimension of the large language model, including evaluation texts of varying difficulty; a model evaluation module for evaluating the large language model based on these datasets, yielding the critical difficulty, capability balance, and capability stability for each capability dimension; the critical difficulty for each capability dimension is defined as the difficulty of the evaluation text at which the accuracy of each capability dimension of the large language model drops to a preset accuracy threshold; capability balance is determined based on the differences in critical difficulty for each capability dimension; and capability stability is determined based on the variance of the large language model's accuracy on evaluation texts within a preset difficulty range. Finally, a capability assessment module is used to obtain a comprehensive capability assessment result for the large language model based on its critical difficulty, capability balance, and capability stability.

[0095] In one possible implementation, the difficulty of the evaluation text is obtained by the following formula:

[0096]

[0097] in, To assess the difficulty of the text, As the first weight, To assess the number of key knowledge points in a text, To evaluate the word count of a text, As the second weight, To evaluate the number of reasoning steps in a text, As the third weight, This is used to evaluate the number of irrelevant information words in a text.

[0098] In one possible implementation, when evaluating the large language model based on the evaluation datasets for each capability dimension, a dynamic difficulty adjustment strategy is employed. This strategy includes: for each capability dimension of the large language model, repeating the adjustment steps until the construction interval of the current capability dimension includes a preset accuracy interval; where the construction interval is the interval between the maximum and minimum evaluation accuracy among all evaluation accuracies for the current capability dimension; the adjustment steps include: when the current evaluation accuracy of the current capability dimension of the large language model is greater than a first preset accuracy, increasing the current difficulty of the current capability dimension and repeating the evaluation steps; when the current evaluation accuracy of the current capability dimension of the large language model is less than a second preset accuracy, decreasing the current difficulty of the current capability dimension and repeating the evaluation steps; and the evaluation steps include: based on the current difficulty of the current capability dimension of the large language model, extracting the current evaluation text for the current capability dimension from the evaluation dataset, and obtaining the accuracy of the current evaluation text based on the large language model to obtain the current evaluation accuracy of the current capability dimension of the large language model.

[0099] In one possible implementation, the step of evaluating the large language model based on the evaluation datasets for each capability dimension to obtain the critical difficulty, capability dimension balance, and capability dimension stability of the large language model includes: for each capability dimension of the large language model, obtaining several evaluation texts of different difficulties for the current capability dimension of the large language model and the accuracy of each evaluation text based on the large language model; and fitting a relational function based on the several evaluation texts of different difficulties for the current capability dimension of the large language model and the accuracy of each evaluation text based on the large language model.

[0100]

[0101] in, For difficulty The evaluation text is based on the expected evaluation accuracy of the large language model. The slope of the curve. For the large language model The critical difficulty of each capability dimension; based on the relation function, obtain the first capability dimension of the large language model. The critical difficulty of each capability dimension.

[0102] In one possible implementation, the step of evaluating the large language model based on the evaluation datasets for each capability dimension to obtain the critical difficulty, capability dimension balance, and capability dimension stability of the large language model includes: obtaining the capability dimension balance of the large language model through the following formula. :

[0103]

[0104] in, This is a vector composed of the critical difficulty of each ability dimension of the large language model. for standard deviation for The average value.

[0105] In one possible implementation, the step of evaluating the large language model based on the evaluation datasets for each capability dimension to obtain the critical difficulty, capability dimension balance, and capability dimension stability of the large language model includes: obtaining the capability dimension stability of the large language model through the following formula. :

[0106]

[0107] in, For all The average value; For the large language model The stability difficulty range of the first capability dimension is included in the large language model. The range of critical difficulty for each ability dimension; For the large language model Difficulty in each ability dimension The variance of the evaluation accuracy of several evaluation texts based on a large language model.

[0108] In one possible implementation, obtaining the comprehensive ability assessment result of the large language model based on the critical difficulty, balance, and stability of each ability dimension includes: obtaining the comprehensive ability assessment result of the large language model using the following formula. :

[0109]

[0110] in, The number of capability dimensions for a large language model. For the first The basic weights of each capability dimension for The normalized value, For the large language model The critical difficulty of each ability dimension. To ensure the balance of capabilities across the large language model, To ensure the stability of the capability dimension of a large language model.

[0111] All relevant content of each step involved in the aforementioned embodiments of the comprehensive ability assessment method for large language models can be referenced from the functional description of the corresponding functional module of the comprehensive ability assessment system for large language models in the embodiments of the present invention, and will not be repeated here.

[0112] The module division in this embodiment of the invention is illustrative and represents only one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in the various embodiments of the invention can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0113] In another embodiment of the present invention, a computer device is provided, comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions from the computer storage medium to achieve a corresponding method flow or function. The processor described in this embodiment of the present invention can be used for the operation of a large language model synthesis capability evaluation method.

[0114] In another embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device used to store programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and extended storage media supported by the computer device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the large language model comprehensive capability evaluation method in the above embodiments.

[0115] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0116] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0117] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0118] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

Claims

1. A method for evaluating the comprehensive ability of a large language model, characterized in that, include: Obtain the evaluation dataset for each capability dimension of the large language model; the evaluation dataset includes several evaluation texts of different difficulty levels; The large language model was evaluated based on the evaluation datasets for each capability dimension, and the critical difficulty, capability balance, and capability stability of each capability dimension of the large language model were obtained. Among them, the critical difficulty of each capability dimension is the difficulty of the evaluation text when the accuracy of each capability dimension of the large language model drops to a preset accuracy threshold. The capability balance is determined based on the difference in the critical difficulty of each capability dimension, and the capability stability is determined based on the variance of the accuracy of the large language model on the evaluation text in the preset difficulty range. Based on the critical difficulty, balance, and stability of each ability dimension of the large language model, the comprehensive ability assessment results of the large language model are obtained. The difficulty of the evaluation text is obtained by the following formula: in, To assess the difficulty of the text, As the first weight, To assess the number of key knowledge points in a text, To evaluate the word count of a text, As the second weight, To evaluate the number of reasoning steps in a text, As the third weight, To measure the number of irrelevant information words in the text; The evaluation of the large language model based on the assessment datasets for each capability dimension yields the following results: critical difficulty, capability dimension balance, and capability dimension stability of the large language model. For each capability dimension of the large language model, obtain the evaluation texts of different difficulties for the current capability dimension of the large language model and the evaluation accuracy of each evaluation text based on the large language model; Based on several evaluation texts of varying difficulty levels across the current capability dimension of the large language model and the evaluation accuracy of each text using the large language model, a fitting relationship function is established: in, For difficulty The evaluation text is based on the expected evaluation accuracy of the large language model. The slope of the curve. For the large language model The critical difficulty of each capability dimension; Based on the relational function, obtain the first... The critical difficulty of each capability dimension.

2. The method for evaluating the comprehensive ability of a large language model according to claim 1, characterized in that, When evaluating a large language model based on assessment datasets for each capability dimension, a dynamic difficulty adjustment strategy is employed. This dynamic difficulty adjustment strategy includes: For each capability dimension of the large language model, the adjustment steps are repeated to ensure that the construction interval of the current capability dimension of the large language model includes a preset accuracy interval. The construction interval is the interval between the maximum and minimum evaluation accuracy among all evaluation accuracies of the current capability dimension. The adjustment steps include: when the current evaluation accuracy of the current capability dimension of the large language model is greater than a first preset accuracy, increasing the current difficulty of the current capability dimension of the large language model and repeating the evaluation steps; when the current evaluation accuracy of the current capability dimension of the large language model is less than a second preset accuracy, decreasing the current difficulty of the current capability dimension of the large language model and repeating the evaluation steps. The evaluation steps include: based on the current difficulty of the current capability dimension of the large language model, extracting the current evaluation text of the current capability dimension from the evaluation dataset, and obtaining the evaluation accuracy of the current evaluation text based on the large language model to obtain the current evaluation accuracy of the current capability dimension of the large language model.

3. The method for evaluating the comprehensive ability of a large language model according to claim 1, characterized in that, The evaluation of the large language model based on the assessment datasets for each capability dimension yields the following results: critical difficulty, capability dimension balance, and capability dimension stability of the large language model. The capability dimension balance of the large language model is obtained through the following formula. : in, This is a vector composed of the critical difficulty of each ability dimension of the large language model. for standard deviation for The average value.

4. The method for evaluating the comprehensive ability of a large language model according to claim 1, characterized in that, The evaluation of the large language model based on the assessment datasets for each capability dimension yields the following results: critical difficulty, capability dimension balance, and capability dimension stability of the large language model. The stability of the capability dimension of the large language model is obtained through the following formula. : in, For all The average value; For the large language model The stability difficulty range of the first capability dimension is included in the large language model. The range of critical difficulty for each ability dimension; For the large language model Difficulty in each ability dimension The variance of the evaluation accuracy of several evaluation texts based on a large language model.

5. The method for evaluating the comprehensive ability of a large language model according to claim 1, characterized in that, The comprehensive ability assessment results of the large language model, based on the critical difficulty, balance, and stability of each ability dimension, include: The comprehensive ability assessment result of the large language model is obtained through the following formula. : in, The number of capability dimensions for a large language model. For the first The basic weights of each capability dimension for The normalized value, For the large language model The critical difficulty of each ability dimension. To ensure the balance of capabilities across the large language model, To ensure the stability of the capability dimension of a large language model.

6. A comprehensive language model assessment system, characterized in that, include: The data acquisition module is used to acquire evaluation datasets for each capability dimension of the large language model; the evaluation datasets include several evaluation texts of different difficulty levels. The model evaluation module is used to evaluate the large language model based on the evaluation datasets of each capability dimension, and to obtain the critical difficulty, capability dimension balance, and capability dimension stability of the large language model for each capability dimension. Among them, the critical difficulty of each capability dimension is the difficulty of the evaluation text when the accuracy of each capability dimension of the large language model drops to a preset accuracy threshold. The capability dimension balance is determined based on the difference in the critical difficulty of each capability dimension, and the capability dimension stability is determined based on the variance of the accuracy of the large language model on the evaluation text in the preset difficulty range. The ability assessment module is used to obtain the comprehensive ability assessment results of the large language model based on the critical difficulty, balance and stability of each ability dimension. The difficulty of the evaluation text is obtained by the following formula: in, To assess the difficulty of the text, As the first weight, To assess the number of key knowledge points in a text, To evaluate the word count of a text, As the second weight, To evaluate the number of reasoning steps in a text, As the third weight, To measure the number of irrelevant information words in the text; The evaluation of the large language model based on the assessment datasets for each capability dimension yields the following results: critical difficulty, capability dimension balance, and capability dimension stability of the large language model. For each capability dimension of the large language model, obtain the evaluation texts of different difficulties for the current capability dimension of the large language model and the evaluation accuracy of each evaluation text based on the large language model; Based on several evaluation texts of varying difficulty levels across the current capability dimension of the large language model and the evaluation accuracy of each text using the large language model, a fitting relationship function is established: in, For difficulty The evaluation text is based on the expected evaluation accuracy of the large language model. The slope of the curve. For the large language model The critical difficulty of each capability dimension; Based on the relational function, obtain the first... The critical difficulty of each capability dimension.

7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the large language model comprehensive capability evaluation method as described in any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the large language model comprehensive capability evaluation method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Power market large language model evaluation method and device based on dynamic scene perception

    CN120910511A

  • Evaluation for large language model

    WO2025102964A1