Large language model capability assessment method and system based on dynamic data assessment
Through dynamic data evaluation and complexity control mechanisms, combined with a multi-dimensional evaluation framework, data pollution and evaluation accuracy problems in the existing technology are solved, and the reliability and fairness of large-language model evaluation is improved.
Patent Information
- Application Number
- CN202510481116.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing large language model evaluation methods face technical challenges such as data pollution, assessment accuracy and complexity control, which have affected the reliability and fairness of the evaluation results.
By introducing dynamic data evaluation methods, dynamically generate evaluation samples, combining complexity control mechanisms and multi-dimensional evaluation frameworks, we ensure the quality and consistency of the evaluation data and improve the reliability and fairness of the evaluation.
It effectively reduces the impact of data pollution on assessment results, improves the accuracy and fairness of assessment, ensures the quality and consistency of assessment data, and is suitable for cross-domain and diverse assessment tasks.
Smart Images

Figure CN119988914A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and in particular to a large language model capability evaluation method and system based on dynamic data evaluation. Background Art
[0002] With the widespread application of large language models (LLMs) in the field of natural language processing, their performance in various tasks has achieved remarkable results. However, in actual use, LLMs face a series of evaluation challenges, especially the problem of data contamination that may be caused during the evaluation process. Current evaluation methods for large LLMs mainly rely on standard datasets, which are usually composed of a large amount of training data. The core task of the evaluation is to verify the performance of the model on a specific task. The evaluation of LLMs generally focuses on factors such as task completion, answer accuracy, and processing time. Common evaluation methods include evaluation based on manual annotation, automatic scoring algorithms, and comparative analysis based on model output. With the continuous expansion of large-scale Internet corpora used in the training process of LLMs, problems with the reliability and fairness of its evaluation have gradually been exposed. The root cause of these problems is mainly the phenomenon of data contamination. In particular, when there is overlap between the evaluation dataset and the training data, the evaluation results may not accurately reflect the generalization ability of the model, and even affect the fair comparison between different models.
[0003] Existing technologies mainly focus on two major directions: data pollution detection and dynamic data evaluation. Among them, data pollution detection methods solve the pollution problem by identifying the overlap between model output and training data. However, since some closed-source language models, such as the GPT series, use a variety of complex filtering mechanisms during the generation process, these methods often find it difficult to capture implicit data pollution. For example, although the model output may appear to be unpolluted, in some cases, the model will still produce output that is highly similar to its training data, resulting in inaccurate evaluation results.
[0004] On the other hand, dynamic data evaluation methods attempt to circumvent the limitations of traditional evaluation methods by generating new, uncontaminated data sets. Representative prior art includes DYVAL, KIEval, LatestEval, SciEval, and the methods of Ying J et al. These methods each explore the potential of dynamic data evaluation from different perspectives and provide valuable tools for LLMs competence evaluation. However, in practical applications, there are still some technical challenges and shortcomings, which are mainly reflected in the narrow scope of application, unstable data generation quality, insufficient control of question complexity, lack of a multi-dimensional evaluation system, data pollution, etc. Therefore, the present invention proposes a large language model capability evaluation method and system based on dynamic data evaluation. Summary of the invention
[0005] The purpose of the present invention is to provide a large language model capability assessment method and system based on dynamic data assessment, which ensures the quality and consistency of the assessment data and improves the reliability and fairness of LLMs capability assessment through dynamic generation, complexity control and multi-dimensional evaluation of evaluation samples.
[0006] To achieve the above object, on the one hand, the present invention provides a large language model capability evaluation method based on dynamic data evaluation, comprising:
[0007] Obtain the topic input by the user, and extract the core knowledge points and main content from the topic;
[0008] Based on the core knowledge points and main content, a pre-trained large language model is used to perform online retrieval to generate detailed knowledge descriptions related to the topic;
[0009] Generate assessment questions based on the core knowledge points, main content and knowledge elaboration;
[0010] Adjust and optimize the difficulty of the assessment questions to obtain final assessment questions;
[0011] Perform multi-dimensional capability assessment and quality inspection on the final assessment questions, obtain assessment results, and complete the capability assessment of the large language model.
[0012] Optionally, extracting core knowledge points and main content from the topic includes:
[0013] Analyze the question by natural language processing technology to obtain analysis information;
[0014] The core knowledge points and main content are extracted based on the parsed information using few-shot learning technology.
[0015] Optionally, based on the core knowledge points and main content, a pre-trained large language model is used to perform online retrieval to generate a detailed description of knowledge related to the topic, including:
[0016] Conduct online search based on the core knowledge points and main content to obtain relevant background information;
[0017] Input the core knowledge points and main content into the pre-trained large language model and output the knowledge elaboration;
[0018] The relevant background information and knowledge elaboration are integrated to obtain the knowledge elaboration.
[0019] Optionally, based on the core knowledge points, main content and knowledge elaboration, generating assessment questions includes:
[0020] Integrate the core knowledge points, main content and knowledge details to obtain integrated information;
[0021] Based on the integrated information, a preliminary question framework is generated according to question type and assessment objectives;
[0022] The complexity of the preliminary question framework is adjusted and the repeatability is tested, and Bloom's cognitive hierarchy is used to obtain a multi-level question framework;
[0023] Perform quality inspection on the multi-level question framework to obtain the evaluation questions.
[0024] Optionally, the difficulty of the evaluation topic is adjusted and optimized to obtain the final evaluation topic, including:
[0025] Performing a complexity assessment on the assessment topic and obtaining an assessment result;
[0026] Performing complexity control based on the evaluation result to obtain the evaluation topic after the control;
[0027] Multiple large language models are used to verify and provide feedback on the adjusted evaluation questions, and the adjusted evaluation questions are optimized based on the feedback results to obtain the final evaluation questions.
[0028] Optionally, adjusting and optimizing the difficulty of the evaluation topic, and obtaining the final evaluation topic further includes:
[0029] Performing a difficulty evaluation on the evaluation questions, and sorting the evaluation questions according to the evaluation results to obtain a difficulty distribution;
[0030] The difficulty span is determined based on the difficulty distribution. When the difficulty span exceeds a preset difficulty span, the difficulty of the assessment questions is smoothed to obtain assessment questions with a balanced difficulty distribution.
[0031] Optionally, performing a multi-dimensional ability assessment on the final assessment topic includes:
[0032] Using Bloom's cognitive hierarchy, generating classification questions of different cognitive levels based on the final assessment questions;
[0033] Evaluate the performance of the large language model at different cognitive levels including classification questions from multiple dimensions to obtain multi-dimensional evaluation results.
[0034] Optionally, performing quality inspection on the final evaluation topic includes:
[0035] Use several large language models to test the quality, accuracy, logic rationality, consistency and diversity of each final evaluation question and obtain the test results;
[0036] The detection result is output, and the final evaluation question is automatically optimized and regenerated based on the detection result.
[0037] On the other hand, the present invention provides a large language model capability evaluation system based on dynamic data evaluation, comprising:
[0038] The knowledge point and main content extraction module is used to obtain the topic input by the user and extract the core knowledge points and main content from the topic;
[0039] An online search and knowledge elaboration module is used to perform online search based on the core knowledge points and main content using a pre-trained large language model to generate knowledge elaborations related to the topic;
[0040] A question design module, for generating assessment questions based on the core knowledge points, main content and knowledge elaboration;
[0041] A complexity control module is used to adjust and optimize the difficulty of the evaluation questions to obtain the final evaluation questions;
[0042] A multi-dimensional assessment module, used to conduct a multi-dimensional ability assessment on the final assessment topic and obtain an assessment result;
[0043] The quality detection module is used to perform quality detection on the final evaluation questions and obtain detection results.
[0044] The beneficial effects of the present invention are:
[0045] The present invention introduces a complexity control mechanism to ensure that the difficulty of the generated questions is consistent with the original data, thereby improving the objectivity and accuracy of the evaluation; through a multi-dimensional evaluation framework, the performance of the model at different cognitive levels is comprehensively examined, especially the evaluation of high-order abilities such as reasoning, evaluation, and creativity; through a quality detection mechanism, the stability and accuracy of the generated data are improved to ensure the reliability of the evaluation results; by introducing a multi-model verification mechanism, the pollution detection capability is optimized to reduce the impact of data pollution on the evaluation results; it can expand the scope of application of existing evaluation methods and provide cross-domain and flexible evaluation solutions; through the above-mentioned technical contents, the present invention ensures the quality and consistency of the evaluation data and improves the reliability and fairness of the LLMs ability evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0047] Figure 1 A flow chart of a large language model capability assessment method based on dynamic data assessment according to an embodiment of the present invention;
[0048] Figure 2 The present invention is a schematic diagram of the structure of a large language model capability assessment system based on dynamic data assessment according to an embodiment of the present invention. DETAILED DESCRIPTION
[0049] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0050] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0051] The prior art has the following defects:
[0052] 1. Narrow scope of application:
[0053] Most existing dynamic data evaluation methods rely on preset data models and fixed question generation mechanisms in specific fields, which limits their applicability in other fields or tasks. Although DYVAL is effective in the field of mathematics, it does not provide extensive support for other fields; KIEval is more suitable for dialogue scenarios and is difficult to adapt to other types of evaluation tasks. Existing technologies lack cross-domain flexibility and cannot provide a unified solution for a variety of evaluation tasks. The evaluation requirements between different tasks vary greatly, and existing methods rely more on a fixed question design framework and fail to fully consider flexibility and scalability issues.
[0054] 2. Unstable data generation quality:
[0055] The generated data often cannot maintain high quality, and may contain redundancy, bias, or even fail to accurately reflect the performance of LLMs in certain specific tasks. For example, when generating mathematical reasoning questions, although DYVAL can generate high-quality questions, there is still room for improvement in the complexity and accuracy control of the task. These methods fail to implement a complete quality control mechanism when generating questions, resulting in the generated data may not accurately reflect the capabilities of the model in complex scenarios, and may even produce evaluation samples that do not meet the original task requirements. This instability affects the reliability of the final evaluation results, especially in scenarios that require high accuracy.
[0056] 3. Insufficient ability to control the complexity of questions:
[0057] Most existing dynamic data evaluation methods lack an effective mechanism to control the complexity of questions, especially how to ensure that the complexity of generated questions is consistent with the difficulty level of the original data. For example, SciEval and KIEval test the complexity of the model by designing new questions, but these methods still lack a sophisticated difficulty control mechanism, resulting in the difficulty of the questions not fully meeting the needs of the target evaluation. Therefore, the evaluation results of existing methods may not objectively reflect the model's capabilities at different difficulty levels, especially in the evaluation of reasoning or creative tasks, it is difficult to accurately evaluate the model's high-level cognitive abilities.
[0058] 4. Lack of multi-dimensional evaluation system:
[0059] Most of the current dynamic data evaluation methods focus on single-dimensional testing, usually focusing on basic dimensions such as model accuracy and speed, and ignoring the comprehensive evaluation of the model on multi-level and complex cognitive tasks. Bloom's taxonomy provides a more refined division of cognitive levels, but existing evaluation methods often fail to fully utilize this framework for comprehensive evaluation, especially for high-level cognitive abilities such as analysis, evaluation, and creation. For example, the simple strategy proposed by Ying J et al. can automatically update the evaluation data set, but it does not consider how to comprehensively evaluate the model's capabilities through questions at different levels.
[0060] 5. Data pollution problem:
[0061] Data contamination remains a serious challenge in the evaluation field. Existing contamination detection methods attempt to address this problem by identifying overlaps between training and evaluation data, but these methods are limited in their effectiveness in capturing implicit contamination due to the special filtering mechanisms used when closed-source language models, such as the GPT series, are generated. Even methods specifically designed to avoid contamination, such as LatestEval, cannot completely eliminate the impact of data contamination, especially when dealing with certain implicit contaminations, where existing technologies are still unable to cope.
[0062] To solve the above technical problems, on the one hand, this embodiment provides a large language model capability evaluation method based on dynamic data evaluation, such as Figure 1 As shown, including:
[0063] Obtain the topic input by the user, and extract the core knowledge points and main content from the topic;
[0064] Based on the core knowledge points and main content, a pre-trained large language model is used to perform online retrieval to generate detailed knowledge descriptions related to the topic;
[0065] Generate assessment questions based on the core knowledge points, main content and knowledge elaboration;
[0066] Adjust and optimize the difficulty of the assessment questions to obtain final assessment questions;
[0067] Perform multi-dimensional capability assessment and quality inspection on the final assessment questions, obtain assessment results, and complete the capability assessment of the large language model.
[0068] Specifically, this embodiment introduces a complexity control mechanism to ensure that the difficulty of the generated questions is consistent with the original data, thereby improving the objectivity and accuracy of the evaluation; through a multi-dimensional evaluation framework, the performance of the model at different cognitive levels is comprehensively examined, especially in the evaluation of high-level abilities such as reasoning, evaluation, and creation; through a quality detection mechanism, the stability and accuracy of the generated data are improved to ensure the reliability of the evaluation results; by introducing a multi-model verification mechanism, the pollution detection capability is optimized to reduce the impact of data pollution on the evaluation results; it can expand the scope of application of existing evaluation methods and provide cross-domain and flexible evaluation solutions. Through the above technical content, the quality and consistency of the evaluation data can be ensured, and the reliability and fairness of the LLMs ability evaluation can be improved.
[0069] Furthermore, extracting core knowledge points and main contents from the topic includes:
[0070] Analyze the question by natural language processing technology to obtain analysis information;
[0071] The core knowledge points and main content are extracted based on the parsed information using few-shot learning technology.
[0072] The specific steps include:
[0073] Question parsing: Receive questions or data sets input by users and parse the questions using natural language processing (NLP) technology.
[0074] Core knowledge point extraction: Through pre-trained models or rule-based algorithms, key concepts or categories related to knowledge are automatically extracted from the questions to ensure that these knowledge points can represent the core content of the questions without containing details or redundant information irrelevant to the core of the questions.
[0075] Main idea extraction: In addition to extracting knowledge points, the main idea of the question is also extracted to summarize the overall meaning of the question and avoid lengthy statements. The main idea extracted should cover the question background, key concepts, and the required core reasoning or knowledge application.
[0076] Few-shot learning technology is used in the extraction process to fine-tune with a small number of examples to further improve the accuracy of knowledge point and main idea extraction. Few-shot learning enables this method to quickly adapt and extract the most representative knowledge points and main ideas when facing various types of questions, ensuring that the generated evaluation questions are consistent with the original data in terms of core ideas and knowledge points, providing a reliable foundation for the subsequent evaluation process.
[0077] Furthermore, based on the core knowledge points and main content, a pre-trained large language model is used to perform online retrieval to generate detailed knowledge descriptions related to the topic, including:
[0078] Conduct online search based on the core knowledge points and main content to obtain relevant background information;
[0079] Input the core knowledge points and main content into the pre-trained large language model and output the knowledge elaboration;
[0080] The relevant background information and knowledge elaboration are integrated to obtain the knowledge elaboration.
[0081] This is achieved through the following steps:
[0082] 1. Online search:
[0083] First, based on the knowledge points output from the knowledge point and theme extraction module, search external databases or the Internet to obtain the latest information, research results, definitions or cases related to these knowledge points. Through the background information obtained by retrieval, the content of the questions is further enriched, so that the evaluation questions are not limited to the surface meaning of the questions, but can also cover a wider knowledge system, improving the depth and accuracy of the model evaluation.
[0084] 2. Knowledge Detail Generation:
[0085] After obtaining relevant background information, the pre-trained large language model is used to elaborate on the retrieved knowledge points. Specifically, it includes:
[0086] Detailed explanation: Based on the core concepts of the knowledge points, generate detailed explanations of related fields, explaining the definitions, application scenarios and related research progress of these concepts.
[0087] Examples and cases: Combined with actual application scenarios or academic research, provide specific examples or cases to help further understand the practical application of knowledge points.
[0088] Domain expansion: Based on the question type and difficulty, the method will also expand the domain knowledge related to the core knowledge points to ensure that the generated questions can cover multiple dimensions of knowledge points.
[0089] 3. Integration and generation of knowledge details:
[0090] Integrate the retrieved relevant background information with the knowledge explanation generated by the model to form a concise, clear and structured text. The text should focus on the core meaning of the knowledge point and exclude irrelevant or redundant information. In this way, each question is provided with richer and deeper content support, providing a solid knowledge foundation for the subsequent generation of evaluation questions.
[0091] 4. Detailed description of output knowledge:
[0092] The generated knowledge description will be output together with the knowledge points and main content as the core basis for the subsequent question generation. Ultimately, these detailed contents will help the diversified generation and complexity control of questions in the subsequent steps to ensure the quality and availability of evaluation data.
[0093] Furthermore, based on the core knowledge points, main content and knowledge elaboration, the generated assessment questions include:
[0094] Integrate the core knowledge points, main content and knowledge details to obtain integrated information;
[0095] Based on the integrated information, a preliminary question framework is generated according to question type and assessment objectives;
[0096] The complexity of the preliminary question framework is adjusted and the repeatability is tested, and Bloom's cognitive hierarchy is used to obtain a multi-level question framework;
[0097] Perform quality inspection on the multi-level question framework to obtain the evaluation questions.
[0098] This is achieved through the following steps:
[0099] 1. Input content integration:
[0100] First, we integrate the core knowledge points and main content, as well as the knowledge background information, as the basic input for question generation. Based on this input information, we can ensure that the generated assessment questions are designed closely around the core knowledge.
[0101] 2. Generate the title framework:
[0102] After integrating the input content, generate a preliminary question framework based on the question type, such as multiple-choice questions, fill-in-the-blank questions, short-answer questions, etc., and the assessment objectives, such as testing memory, comprehension, application, analysis, etc. Specifically, it includes:
[0103] Determine the question type: Select the appropriate question type based on the requirements of the task, such as multiple-choice questions, true-or-false questions, or short-answer questions.
[0104] Question design: Design specific questions based on the integrated input content. The content of the questions should be closely centered around the knowledge points to ensure that the core knowledge of the test can be accurately covered.
[0105] Option generation: This embodiment is only for multiple-choice questions. When generating multiple-choice questions, multiple options are generated according to the content of the question, including a correct answer and several distracting wrong options. The generation of distracting options requires not only that they have a certain logical similarity with the correct answer, but also that they conform to the context of the original question.
[0106] 3. Complexity control and alignment:
[0107] In order to ensure that the generated questions are consistent with the original questions in terms of difficulty or meet the predetermined requirements, the complexity of the questions is controlled and aligned. Specifically, it includes:
[0108] Difficulty assessment of knowledge points: Generate matching questions based on the difficulty of the knowledge points to ensure that the difficulty level of the questions is reasonable.
[0109] Problem complexity adjustment: Automatically adjust the complexity of the questions according to the task requirements. For example, in situations where a higher level of cognition needs to be tested, such as analysis, evaluation, and creativity, the design of the questions will require more complex reasoning or application of real-world scenarios.
[0110] Generate multi-level questions: Generate corresponding multi-dimensional questions based on the six cognitive levels of Bloom's taxonomy, including questions in six dimensions: memory, understanding, application, analysis, evaluation and creation, and ensure that the questions in each dimension have appropriate difficulty and question type.
[0111] 4. Innovative and diverse topic design:
[0112] In addition to ensuring the basic requirements of the topic, we also use innovative design to ensure that the generated topics have a certain degree of diversity. For example:
[0113] Examining knowledge points from different angles: By changing the question types and applying different backgrounds or situations, students' multi-dimensional understanding of knowledge points can be examined.
[0114] Innovation in question format: You can design some innovative questions, such as case analysis questions, reasoning questions, comparison questions, etc., to increase the complexity and depth of the questions.
[0115] Avoid repetitive topics: Avoid generating topics that are too similar to the original topics, and ensure that the newly generated topics are innovative and unique in form and content.
[0116] 5.Quality inspection and optimization:
[0117] After generating the questions, we will conduct a quality check on the questions to ensure that each question meets the following standards:
[0118] Clear logic: The question wording of each question should be clear and unambiguous to ensure that the subjects can understand and answer accurately.
[0119] The options should be reasonable and discriminatory: In multiple-choice questions, the options should be distracting to avoid correct answers that are too obvious and ensure the depth of the examination.
[0120] Only one correct answer: Make sure there is only one correct answer among the options of multiple-choice questions, and that the other options can reasonably interfere with students' judgment.
[0121] Core knowledge point coverage: Each question should accurately reflect the core knowledge point and be able to effectively test students' mastery of the knowledge point.
[0122] Furthermore, the difficulty of the evaluation questions is adjusted and optimized to obtain the final evaluation questions, including:
[0123] Performing a complexity assessment on the assessment topic and obtaining an assessment result;
[0124] Performing complexity control based on the evaluation result to obtain the evaluation topic after the control;
[0125] Multiple large language models are used to verify and provide feedback on the adjusted evaluation questions, and the adjusted evaluation questions are optimized based on the feedback results to obtain the final evaluation questions.
[0126] The method also includes: performing difficulty evaluation on the evaluation questions, and sorting the evaluation questions according to the evaluation results to obtain a difficulty distribution; judging a difficulty span based on the difficulty distribution, and when the difficulty span exceeds a preset difficulty span, smoothing the difficulty of the evaluation questions to obtain evaluation questions with a balanced difficulty distribution.
[0127] The specific steps are as follows:
[0128] 1. Assessment of topic complexity:
[0129] After generating the assessment topic, the complexity control module first performs a preliminary assessment of the complexity of the topic. The assessment is performed in the following ways:
[0130] Difficulty score: Each question is assigned a difficulty score based on the type, structure and difficulty of the knowledge points. The scoring criteria include factors such as the depth of the question, the complexity of reasoning, and the abstractness of the concepts involved in the question.
[0131] Benchmark data comparison: By comparing with the questions in the original dataset, we evaluate whether the difficulty of the generated questions is consistent with the level of the original questions. For example, if the difficulty of the original questions is high, the generated questions also need to have a corresponding complexity; if the difficulty of the original dataset is low, the method will adjust the difficulty of the generated questions to meet the predetermined difficulty requirements.
[0132] Multi-dimensional difficulty assessment: In addition to the difficulty of the knowledge points, the complexity control module also considers the cognitive level of the questions. For example, the high-level cognitive dimensions of Bloom's taxonomy, such as analysis, evaluation, and creation, usually correspond to higher complexity, while memory and comprehension questions are relatively simple. The method assigns different difficulty scores to questions based on the cognitive level and adjusts them based on the overall assessment results.
[0133] 2. Complexity control mechanism:
[0134] After completing the preliminary complexity assessment, adopt corresponding complexity control strategies based on the assessment results to ensure that the difficulty of the generated questions meets expectations:
[0135] Complexity alignment: Automatically adjust the complexity of questions based on the evaluation results. For example, if a question is too difficult after it is generated, the method will reduce the difficulty of the question by simplifying the question's statement, reducing the number of reasoning steps in the question, or reducing the number of concepts involved in the question; if the complexity of the question is too low, the method will increase the difficulty of the question and require students to conduct more in-depth analysis or reasoning.
[0136] Difficulty offset calculation: In order to ensure that the difficulty of the questions is consistent with the original data set, the accuracy difference between the generated data set and the original data set is calculated, that is, the offset. By calculating the offset, the difficulty gap between the generated questions and the original questions can be quantified, and the content of the generated questions can be adjusted according to the gap. For example, if the average accuracy of the generated data set is too high, the difficulty of the questions can be increased by modifying some questions or adding more challenging questions; conversely, if the accuracy is low, the method will simplify the questions appropriately.
[0137] Automatic adjustment strategy: For questions with high or low complexity, the corresponding strategy will be automatically selected for adjustment. Specific adjustments include:
[0138] Question simplification: For questions that are too difficult, simplify the stem, reduce multiple reasoning in the questions, reduce the distractors in the options, or reduce the difficulty by rephrasing the questions.
[0139] Question strengthening: For questions that are too easy, add more challenging situations or requirements, add more analysis, reasoning or interdisciplinary knowledge points, and ensure that the questions can cover more advanced cognitive levels.
[0140] 3. Multi-model verification and feedback mechanism:
[0141] We also introduced a multi-model verification mechanism, which uses multiple large language models (LLMs) to verify and provide feedback on the difficulty of the generated questions. The specific operations are as follows:
[0142] Multi-model prediction: The generated questions are input into multiple preset LLMs, and the difficulty of the questions is predicted by these models. Based on the feedback results of the model, the difficulty of the questions can be further optimized to meet the predetermined requirements.
[0143] Feedback adjustment: fine-tune the complexity of questions based on feedback from multiple models. If multiple models think a question is too simple or too complex, make corresponding adjustments based on the feedback to ensure that the difficulty of the question is within a reasonable range.
[0144] 4. Balanced Difficulty Distribution:
[0145] In order to ensure the balance of difficulty of the generated data set, it is also necessary to monitor the difficulty distribution of the entire data set. By analyzing the distribution of question difficulty, it can ensure that the difficulty level of the data set is reasonable, and questions from simple to complex can fully cover all cognitive levels. Specific operations include:
[0146] Question difficulty distribution: Based on the predetermined difficulty level requirements, the generated questions are sorted according to the difficulty distribution, and ensure that questions of different levels are covered evenly.
[0147] Difficulty span adjustment: When the difficulty span of the generated questions is large, the difficulty of the questions is smoothed to ensure that the overall difficulty distribution of the questions does not have any extreme questions that are too prominent.
[0148] Furthermore, the multi-dimensional ability assessment of the final assessment topic includes:
[0149] Using Bloom's cognitive hierarchy, generating classification questions of different cognitive levels based on the final assessment questions;
[0150] Evaluate the performance of the large language model at different cognitive levels including classification questions from multiple dimensions to obtain multi-dimensional evaluation results.
[0151] The specific steps are as follows:
[0152] 1. Cognitive level division and task definition:
[0153] In the multi-dimensional assessment module, each question is first divided into cognitive levels according to Bloom's taxonomy. Bloom's taxonomy divides the degree of knowledge mastery into six cognitive levels, from the most basic memory to the most complex creation, including:
[0154] Remembering: This measures whether students can recall and recognize learned facts, terms, definitions and other basic knowledge.
[0155] Understanding: Tests whether students can understand the meaning of information and can express or explain it in their own words.
[0156] Application: Students are required to apply what they have learned to new situations to solve practical problems.
[0157] Analyzing: Examines whether students can break down complex content into basic elements and understand the relationship between the parts.
[0158] Evaluating: Assessing whether students can judge and evaluate the validity, accuracy and value of information based on certain criteria.
[0159] Creating: examines whether students can construct new ideas, methods or models based on existing knowledge.
[0160] 2. Matching of topics and cognitive levels:
[0161] A cognitive level is automatically assigned to each question based on the knowledge points, complexity, and ability dimensions of each question. For example, simple fact recall questions will be marked as the "memory" level, while questions that require complex reasoning or comprehensive analysis will be marked as the "analysis" or "creativity" level. Through this process, it is ensured that each question can not only test basic knowledge, but also involve higher-level thinking skills, thereby comprehensively evaluating the performance of the model.
[0162] 3. Level coverage and topic distribution:
[0163] In order to ensure the comprehensiveness of the assessment, the number of questions for each level is reasonably allocated according to the six cognitive levels of Bloom's taxonomy. The specific operations include:
[0164] Level coverage: Ensure that there are questions for each level to test, so that certain cognitive levels, such as creativity or evaluation, are not neglected. Questions at each cognitive level will cover certain knowledge points and ensure that the questions at that level can fully reflect the students' relevant abilities.
[0165] Balanced number of questions: When generating assessment data, the number of questions for each level is evenly distributed according to the predetermined difficulty and level requirements. For example, for more complex levels, such as analysis and creation, the number of questions is appropriately reduced, while for basic levels, such as memory and understanding, the number of questions is increased to ensure that the difficulty distribution is reasonable and meets the assessment objectives.
[0166] 4. Multi-dimensional evaluation generation:
[0167] After completing the matching and allocation of questions and cognitive levels, questions are generated based on different cognitive levels through a multi-dimensional assessment generation mechanism. These questions cover the ability assessment of each level, ensuring that the assessment can fully reflect the student's performance in each cognitive dimension.
[0168] Memory and comprehension questions: These are generally tests of basic knowledge, requiring students to memorize and understand core concepts. Questions may include basic forms such as fill-in-the-blank questions and multiple-choice questions.
[0169] Application questions: Students are required to apply the knowledge they have learned to actual scenarios or new situations. The questions are usually in the form of case analysis, problem solving, etc.
[0170] Analytical, evaluative and creative questions: These questions are more complex and require students to conduct in-depth analysis, comparison or creative thinking. They usually involve longer texts, reasoning processes and program design.
[0171] 5. Model capability scoring and evaluation:
[0172] In the multidimensional assessment process, LLMs are scored for their performance at each cognitive level. By comparing the model's answers with the standard answers, the model's capabilities at each level can be assessed. For example:
[0173] At the "memory" level, the model is mainly evaluated on whether it can accurately recall and recognize basic facts and concepts.
[0174] At the “understanding” level, evaluate whether the model can accurately explain and restate the knowledge points.
[0175] At the “application” level, the model is evaluated to see whether it can apply knowledge to new scenarios and solve practical problems.
[0176] At the "analysis", "evaluation" and "creation" levels, the model's high-level cognitive abilities are assessed by examining its reasoning process and creative output.
[0177] 6. Output multi-dimensional evaluation results:
[0178] After completing the above evaluation, a multi-dimensional evaluation report is generated, covering the performance of the model at each cognitive level. These results will help researchers and developers understand the capabilities of the model at different levels and provide optimization directions. For example, a model with a low score at the "creativity" level may mean that the model has limitations in generating innovative answers and needs to further improve its innovation capabilities.
[0179] Furthermore, the quality inspection of the final evaluation topic includes:
[0180] Use several large language models to test the quality, accuracy, logic rationality, consistency and diversity of each final evaluation question and obtain the test results;
[0181] The detection result is output, and the final evaluation question is automatically optimized and regenerated based on the detection result.
[0182] The specific steps are as follows:
[0183] 1. Multi-language model voting mechanism:
[0184] First, multiple large language models (LLMs) are used to vote on the generated evaluation questions. These language models include multiple preset open source and closed source models, such as the GPT series, Qwen, and Doubao. The quality of each question is judged separately to evaluate the accuracy and consistency of its content.
[0185] Voting mechanism: Each generated question will be input into at least three different LLMs, and a comprehensive analysis will be performed based on the output results of these models. If at least two of the three models consider the question qualified, it will be judged as a qualified question; if it fails to meet this standard, the method will regenerate or optimize the question.
[0186] Error tolerance setting: During the voting process, error tolerance parameters are set. When the judgment results between models differ greatly, automatic adjustments or re-evaluations are made to ensure that the quality of the questions meets expectations.
[0187] 2. Accuracy check:
[0188] Conduct a detailed check on the accuracy of the generated questions to ensure that the question content and options meet the definition and requirements of the knowledge points and avoid errors or deviations. Specifically include:
[0189] Question content check: Check the wording of the questions to ensure that each question can accurately and clearly express the knowledge points tested and is consistent with the extracted knowledge points and main content.
[0190] Option consistency check: For multiple-choice questions, ensure the rationality and consistency of the options to avoid logical contradictions or erroneous information. In particular, in the design of interference options, the rationality of the options is judged through the model to ensure that at least one interference option can effectively interfere with the choice of the correct answer, rather than being too obvious or irrelevant.
[0191] 3. Logical rationality test:
[0192] The generated questions must not only be accurate, but also logically reasonable. The logic and internal consistency of each question are verified by the following methods:
[0193] Logical consistency check: Analyze the logical structure of the question to ensure that there is no conflict between the various parts of the question. For multiple-choice questions, check the consistency of each option with the question requirements to avoid inconsistencies between options or between the question and the options.
[0194] Reasoning process verification: For questions involving reasoning or calculation, the reasoning chain is verified to ensure that the reasoning process of the question is logical. Any illogical or ambiguous questions will be automatically marked as "pending revision".
[0195] 4.Consistency and diversity check:
[0196] The generated questions are checked for consistency to ensure that the questions in the dataset are balanced and consistent in style, difficulty, and coverage. Specifically, they include:
[0197] Difficulty consistency check: By comparing the difficulty scores of the questions, ensure that the difficulty distribution of the questions is reasonable and in line with expectations. If there are too simple or too complex questions in the data set, adjust the question content to ensure the overall balance of difficulty.
[0198] Knowledge point coverage check: Check whether the questions can cover various knowledge points, especially for those questions involving multiple subjects or cross-domain knowledge, to ensure the diversity and comprehensiveness of question generation.
[0199] 5. Automatic optimization and regeneration:
[0200] When it is detected that a question does not meet the standards, the question will be automatically optimized or regenerated based on the detection results:
[0201] Automatic optimization: Automatically adjust or optimize some minor quality issues, such as unclear expressions, inconsistent options, etc., to ensure the accuracy and logic of the questions.
[0202] Regeneration mechanism: For serious problems that cannot be corrected through optimization, the regeneration mechanism is triggered to regenerate qualified assessment questions. When regenerating, the core knowledge points and structure of the original questions are referenced, and the quality inspection feedback is combined to ensure that the generated questions meet high standards in accuracy, logic and complexity.
[0203] 6. Output quality assessment report:
[0204] After the quality inspection is completed, a detailed quality assessment report is generated, including the evaluation results of each question in terms of accuracy, consistency, logical rationality, etc. These reports can not only help developers understand the quality of the generated questions, but also provide data support for further optimization methods.
[0205] Quality indicators: The report will list various quality detection indicators, such as accuracy, logical consistency, option interference, etc., to help users fully understand the quality of the questions.
[0206] Quality improvement suggestions: Based on the detection results, targeted improvement suggestions are provided to help users quickly correct low-quality questions.
[0207] On the other hand, this embodiment provides a large language model capability evaluation system based on dynamic data evaluation, such as Figure 2 As shown, including:
[0208] The knowledge point and main content extraction module is used to obtain the topic input by the user and extract the core knowledge points and main content from the topic;
[0209] An online search and knowledge elaboration module is used to perform online search based on the core knowledge points and main content using a pre-trained large language model to generate knowledge elaborations related to the topic;
[0210] A question design module, for generating assessment questions based on the core knowledge points, main content and knowledge elaboration;
[0211] A complexity control module is used to adjust and optimize the difficulty of the evaluation questions to obtain the final evaluation questions;
[0212] A multi-dimensional assessment module, used to conduct a multi-dimensional ability assessment on the final assessment topic and obtain an assessment result;
[0213] The quality detection module is used to perform quality detection on the final evaluation questions and obtain detection results.
[0214] In this embodiment, a specific workflow of a large language model capability assessment system based on dynamic data assessment is as follows:
[0215] 1. Knowledge point and main idea extraction module;
[0216] Function: This module is responsible for extracting the core knowledge points and main content from the input questions to ensure that the generated evaluation samples are consistent with the core ideas of the original questions.
[0217] Working principle:
[0218] The question entered is: "Which of the following regular expressions is equivalent to (a* + b)*(c + d)?"
[0219] Analyze the questions through natural language processing (NLP) technology to extract the core knowledge points and themes:
[0220] Knowledge points: regular expressions, operators, Kleene star, union operator, etc.
[0221] Purpose: Find a regular expression equivalent to the expression (a* + b)*(c + d).
[0222] Output:
[0223] Knowledge points: ["Regular expressions and their operators", "Concatenation inregular expressions", "Union operator (+) in regular expressions", "Kleenestar (*) in regular expressions", "Equivalence of regular expressions"].
[0224] Purpose: Finding the regular expression equivalent to (a + b)(c + d).
[0225] 2. Online search and knowledge elaboration module;
[0226] Function: Based on the extracted knowledge points, relevant background information is obtained through online retrieval to generate detailed knowledge elaboration to assist in the subsequent question design and evaluation.
[0227] Working principle:
[0228] Knowledge point: Regular expressions and their operators.
[0229] Purpose: Finding the regular expression equivalent to (a* + b)*(c + d).
[0230] The system obtains relevant information about regular expressions through external search engines and generates detailed background information based on it, such as the definition of Kleene asterisks, operators of regular expressions, etc.
[0231] Output:
[0232] Background information: "Regular expressions are sequences of characters that define search patterns in strings. Basic operators include concatenation, which puts expressions in sequence; selection, using | or + to represent different choices; and the Kleene star, *, which represents zero or more repetitions. Mastering these operators is essential to building expressions that effectively represent specific string patterns."
[0233] 3. Question design module;
[0234] Function: Dynamically generate new assessment questions with the same complexity as the original questions based on knowledge points and main content to ensure the accuracy and fairness of the assessment tasks.
[0235] Working principle:
[0236] Based on the extracted knowledge points and themes, the system generates similar assessment questions. For example:
[0237] Original question: Which of the following regular expressions is equivalent to(a* + b)*(c + d)?
[0238] New question: Which of the following regular expressions correctly represents all strings that start with zero or more repetitions of 'x' or 'y', followed by exactly two repetitions of 'm' or 'n', and end with a singleoccurrence of 'p' or 'q'?
[0239] Generated question: Which of the following regular expressions correctly represents all strings that start with zero or more repetitions of 'x' or 'y', followed by exactly two repetitions of 'm' or 'n', and end with a singleoccurrence of 'p' or 'q'?
[0240] Options:
[0241] A: (x|y)*((m|n)(m|n))(p|q);
[0242] B: ((x|y)*(m|n))^2(p|q);
[0243] C: (x* + y*)((m + n)(m + n))(p + q);
[0244] D: (x|y)^2((m|n)(m|n))(p|q);
[0245] The answer is: A.
[0246] 4. Complexity control module;
[0247] Function: This module is responsible for evaluating the difficulty of the generated questions and dynamically adjusting the difficulty to ensure that the complexity of the questions remains consistent with the original questions.
[0248] Working principle:
[0249] After evaluating the complexity of the original question, the complexity of generating the new question is controlled according to the evaluation result.
[0250] For example, when the original question is of high complexity, the generated new question will maintain a certain complexity and increase the reasoning steps.
[0251] If the complexity of the question is too low, the system will automatically increase the difficulty of the question and add more reasoning and analysis requirements.
[0252] Effect:
[0253] Ensure that the complexity of the generated questions is aligned with the difficulty of the original dataset and ensure fair evaluation.
[0254] 5. Multi-dimensional evaluation module;
[0255] Function: Through Bloom's taxonomy, LLMs are comprehensively evaluated from six dimensions: memory, comprehension, application, analysis, evaluation and creation.
[0256] Working principle:
[0257] Each generated question will be evaluated according to its cognitive level, ensuring that the model's capabilities at all levels are fully tested. For example:
[0258] Memory: Tests whether the model can remember the definition of a regular expression.
[0259] Comprehension: Evaluate whether the model can explain the meaning of the expression.
[0260] Application: Examine whether the model can apply the learned expressions to new situations.
[0261] Output:
[0262] Levels: memory, understanding, application, analysis, evaluation, creation;
[0263] Question: Each level generates a corresponding question and gives the correct answer.
[0264] 6.Quality inspection module;
[0265] Function: Through the voting mechanism of multilingual models, the quality of generated data is ensured, and unqualified questions are corrected or regenerated.
[0266] Working principle:
[0267] The quality of the generated questions is checked by calling multiple LLMs to ensure the logic, consistency and accuracy of the questions.
[0268] If the detection fails, the system will automatically optimize or regenerate the question.
[0269] Effect:
[0270] Efficiently ensure the quality of generated questions and avoid erroneous or unreasonable questions.
[0271] This embodiment can effectively solve the problems of data pollution, evaluation accuracy and complexity control in the prior art, and can significantly improve the reliability, fairness and comprehensiveness of LLMs evaluation. The following are the main advantages of this embodiment:
[0272] 1. Reduce data pollution and improve the reliability and fairness of evaluation results;
[0273] Existing evaluation methods generally have data contamination problems, especially when the evaluation data overlaps with the model's training data, the reliability of the evaluation results is seriously affected. By adopting dynamic data generation and complexity control mechanisms, this embodiment can generate new pollution-free data sets during the evaluation process, thereby effectively avoiding the interference of evaluation data contamination on the results and ensuring the objectivity and fairness of the evaluation results.
[0274] Specifically, this embodiment ensures that the generated evaluation questions are consistent with the core ideas of the original questions through the knowledge point and main idea extraction module, and introduces a multi-dimensional quality detection module to further ensure the accuracy and consistency of the evaluation samples through the multi-language model voting mechanism, thereby reducing the evaluation bias caused by data pollution.
[0275] 2. Improve the ability to control complexity and achieve precise difficulty matching;
[0276] When evaluating LLMs, it is very important to ensure that the difficulty of the evaluation questions is consistent with that of the original data set, especially in high-level cognitive tasks such as reasoning, analysis, and evaluation, where the complexity control of the questions is particularly critical. In the existing technology, although some methods can generate new evaluation samples, they often lack an effective complexity control mechanism, resulting in the inability to accurately align the difficulty of the questions with the original data, which in turn affects the accuracy of the evaluation.
[0277] This embodiment introduces a complexity control module to dynamically adjust the difficulty during the process of generating questions, ensuring that the generated questions not only meet the knowledge point requirements, but also remain consistent with the original data set in terms of complexity. Specifically, when the complexity of the generated questions is too high, the system will simplify the question design; when the complexity is too low, the difficulty of the questions will be increased by increasing the reasoning and analysis requirements.
[0278] 3. Provide multi-dimensional competency assessment to comprehensively measure the performance of LLMs;
[0279] Existing evaluation methods mostly focus on the basic ability assessment of models, such as memory and understanding, but lack the examination of models at higher cognitive levels, such as analysis, evaluation, and creation. Bloom's taxonomy provides an effective framework that can comprehensively evaluate model capabilities from six cognitive levels: memory, understanding, application, analysis, evaluation, and creation. However, most existing technologies fail to make full use of this framework.
[0280] This embodiment introduces a multi-dimensional assessment module to comprehensively assess LLMs from six cognitive levels based on Bloom's taxonomy, ensuring that the model's capabilities are not only measured in terms of basic knowledge, but also accurately assessed in terms of higher-level reasoning, analysis, and creative abilities.
[0281] 4. Improve the automation and stability of data generation and optimize the evaluation process;
[0282] In the prior art, the data generation process often requires a lot of manual intervention, and the quality and stability of the generated data are difficult to guarantee. This embodiment uses an automated data generation process and a multi-language model voting mechanism to automatically generate high-quality evaluation data without manual intervention, and ensures the consistency and accuracy of the data through multi-dimensional quality testing.
[0283] Specifically, the quality detection module of this embodiment uses multiple LLMs to evaluate data quality, and ensures that the generated data meets the requirements through a voting mechanism, thereby further optimizing the accuracy and stability of the generated data.
[0284] 5. Improve the scalability and adaptability of assessment;
[0285] This embodiment adopts a flexible modular design in the evaluation process, so that the evaluation method can adapt to different tasks and application scenarios. Whether it is for mathematical reasoning tasks, dialogue tasks, or scientific research ability assessment, the technical solution of this embodiment can flexibly adapt to various evaluation needs by adjusting the input of the evaluation task, the type and complexity of the generated questions.
[0286] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.
Claims
1. A large language model capability evaluation method based on dynamic data evaluation, characterized in that: include: Obtain the topic input by the user, and extract the core knowledge points and main content from the topic; Based on the core knowledge points and main content, a pre-trained large language model is used to perform online retrieval to generate detailed knowledge descriptions related to the topic; Generate assessment questions based on the core knowledge points, main content and knowledge elaboration; Adjust and optimize the difficulty of the assessment questions to obtain final assessment questions; Perform multi-dimensional capability assessment and quality inspection on the final assessment questions, obtain assessment results, and complete the capability assessment of the large language model.
2. The large language model capability assessment method based on dynamic data assessment according to claim 1, characterized in that: The core knowledge points and main contents extracted from the above topics include: Analyze the question by natural language processing technology to obtain analysis information; The core knowledge points and main content are extracted based on the parsed information using few-shot learning technology.
3. The large language model capability assessment method based on dynamic data assessment according to claim 1, characterized in that: Based on the core knowledge points and main content, a pre-trained large language model is used to perform online retrieval to generate detailed knowledge related to the topic, including: Conduct online search based on the core knowledge points and main content to obtain relevant background information; Input the core knowledge points and main content into the pre-trained large language model and output the knowledge elaboration; The relevant background information and knowledge elaboration are integrated to obtain the knowledge elaboration.
4. The large language model capability assessment method based on dynamic data assessment according to claim 1, characterized in that: Based on the core knowledge points, main content and knowledge elaboration, the generated assessment questions include: Integrate the core knowledge points, main content and knowledge details to obtain integrated information; Based on the integrated information, a preliminary question framework is generated according to question type and assessment objectives; The complexity of the preliminary question framework is adjusted and the repeatability is tested, and Bloom's cognitive hierarchy is used to obtain a multi-level question framework; Perform quality inspection on the multi-level question framework to obtain the evaluation questions.
5. The large language model capability assessment method based on dynamic data assessment according to claim 1, characterized in that: The difficulty of the assessment questions is adjusted and optimized, and the final assessment questions include: Performing a complexity assessment on the assessment topic and obtaining an assessment result; Performing complexity control based on the evaluation result to obtain the evaluation topic after the control; Multiple large language models are used to verify and provide feedback on the adjusted evaluation questions, and the adjusted evaluation questions are optimized based on the feedback results to obtain the final evaluation questions.
6. The large language model capability assessment method based on dynamic data assessment according to claim 5, characterized in that: The difficulty of the evaluation questions is adjusted and optimized, and obtaining the final evaluation questions also includes: Performing a difficulty evaluation on the evaluation questions, and sorting the evaluation questions according to the evaluation results to obtain a difficulty distribution; The difficulty span is determined based on the difficulty distribution. When the difficulty span exceeds a preset difficulty span, the difficulty of the assessment questions is smoothed to obtain assessment questions with a balanced difficulty distribution.
7. The large language model capability assessment method based on dynamic data assessment according to claim 1, characterized in that: The multi-dimensional ability assessment of the final assessment topic includes: Using Bloom's cognitive hierarchy, generating classification questions of different cognitive levels based on the final assessment questions; Evaluate the performance of the large language model at different cognitive levels including classification questions from multiple dimensions to obtain multi-dimensional evaluation results.
8. The large language model capability assessment method based on dynamic data assessment according to claim 1, characterized in that: The quality check of the final assessment questions includes: Use several large language models to test the quality, accuracy, logic rationality, consistency and diversity of each final evaluation question and obtain the test results; The detection result is output, and the final evaluation question is automatically optimized and regenerated based on the detection result.
9. A large language model capability evaluation system based on dynamic data evaluation, characterized in that: include: The knowledge point and main content extraction module is used to obtain the topic input by the user and extract the core knowledge points and main content from the topic; An online search and knowledge elaboration module is used to perform online search based on the core knowledge points and main content using a pre-trained large language model to generate knowledge elaborations related to the topic; A question design module, for generating assessment questions based on the core knowledge points, main content and knowledge elaboration; A complexity control module is used to adjust and optimize the difficulty of the evaluation questions to obtain the final evaluation questions; A multi-dimensional assessment module, used to conduct a multi-dimensional ability assessment on the final assessment topic and obtain an assessment result; The quality detection module is used to perform quality detection on the final evaluation questions and obtain detection results.
Citation Information
Patent Citations
Domain knowledge mastery degree self-testing method and system of large language model
CN118095257A
Assessment method and device of large language model and computer equipment
CN118535443A
Automatic defect detection system and method for large language model
CN118733455A
Semantic model-based power grid dispatching adaptive evaluation question generation method and system
CN118939789A
Model evaluation method and device, electronic equipment and storage medium
CN119760376A