A large language model capability evaluation method and system based on dynamic data evaluation
By extracting core knowledge points and main content, generating assessment questions and conducting multi-dimensional assessments, we solve the problems of data pollution and narrow scope of application in the evaluation of large language models, and achieve more accurate and reliable evaluation results.
Patent Information
- Application Number
- CN202510481116.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Existing large language model evaluation methods have data pollution problems, inaccurate evaluation results, narrow scope of application, unstable data generation quality, lack of a multi-dimensional evaluation system, and lack of effective question complexity control capabilities.
By obtaining user input questions, extracting core knowledge points and main content, using pre-trained large language models to conduct online retrieval to generate knowledge descriptions, generating assessment questions, and performing difficulty adjustment and multi-dimensional ability assessment, introducing complexity control and multi-model verification mechanisms to ensure the quality and consistency of assessment data.
It improves the objectivity and accuracy of the evaluation, comprehensively examines the performance of the model at different cognitive levels, enhances the reliability and fairness of the evaluation results, expands the scope of application of the evaluation method, and reduces the impact of data contamination.
Smart Images

Figure CN119988914B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine learning technology, and in particular to a large language model capability assessment method and system based on dynamic data assessment. Background Art
[0002] With the widespread application of large language models (LLMs) in natural language processing, they have achieved remarkable results in various tasks. However, in practical use, LLMs face a series of evaluation challenges, particularly the potential for data contamination during the evaluation process. Current evaluation methods for large LLMs primarily rely on standard datasets, which typically consist of large amounts of training data. The core task of evaluation is to verify the model's performance on specific tasks. Evaluation of LLMs generally focuses on factors such as task completion, answer accuracy, and processing time. Common evaluation methods include manual annotation-based evaluation, automatic scoring algorithms, and comparative analysis of model outputs. As the large-scale internet corpora used in LLM training continue to expand, issues with the reliability and fairness of their evaluation have gradually emerged. The root cause of these issues is data contamination. In particular, when the evaluation dataset overlaps with the training data, the evaluation results may not accurately reflect the model's generalization ability and may even affect fair comparisons between different models.
[0003] Existing technologies focus primarily on two main areas: data contamination detection and dynamic data evaluation. Data contamination detection methods address the contamination problem by identifying overlaps between model outputs and training data. However, because some closed-source language models, such as the GPT family, employ complex filtering mechanisms during their generation process, these methods often struggle to detect hidden data contamination. For example, while model outputs may appear uncontaminated, in some cases the model may still produce outputs highly similar to its training data, leading to inaccurate evaluation results.
[0004] On the other hand, dynamic data evaluation methods attempt to circumvent the limitations of traditional evaluation methods by generating new, uncontaminated datasets. Representative existing technologies include DYVAL, KIEval, LatestEval, SciEval, and the methods of Ying J et al. These methods each explore the potential of dynamic data evaluation from different perspectives and provide valuable tools for LLMs ability assessment. However, in practical applications, they still face some technical challenges and shortcomings, mainly reflected in the narrow scope of application, unstable data generation quality, insufficient control over question complexity, lack of a multi-dimensional evaluation system, and data contamination. Therefore, the present invention proposes a large language model ability assessment method and system based on dynamic data evaluation. Summary of the Invention
[0005] The purpose of the present invention is to provide a large language model capability assessment method and system based on dynamic data evaluation. By dynamically generating evaluation samples, controlling complexity and performing multi-dimensional evaluation, the quality and consistency of evaluation data are ensured, and the reliability and fairness of LLMs capability assessment are improved.
[0006] To achieve the above objectives, the present invention provides a method for evaluating the capability of a large language model based on dynamic data evaluation, comprising:
[0007] Obtain the topic input by the user and extract the core knowledge points and main content from the topic;
[0008] Based on the core knowledge points and main content, a pre-trained large language model is used to conduct online retrieval to generate detailed knowledge related to the topic;
[0009] Generate assessment questions based on the core knowledge points, main content and knowledge elaboration;
[0010] Adjusting and optimizing the difficulty of the assessment questions to obtain final assessment questions;
[0011] Conduct multi-dimensional capability assessment and quality testing on the final assessment questions, obtain assessment results, and complete the capability assessment of the large language model.
[0012] Optionally, extracting core knowledge points and main content from the topic includes:
[0013] Analyze the question using natural language processing technology to obtain analytical information;
[0014] The core knowledge points and main content are extracted based on the parsed information using few-shot learning technology.
[0015] Optionally, based on the core knowledge points and main content, a pre-trained large language model is used to perform online retrieval to generate detailed knowledge related to the topic, including:
[0016] Conduct online searches based on the core knowledge points and main content to obtain relevant background information;
[0017] Input the core knowledge points and main content into the pre-trained large language model and output knowledge elaboration;
[0018] The relevant background information and knowledge elaboration are integrated to obtain the knowledge elaboration.
[0019] Optionally, based on the core knowledge points, main content and knowledge elaboration, generating assessment questions includes:
[0020] Integrate the core knowledge points, main content and knowledge details to obtain integrated information;
[0021] Based on the integrated information, a preliminary question framework is generated according to question type and assessment objectives;
[0022] The complexity of the preliminary question framework is adjusted and the repeatability is tested, and Bloom's cognitive hierarchy is used to obtain a multi-level question framework;
[0023] Perform quality inspection on the multi-level question framework to obtain the evaluation questions.
[0024] Optionally, adjusting and optimizing the difficulty of the assessment topic to obtain the final assessment topic includes:
[0025] Performing a complexity assessment on the assessment topic and obtaining an assessment result;
[0026] Performing complexity control based on the evaluation results to obtain the adjusted evaluation questions;
[0027] Multiple large language models are used to verify and provide feedback on the adjusted evaluation questions, and the adjusted evaluation questions are optimized based on the feedback results to obtain the final evaluation questions.
[0028] Optionally, adjusting and optimizing the difficulty of the assessment questions to obtain the final assessment questions further includes:
[0029] Performing a difficulty evaluation on the assessment questions, and sorting the assessment questions according to the evaluation results to obtain a difficulty distribution;
[0030] The difficulty span is determined based on the difficulty distribution. When the difficulty span exceeds a preset difficulty span, the difficulty of the assessment questions is smoothed to obtain assessment questions with a balanced difficulty distribution.
[0031] Optionally, performing a multi-dimensional ability assessment on the final assessment topic includes:
[0032] Using Bloom's cognitive hierarchy, generate classification questions of different cognitive levels based on the final assessment questions;
[0033] Evaluate the performance of large language models at different cognitive levels, including classification questions, from multiple dimensions to obtain multi-dimensional evaluation results.
[0034] Optionally, performing quality inspection on the final assessment topic includes:
[0035] Use several large language models to test the quality, accuracy, logic rationality, consistency and diversity of each final assessment question and obtain the test results;
[0036] The detection result is output, and the final assessment question is automatically optimized and regenerated based on the detection result.
[0037] In another aspect, the present invention provides a large language model capability assessment system based on dynamic data assessment, comprising:
[0038] The knowledge point and main theme extraction module is used to obtain the topic input by the user and extract the core knowledge points and main theme content from the topic;
[0039] An online search and knowledge elaboration module is used to perform online search based on the core knowledge points and main content using a pre-trained large language model to generate knowledge elaboration related to the topic;
[0040] A question design module, for generating assessment questions based on the core knowledge points, main content and knowledge elaboration;
[0041] A complexity control module is used to adjust and optimize the difficulty of the assessment questions to obtain the final assessment questions;
[0042] A multi-dimensional assessment module, used to perform a multi-dimensional ability assessment on the final assessment topic and obtain an assessment result;
[0043] The quality detection module is used to perform quality detection on the final evaluation questions and obtain detection results.
[0044] The beneficial effects of the present invention are:
[0045] The present invention introduces a complexity control mechanism to ensure that the difficulty of the generated questions is consistent with the original data, thereby improving the objectivity and accuracy of the evaluation; through a multi-dimensional evaluation framework, the performance of the model at different cognitive levels is comprehensively examined, especially the evaluation of high-order abilities such as reasoning, evaluation, and creativity; through a quality detection mechanism, the stability and accuracy of the generated data are improved to ensure the reliability of the evaluation results; by introducing a multi-model verification mechanism, the pollution detection capability is optimized to reduce the impact of data pollution on the evaluation results; it can expand the scope of application of existing evaluation methods and provide cross-domain and flexible evaluation solutions; through the above technical content, the present invention ensures the quality and consistency of the evaluation data and improves the reliability and fairness of LLMs ability evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0047] Figure 1 This is a flow chart of a method for evaluating the capability of a large language model based on dynamic data evaluation according to an embodiment of the present invention;
[0048] Figure 2 This is a schematic diagram of the structure of a large language model capability assessment system based on dynamic data assessment in an embodiment of the present invention. DETAILED DESCRIPTION
[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0050] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0051] The existing technology has the following defects:
[0052] 1. Narrow scope of application:
[0053] Most existing dynamic data evaluation methods rely on pre-set data models and fixed question generation mechanisms for specific fields, which limits their applicability in other fields or tasks. Although DYVAL is effective in the field of mathematics, it does not provide broad support for other fields; KIEval is more suitable for conversational scenarios and is difficult to adapt to other types of evaluation tasks. Existing technologies lack cross-domain flexibility and cannot provide a unified solution for diverse evaluation tasks. The evaluation requirements between different tasks vary greatly, and existing methods rely more on fixed question design frameworks and fail to fully consider flexibility and scalability issues.
[0054] 2. Unstable data generation quality:
[0055] The generated data often lacks high quality and may contain redundancy, bias, or even fail to accurately reflect the performance of LLMs in certain specific tasks. For example, when generating mathematical reasoning questions, while DYVAL can generate high-quality questions, there is still room for improvement in terms of task complexity and accuracy control. These methods fail to implement a complete quality control mechanism when generating questions, resulting in the generated data not accurately reflecting the model's capabilities in complex scenarios, and may even produce evaluation samples that do not meet the original task requirements. This instability affects the reliability of the final evaluation results, especially in scenarios that require high accuracy.
[0056] 3. Insufficient ability to control the complexity of the questions:
[0057] Most existing dynamic data evaluation methods lack effective mechanisms for controlling the complexity of questions, particularly ensuring that the complexity of generated questions aligns with the difficulty level of the original data. For example, SciEval and KIEval test a model's complex capabilities by designing new questions, but these methods still lack sophisticated mechanisms for difficulty control, resulting in the difficulty of questions not fully meeting the requirements of the target assessment. Consequently, the evaluation results of existing methods may not objectively reflect the model's capabilities at different difficulty levels, making it difficult to accurately assess the model's high-level cognitive abilities, especially in the evaluation of reasoning or creative tasks.
[0058] 4. Lack of a multi-dimensional evaluation system:
[0059] Most current dynamic data evaluation methods focus on single-dimensional testing, typically focusing on basic dimensions such as model accuracy and speed, while neglecting comprehensive assessment of models across multiple levels and complex cognitive tasks. Bloom's taxonomy provides a more refined classification of cognitive hierarchies, but existing evaluation methods often fail to fully leverage this framework for comprehensive assessment, particularly for higher-level cognitive abilities such as analysis, evaluation, and creativity. For example, a simple strategy proposed by Ying J et al. can automatically update evaluation datasets, but it fails to consider how to comprehensively evaluate model capabilities through questions at different levels.
[0060] 5. Data pollution problem:
[0061] Data contamination remains a serious challenge in the evaluation field. Existing contamination detection methods attempt to address this issue by identifying overlaps between training and evaluation data. However, due to the specialized filtering mechanisms used in the generation of closed-source language models, such as the GPT series, these methods are limited in their effectiveness in catching implicit contamination. Even methods specifically designed to avoid contamination, such as LatestEval, cannot completely eliminate the impact of data contamination. Existing technologies are particularly limited when dealing with certain types of implicit contamination.
[0062] To solve the above technical problems, on the one hand, this embodiment provides a large language model capability evaluation method based on dynamic data evaluation, such as Figure 1 As shown, including:
[0063] Obtain the topic input by the user and extract the core knowledge points and main content from the topic;
[0064] Based on the core knowledge points and main content, a pre-trained large language model is used to conduct online retrieval to generate detailed knowledge related to the topic;
[0065] Generate assessment questions based on the core knowledge points, main content and knowledge elaboration;
[0066] Adjusting and optimizing the difficulty of the assessment questions to obtain final assessment questions;
[0067] Conduct multi-dimensional capability assessment and quality testing on the final assessment questions, obtain assessment results, and complete the capability assessment of the large language model.
[0068] Specifically, this embodiment introduces a complexity control mechanism to ensure that the difficulty of generated questions is consistent with the original data, thereby improving the objectivity and accuracy of the assessment. A multi-dimensional assessment framework comprehensively examines the model's performance at different cognitive levels, particularly in the assessment of higher-level abilities such as reasoning, evaluation, and creativity. A quality inspection mechanism improves the stability and accuracy of generated data, ensuring the reliability of the assessment results. A multi-model verification mechanism optimizes contamination detection capabilities and reduces the impact of data contamination on assessment results. This expands the scope of application of existing assessment methods and provides a cross-domain, flexible assessment solution. Through the above technical content, the quality and consistency of assessment data can be ensured, improving the reliability and fairness of LLMs' ability assessment.
[0069] Furthermore, extracting core knowledge points and main content from the topic includes:
[0070] Analyze the question using natural language processing technology to obtain analytical information;
[0071] The core knowledge points and main content are extracted based on the parsed information using few-shot learning technology.
[0072] The specific steps include:
[0073] Question parsing: Receive questions or data sets input by users and parse the questions using natural language processing (NLP) technology.
[0074] Core knowledge point extraction: Through pre-trained models or rule-based algorithms, key concepts or categories related to knowledge are automatically extracted from the questions, ensuring that these knowledge points can represent the core content of the questions without containing details or redundant information irrelevant to the core of the questions.
[0075] Main idea extraction: In addition to extracting key knowledge points, we also extract the main content of the question, summarizing the overall meaning of the question and avoiding lengthy statements. The extracted main content should cover the question background, key concepts, and the core reasoning or knowledge application required.
[0076] During the extraction process, we used a few-shot learning technique, fine-tuning the method using a small number of examples to further improve the accuracy of knowledge point and main idea extraction. Few-shot learning enables this method to quickly adapt to and extract the most representative knowledge points and main ideas across a wide range of question types, ensuring that the generated assessment questions align with the original data in terms of core ideas and knowledge points, providing a reliable foundation for subsequent assessment processes.
[0077] Furthermore, based on the core knowledge points and main content, a pre-trained large language model is used to perform online retrieval to generate detailed knowledge related to the topic, including:
[0078] Conduct online searches based on the core knowledge points and main content to obtain relevant background information;
[0079] Input the core knowledge points and main content into the pre-trained large language model and output knowledge elaboration;
[0080] The relevant background information and knowledge elaboration are integrated to obtain the knowledge elaboration.
[0081] This is achieved through the following steps:
[0082] 1. Online search:
[0083] First, based on the knowledge points output from the Knowledge Point and Theme Extraction module, we search external databases or the internet to obtain the latest relevant information, research results, definitions, or case studies. This background information further enriches the content of the questions, allowing the assessment questions to go beyond the superficial meaning and encompass a broader range of knowledge, enhancing the depth and accuracy of the model assessment.
[0084] 2. Knowledge Detail Generation:
[0085] After obtaining relevant background information, the pre-trained large language model is used to elaborate on the retrieved knowledge points. Specifically, it includes:
[0086] Detailed explanation: Based on the core concepts of the knowledge points, generate detailed explanations of related fields, explaining the definitions, application scenarios and related research progress of these concepts.
[0087] Examples and cases: Provide specific examples or cases based on actual application scenarios or academic research to help further understand the practical application of knowledge points.
[0088] Domain expansion: Based on the question type and difficulty, the method will also expand the domain knowledge related to the core knowledge points to ensure that the generated questions can cover multiple dimensions of the knowledge points.
[0089] 3. Detailed description of integrated knowledge generation:
[0090] Integrate the retrieved relevant background information with the model-generated knowledge exposition to form a concise, clear, and structured text. This text should focus on the core meaning of the knowledge point and exclude irrelevant or redundant information. This approach provides richer and deeper content support for each question, laying a solid knowledge foundation for subsequent assessment question generation.
[0091] 4. Detailed description of output knowledge:
[0092] The generated knowledge descriptions will be output along with the knowledge points and main content, serving as the core basis for subsequent question generation. Ultimately, these detailed descriptions will help diversify and control the complexity of questions in subsequent steps, ensuring the quality and usability of assessment data.
[0093] Furthermore, based on the core knowledge points, main content and knowledge elaboration, assessment questions are generated including:
[0094] Integrate the core knowledge points, main content and knowledge details to obtain integrated information;
[0095] Based on the integrated information, a preliminary question framework is generated according to question type and assessment objectives;
[0096] The complexity of the preliminary question framework is adjusted and the repeatability is tested, and Bloom's cognitive hierarchy is used to obtain a multi-level question framework;
[0097] Perform quality inspection on the multi-level question framework to obtain the evaluation questions.
[0098] This is achieved through the following steps:
[0099] 1. Input content integration:
[0100] First, we integrate core knowledge points, main content, and background information as the basic input for question generation. Based on this input, we can ensure that the generated assessment questions are closely designed around core knowledge.
[0101] 2. Generate the title framework:
[0102] After integrating the input content, a preliminary question framework is generated based on the question type, such as multiple-choice, fill-in-the-blank, short-answer, etc., and the assessment objectives, such as testing memory, comprehension, application, analysis, etc. Specifically, it includes:
[0103] Determine the question type: Select the appropriate question type based on the requirements of the task, such as multiple-choice questions, true-or-false questions, or short-answer questions.
[0104] Question design: Design specific questions based on the integrated input content. The question content should be closely centered around the knowledge points to ensure that the core knowledge of the test can be accurately covered.
[0105] Option Generation: This embodiment targets multiple-choice questions only. When generating multiple-choice questions, multiple options are generated based on the question content, including one correct answer and several distracting incorrect options. The distracting options must not only have a certain logical similarity to the correct answer but also fit the context of the original question.
[0106] 3. Complexity control and alignment:
[0107] In order to ensure that the difficulty of the generated questions is consistent with the original questions or meets the predetermined requirements, the complexity of the questions is controlled and aligned. Specifically,
[0108] Difficulty assessment of knowledge points: Generate matching questions based on the difficulty of the knowledge points to ensure that the difficulty level of the questions is reasonable.
[0109] Question complexity adjustment: Automatically adjust the complexity of questions based on task requirements. For example, when testing higher-level cognitive skills, such as analysis, evaluation, and creativity, questions may require more complex reasoning or be applied in real-world scenarios.
[0110] Generate multi-level questions: Generate multi-dimensional questions based on the six cognitive levels of Bloom's Taxonomy. Specifically, these questions cover the six dimensions of memory, comprehension, application, analysis, evaluation, and creativity, ensuring that questions in each dimension are of appropriate difficulty and type.
[0111] 4. Innovative and diverse topic design:
[0112] In addition to ensuring the basic requirements of the questions, we also use innovative design to ensure that the generated questions have a certain degree of diversity. For example:
[0113] Examining knowledge points from different angles: By changing the question types and applying different backgrounds or situations, students' multi-dimensional understanding of knowledge points can be examined.
[0114] Innovation in question format: You can design some innovative questions, such as case analysis questions, reasoning questions, comparison questions, etc., to increase the complexity and depth of the questions.
[0115] Avoid repetitive questions: Avoid generating questions that are too similar to the original questions, and ensure that the newly generated questions are innovative and unique in form and content.
[0116] 5. Quality inspection and optimization:
[0117] After generating the questions, we conduct a quality check to ensure that each question meets the following standards:
[0118] Clear logic: The question wording of each question should be clear and unambiguous to ensure that the subjects can understand and answer accurately.
[0119] The options should be reasonable and discriminatory: In multiple-choice questions, the options should be distracting to avoid overly obvious correct answers and ensure the depth of the examination.
[0120] Only one correct answer: Ensure that there is only one correct answer among the options for multiple-choice questions, and that the other options can reasonably interfere with students' judgment.
[0121] Core knowledge point coverage: Each question should accurately reflect the core knowledge point and be able to effectively test students' mastery of the knowledge point.
[0122] Furthermore, the difficulty of the assessment questions is adjusted and optimized to obtain the final assessment questions, including:
[0123] Performing a complexity assessment on the assessment topic and obtaining an assessment result;
[0124] Performing complexity control based on the evaluation results to obtain the adjusted evaluation questions;
[0125] Multiple large language models are used to verify and provide feedback on the adjusted evaluation questions, and the adjusted evaluation questions are optimized based on the feedback results to obtain the final evaluation questions.
[0126] It also includes: performing difficulty evaluation on the evaluation questions, and sorting the evaluation questions according to the evaluation results to obtain a difficulty distribution; judging a difficulty span based on the difficulty distribution, and when the difficulty span exceeds a preset difficulty span, smoothing the difficulty of the evaluation questions to obtain evaluation questions with a balanced difficulty distribution.
[0127] The specific steps are as follows:
[0128] 1. Assessment of topic complexity:
[0129] After generating the assessment questions, the complexity control module first performs a preliminary assessment of the complexity of the questions. This assessment is performed in the following ways:
[0130] Difficulty Rating: Each question is assigned a difficulty rating based on the type, structure, and difficulty of the knowledge points covered. The rating is based on factors such as the depth of the question, the complexity of the reasoning, and the abstractness of the concepts covered.
[0131] Benchmark data comparison: By comparing the generated questions with the questions in the original dataset, we evaluate whether the difficulty of the generated questions matches the original questions' level. For example, if the original questions are difficult, the generated questions need to be of a corresponding level of complexity. If the original dataset is difficult, the method adjusts the difficulty of the generated questions to meet the predetermined difficulty requirements.
[0132] Multi-dimensional Difficulty Assessment: In addition to the difficulty of the knowledge point, the complexity control module also considers the cognitive level of the question. For example, high-level cognitive dimensions of Bloom's Taxonomy, such as analysis, evaluation, and creation, typically correspond to higher complexity, while memory and comprehension questions are relatively simple. The method assigns different difficulty scores to questions based on cognitive level and adjusts them based on the overall assessment results.
[0133] 2. Complexity control mechanism:
[0134] After completing the initial complexity assessment, adopt corresponding complexity control strategies based on the assessment results to ensure that the difficulty of the generated questions meets expectations:
[0135] Complexity alignment: Automatically adjust the complexity of questions based on assessment results. For example, if a question is generated to be too difficult, the method will reduce its difficulty by simplifying the question's statement, reducing the number of reasoning steps, or reducing the number of concepts involved. If the question is too complex, the method will increase its difficulty, requiring students to conduct more in-depth analysis or reasoning.
[0136] Difficulty offset calculation: To ensure that the difficulty of the questions remains consistent with the original dataset, the difference in accuracy between the generated dataset and the original dataset is calculated, i.e., the offset. By calculating the offset, we can quantify the difference in difficulty between the generated questions and the original questions and adjust the generated questions accordingly. For example, if the average accuracy of the generated dataset is too high, we can increase the difficulty by modifying some questions or adding more challenging ones. Conversely, if the accuracy is too low, we can simplify the questions appropriately.
[0137] Automatic adjustment strategy: For questions with high or low complexity, the corresponding strategy will be automatically selected for adjustment. Specific adjustments include:
[0138] Question simplification: For questions that are too difficult, simplify the question stem, reduce multiple reasoning in the question, reduce the number of distractors in the options, or reduce the difficulty by rephrasing the question.
[0139] Question enhancement: For questions that are too easy, add more challenging situations or requirements, and add more analysis, reasoning or interdisciplinary knowledge points to ensure that the questions can cover more advanced cognitive levels.
[0140] 3. Multi-model verification and feedback mechanism:
[0141] We also introduced a multi-model verification mechanism, using multiple large language models (LLMs) to verify and provide feedback on the difficulty of generated questions. The specific operation is as follows:
[0142] Multi-model prediction: Generated questions are fed into multiple pre-set LLMs, and these models are used to predict the difficulty of the questions. Based on the feedback from the models, the difficulty of the questions can be further optimized to meet the predetermined requirements.
[0143] Feedback Adjustment: Fine-tune the complexity of questions based on feedback from multiple models. If multiple models find a question too easy or too complex, we make adjustments based on the feedback to ensure the question's difficulty remains within a reasonable range.
[0144] 4. Balanced Difficulty Distribution:
[0145] In order to ensure the balance of difficulty in the generated dataset, it is also necessary to monitor the difficulty distribution of the entire dataset. By analyzing the distribution of question difficulty, we can ensure that the difficulty level of the dataset is reasonable, and that questions from simple to complex can fully cover all cognitive levels. Specific operations include:
[0146] Question difficulty distribution: Based on the predetermined difficulty level requirements, the generated questions will be sorted according to the difficulty distribution, and ensure that questions of different levels are evenly covered.
[0147] Difficulty span adjustment: When the difficulty span of the generated questions is large, the difficulty of the questions is smoothed to ensure that the overall difficulty distribution of the questions does not have any extreme questions that are too prominent.
[0148] Furthermore, the multi-dimensional ability assessment of the final assessment topic includes:
[0149] Using Bloom's cognitive hierarchy, generate classification questions of different cognitive levels based on the final assessment questions;
[0150] Evaluate the performance of large language models at different cognitive levels, including classification questions, from multiple dimensions to obtain multi-dimensional evaluation results.
[0151] The specific steps are as follows:
[0152] 1. Cognitive level division and task definition:
[0153] In the multi-dimensional assessment module, each question is first divided into cognitive levels according to Bloom's taxonomy. Bloom's taxonomy divides the degree of knowledge mastery into six cognitive levels, from the most basic memory to the most complex creation, including:
[0154] Remembering: This measures students’ ability to recall and identify learned facts, terms, definitions and other basic knowledge.
[0155] Understanding: Tests whether students can understand the meaning of information and express or explain it in their own words.
[0156] Applying: Students are required to apply the knowledge they have learned to new situations to solve practical problems.
[0157] Analyzing: This examines whether students can break down complex content into its basic elements and understand the relationships between the parts.
[0158] Evaluating: Assessing whether students can judge and evaluate the validity, accuracy and value of information based on certain criteria.
[0159] Creating: Examines whether students can construct new ideas, methods or models based on existing knowledge.
[0160] 2. Matching the topic with the cognitive level:
[0161] Each question is automatically assigned a cognitive level based on its knowledge points, complexity, and the ability dimensions it tests. For example, simple fact recall questions are labeled as "Memory," while questions requiring complex reasoning or comprehensive analysis are labeled as "Analysis" or "Creation." This process ensures that each question not only tests basic knowledge but also involves higher-level thinking skills, thereby comprehensively evaluating the model's performance.
[0162] 3. Level coverage and topic distribution:
[0163] To ensure the comprehensiveness of the assessment, the number of questions for each level is reasonably allocated according to the six cognitive levels of Bloom's taxonomy. Specific operations include:
[0164] Level Coverage: Ensure that questions are tested at each level to avoid neglecting certain cognitive levels, such as creativity or evaluation. Questions at each cognitive level will cover a certain number of knowledge points and ensure that questions at that level fully reflect students' relevant abilities.
[0165] Balanced question count: When generating assessment data, the number of questions at each level is evenly distributed based on the predetermined difficulty and level requirements. For example, for more complex levels such as analysis and creativity, the number of questions should be appropriately reduced, while for basic levels such as memory and comprehension, the number of questions should be increased to ensure a reasonable difficulty distribution and meet the assessment objectives.
[0166] 4. Multi-dimensional evaluation generation:
[0167] After matching and assigning questions to cognitive levels, a multi-dimensional assessment generation mechanism generates questions based on different cognitive levels. These questions cover abilities at all levels, ensuring that the assessment fully reflects students' performance in each cognitive dimension.
[0168] Memory and comprehension questions: These are generally tests of basic knowledge, requiring students to memorize and understand core concepts. These questions may include basic formats such as fill-in-the-blank questions and multiple-choice questions.
[0169] Application questions: These require students to apply the knowledge they have learned to real-world scenarios or new situations. The questions are usually in the form of case analysis, problem solving, etc.
[0170] Analysis, evaluation and creativity questions: These questions are more complex and require students to conduct in-depth analysis, comparison or creative thinking. They usually involve longer texts, reasoning processes and solution designs.
[0171] 5. Model capability scoring and evaluation:
[0172] In the multidimensional assessment process, LLMs are scored for their performance at each cognitive level. By comparing the model's answers with the standard answers, the model's capabilities at each level can be assessed. For example:
[0173] At the "memory" level, the model is mainly evaluated on whether it can accurately recall and recognize basic facts and concepts.
[0174] At the “understanding” level, evaluate whether the model can accurately explain and restate the knowledge points.
[0175] At the "application" level, the model is evaluated to see whether it can apply knowledge to new scenarios and solve practical problems.
[0176] At the "analysis", "evaluation" and "creation" levels, the model's high-level cognitive ability is assessed by examining its reasoning process and creative output.
[0177] 6. Output multi-dimensional evaluation results:
[0178] After completing the above assessment, a multi-dimensional evaluation report is generated, covering the model's performance at each cognitive level. These results will help researchers and developers understand the model's capabilities at different levels and provide optimization strategies. For example, a model's low score at the "Creativity" level may indicate limitations in generating innovative answers and require further improvement.
[0179] Furthermore, the quality inspection of the final assessment questions includes:
[0180] Use several large language models to test the quality, accuracy, logic rationality, consistency and diversity of each final assessment question and obtain the test results;
[0181] The detection result is output, and the final assessment question is automatically optimized and regenerated based on the detection result.
[0182] The specific steps are as follows:
[0183] 1. Multilingual model voting mechanism:
[0184] First, we use multiple large language models (LLMs) to perform voting tests on the generated evaluation questions. These language models include multiple preset open-source and closed-source models, such as the GPT series, Qwen, and Doubao. We perform quality assessments on each question to evaluate the accuracy and consistency of its content.
[0185] Voting mechanism: Each generated question is fed into at least three different LLMs, and a comprehensive analysis is performed based on the outputs of these models. If at least two of the three models find the question qualified, it is considered qualified. If this criterion is not met, the method will regenerate or optimize the question.
[0186] Error tolerance settings: During the voting process, error tolerance parameters are set. When the judgment results between models differ significantly, automatic adjustments or re-evaluations are made to ensure that the quality of the questions meets expectations.
[0187] 2. Accuracy check:
[0188] Conduct detailed checks on the accuracy of generated questions to ensure that the question content and options conform to the definition and requirements of the knowledge points and avoid errors or deviations. Specifically, this includes:
[0189] Question content check: Check the wording of the questions to ensure that each question can accurately and clearly express the knowledge points tested and is consistent with the extracted knowledge points and main content.
[0190] Checking option consistency: For multiple-choice questions, ensure the rationality and consistency of the options to avoid logical contradictions or erroneous information. Specifically, when designing distractor options, use the model to assess their rationality, ensuring that at least one distractor effectively influences the correct answer and is not overly obvious or irrelevant.
[0191] 3. Logical rationality test:
[0192] The generated questions must not only meet the requirements of accuracy, but also logical rationality. The logic and internal consistency of each question are verified by the following methods:
[0193] Logical consistency check: Analyze the internal logical structure of the question to ensure there are no conflicts between the various parts of the question. For multiple-choice questions, check the consistency of each option with the question requirements to avoid inconsistencies between options or between the question and the options.
[0194] Reasoning Verification: For questions involving reasoning or calculation, we verify the reasoning chain to ensure that the reasoning process is logical. Any illogical or ambiguous questions will be automatically marked as "pending revision".
[0195] 4.Consistency and diversity check:
[0196] Perform consistency checks on the generated questions to ensure that the questions in the dataset maintain a certain balance and consistency in style, difficulty, and coverage. Specifically, this includes:
[0197] Difficulty consistency check: By comparing the difficulty scores of the questions, we ensure that the difficulty distribution of the questions is reasonable and in line with expectations. If there are questions in the dataset that are too easy or too complex, we adjust the question content to ensure the overall difficulty balance.
[0198] Knowledge point coverage detection: Check whether the questions can cover various knowledge points, especially for those involving multiple subjects or cross-domain knowledge, to ensure the diversity and comprehensiveness of question generation.
[0199] 5. Automatic optimization and regeneration:
[0200] When a question is detected to have non-compliant issues, the question will be automatically optimized or regenerated based on the detection results:
[0201] Automatic optimization: Automatically adjust or optimize minor quality issues, such as unclear expression and inconsistent options, to ensure the accuracy and logic of the questions.
[0202] Regeneration Mechanism: For serious issues that cannot be corrected through optimization, the regeneration mechanism is triggered to regenerate qualified assessment questions. During regeneration, the core knowledge points and structure of the original questions are referenced, and quality inspection feedback is also incorporated to ensure that the generated questions meet high standards in accuracy, logic, and complexity.
[0203] 6. Output quality assessment report:
[0204] After completing the quality test, a detailed quality assessment report is generated, including the evaluation results of each question in terms of accuracy, consistency, logical rationality, etc. These reports not only help developers understand the quality of the generated questions, but also provide data support for further optimization methods.
[0205] Quality indicators: The report will list various quality detection indicators, such as accuracy, logical consistency, option interference, etc., to help users fully understand the quality of the questions.
[0206] Quality improvement suggestions: Based on the test results, targeted improvement suggestions are provided to help users quickly correct low-quality questions.
[0207] On the other hand, this embodiment provides a large language model capability evaluation system based on dynamic data evaluation, such as Figure 2 As shown, including:
[0208] The knowledge point and main theme extraction module is used to obtain the topic input by the user and extract the core knowledge points and main theme content from the topic;
[0209] An online search and knowledge elaboration module is used to perform online search based on the core knowledge points and main content using a pre-trained large language model to generate knowledge elaboration related to the topic;
[0210] A question design module, for generating assessment questions based on the core knowledge points, main content and knowledge elaboration;
[0211] A complexity control module is used to adjust and optimize the difficulty of the assessment questions to obtain the final assessment questions;
[0212] A multi-dimensional assessment module, used to perform a multi-dimensional ability assessment on the final assessment topic and obtain an assessment result;
[0213] The quality detection module is used to perform quality detection on the final evaluation questions and obtain detection results.
[0214] The specific workflow of a large language model capability assessment system based on dynamic data assessment in this embodiment is as follows:
[0215] 1. Knowledge point and main idea extraction module;
[0216] Function: This module is responsible for extracting core knowledge points and main content from the input questions, ensuring that the generated evaluation samples are consistent with the core ideas of the original questions.
[0217] Working principle:
[0218] The input question is: "Which of the following regular expressions is equivalent to (a* + b)*(c + d)?"
[0219] Analyze questions using natural language processing (NLP) technology to extract core knowledge points and themes:
[0220] Knowledge points: regular expressions, operators, Kleene star, union operator, etc.
[0221] Purpose: Find a regular expression equivalent to the expression (a* + b)*(c + d).
[0222] Output:
[0223] Knowledge points: ["Regular expressions and their operators", "Concatenation inregular expressions", "Union operator (+) in regular expressions", "Kleenestar (*) in regular expressions", "Equivalence of regular expressions"].
[0224] Purpose: Finding the regular expression equivalent to (a + b)(c + d).
[0225] 2. Online search and knowledge elaboration module;
[0226] Function: Based on the extracted knowledge points, relevant background information is obtained through online retrieval to generate detailed knowledge elaboration to assist in subsequent question design and evaluation.
[0227] Working principle:
[0228] Knowledge point: Regular expressions and their operators.
[0229] Purpose: Finding the regular expression equivalent to (a* + b)*(c + d).
[0230] The system obtains relevant information about regular expressions through external search engines and generates detailed background information based on this, such as the definition of Kleene asterisks, regular expression operators, etc.
[0231] Output:
[0232] Background information: "Regular expressions are sequences of characters that define search patterns within strings. Basic operators include concatenation, which sequences expressions together; selection, using | or + to represent different choices; and the Kleene asterisk, *, which represents zero or more repetitions. Mastering these operators is essential for constructing expressions that effectively represent specific string patterns."
[0233] 3. Question design module;
[0234] Function: Dynamically generate new assessment questions with the same complexity as the original questions based on knowledge points and main content, ensuring the accuracy and fairness of the assessment tasks.
[0235] Working principle:
[0236] Based on the extracted knowledge points and themes, the system generates similar assessment questions. For example:
[0237] Original question: Which of the following regular expressions is equivalent to(a* + b)*(c + d)?
[0238] New question: Which of the following regular expressions correctly represents all strings that start with zero or more repetitions of 'x' or 'y', followed by exactly two repetitions of 'm' or 'n', and end with a singleoccurrence of 'p' or 'q'?
[0239] Generated question: Which of the following regular expressions correctly represents all strings that start with zero or more repetitions of 'x' or 'y', followed by exactly two repetitions of 'm' or 'n', and end with a singleoccurrence of 'p' or 'q'?
[0240] Options:
[0241] A: (x|y)*((m|n)(m|n))(p|q);
[0242] B: ((x|y)*(m|n))^2(p|q);
[0243] C: (x* + y*)((m + n)(m + n))(p + q);
[0244] D: (x|y)^2((m|n)(m|n))(p|q);
[0245] Answer: A.
[0246] 4. Complexity control module;
[0247] Function: This module is responsible for evaluating the difficulty of generated questions and dynamically adjusting the difficulty to ensure that the complexity of the questions remains consistent with the original questions.
[0248] Working principle:
[0249] After evaluating the complexity of the original question, the complexity of generating the new question is controlled based on the evaluation results.
[0250] For example, when the original question is of high complexity, the generated new question will maintain a certain complexity and increase the reasoning steps.
[0251] If the complexity of the question is too low, the system will automatically increase the difficulty of the question and add more reasoning and analysis requirements.
[0252] Effect:
[0253] Ensure that the complexity of the generated questions is aligned with the difficulty of the original dataset and ensure fair evaluation.
[0254] 5. Multi-dimensional evaluation module;
[0255] Function: Through Bloom's taxonomy, LLMs are comprehensively evaluated from six dimensions: memory, comprehension, application, analysis, evaluation and creation.
[0256] Working principle:
[0257] Each generated question will be evaluated according to its cognitive level, ensuring that the model's capabilities at all levels are fully tested. For example:
[0258] Memory: Tests whether the model can remember the definition of a regular expression.
[0259] Comprehension: Evaluate whether the model can explain the meaning of an expression.
[0260] Application: Examine whether the model can apply the learned expressions to new situations.
[0261] Output:
[0262] Levels: memory, understanding, application, analysis, evaluation, creation;
[0263] Question: Each level generates a corresponding question and gives the correct answer.
[0264] 6.Quality inspection module;
[0265] Function: Through the voting mechanism of multilingual models, the quality of generated data is ensured, and unqualified questions are corrected or regenerated.
[0266] Working principle:
[0267] The generated questions are quality checked by calling multiple LLMs to ensure the logic, consistency and accuracy of the questions.
[0268] If the detection fails, the system will automatically optimize or regenerate the question.
[0269] Effect:
[0270] Efficiently ensure the quality of generated questions and avoid incorrect or unreasonable questions.
[0271] This embodiment can effectively solve the problems of data pollution, evaluation accuracy, and complexity control in the existing technology, and can significantly improve the reliability, fairness, and comprehensiveness of LLMs evaluation. The following are the main advantages of this embodiment:
[0272] 1. Reduce data pollution and improve the reliability and fairness of evaluation results;
[0273] Existing evaluation methods are commonly plagued by data contamination, particularly when the evaluation data overlaps with the model's training data, which can severely impact the reliability of the evaluation results. By employing dynamic data generation and complexity control mechanisms, this embodiment generates new, uncontaminated datasets during the evaluation process, effectively preventing the impact of data contamination on the results and ensuring the objectivity and fairness of the evaluation results.
[0274] Specifically, this embodiment uses a knowledge point and main idea extraction module to ensure that the generated evaluation questions are consistent with the core idea of the original questions. At the same time, it introduces a multi-dimensional quality detection module and further ensures the accuracy and consistency of the evaluation samples through a multi-language model voting mechanism, thereby reducing the evaluation bias caused by data contamination.
[0275] 2. Improve complexity control capabilities and achieve precise difficulty matching;
[0276] When evaluating LLMs, it's crucial to ensure that the difficulty of the assessment questions matches that of the original dataset. This is especially true for high-level cognitive tasks like reasoning, analysis, and evaluation, where controlling the complexity of the questions is crucial. While some existing methods can generate new assessment samples, they often lack effective complexity control mechanisms, resulting in inaccurate alignment of the difficulty of the questions with the original data, which in turn affects the accuracy of the assessment.
[0277] This embodiment introduces a complexity control module to dynamically adjust the difficulty of questions during the question generation process, ensuring that the generated questions meet the knowledge requirements and maintain the same complexity as the original dataset. Specifically, if the complexity of the generated questions is too high, the system will simplify the question design; if the complexity is too low, the difficulty of the questions will be increased by increasing the reasoning and analysis requirements.
[0278] 3. Provide multi-dimensional competency assessment to comprehensively measure LLMs’ performance;
[0279] Existing evaluation methods often focus on assessing basic model capabilities, such as memory and comprehension, but lack an examination of higher cognitive levels, such as analysis, evaluation, and creativity. Bloom's taxonomy provides an effective framework for comprehensively evaluating model capabilities across six cognitive levels: memory, comprehension, application, analysis, evaluation, and creativity. However, most existing technologies fail to fully utilize this framework.
[0280] This embodiment introduces a multi-dimensional assessment module and comprehensively evaluates LLMs from six cognitive levels based on Bloom's taxonomy, ensuring that the model's capabilities are not only measured in terms of basic knowledge, but also accurately evaluated in terms of higher-level reasoning, analysis, and creativity.
[0281] 4. Improve the automation and stability of data generation and optimize the evaluation process;
[0282] In existing technologies, the data generation process often requires extensive manual intervention, and the quality and stability of the generated data are difficult to guarantee. This embodiment, through an automated data generation process and a multilingual model voting mechanism, can automatically generate high-quality evaluation data without manual intervention, and ensure the consistency and accuracy of the data through multi-dimensional quality testing.
[0283] Specifically, the quality detection module of this embodiment uses multiple LLMs to evaluate data quality, and ensures that the generated data meets the requirements through a voting mechanism, thereby further optimizing the accuracy and stability of the generated data.
[0284] 5. Improve the scalability and adaptability of assessments;
[0285] This embodiment employs a flexible modular design during the assessment process, allowing the assessment method to adapt to different tasks and application scenarios. Whether it's for mathematical reasoning tasks, conversational tasks, or scientific research ability assessments, the technical solution of this embodiment can flexibly adapt to various assessment needs by adjusting the input of the assessment task, the type of questions generated, and their complexity.
[0286] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A large language model capability evaluation method based on dynamic data evaluation, characterized in that: include: Obtain the question entered by the user, parse it using natural language processing technology, and use few-shot learning technology to extract core knowledge points and main content based on the parsed information; Conduct online searches based on the core knowledge points and main content to obtain relevant background information, including the latest data, research results, definitions, and cases; Input the core knowledge points and main content into a pre-trained large language model and output knowledge elaboration, wherein the knowledge elaboration includes detailed explanations, examples and cases, and domain expansion; Integrate the relevant background information and knowledge elaboration to obtain knowledge elaboration; Integrate the core knowledge points, main content and knowledge details to obtain integrated information; Based on the integrated information, a preliminary question framework is generated according to question type and assessment objectives; The preliminary question framework is adjusted for complexity and tested for repeatability, and corresponding multi-dimensional questions are generated according to the six cognitive levels of Bloom's taxonomy, including questions in the six dimensions of memory, comprehension, application, analysis, evaluation, and creation; Conduct quality checks on the multi-level question framework to obtain assessment questions; Performing a complexity evaluation on the evaluation questions, performing complexity control based on the evaluation results, calculating the accuracy difference between the generated data set and the original data set as an offset, quantifying the difficulty difference between the generated questions and the original questions using the offset, and adjusting the content of the generated questions based on the difficulty difference; Utilize multiple large language models to verify and provide feedback on the adjusted evaluation questions, and optimize the adjusted evaluation questions based on the feedback results to obtain the final evaluation questions, wherein the verification and feedback include: Multi-model prediction: The generated questions are input into multiple preset LLMs models, and the difficulty of the questions is predicted by the models. The difficulty of the questions is further optimized based on the feedback results of the models to ensure that the questions meet the predetermined requirements. Feedback adjustment: fine-tune the complexity of the questions based on feedback from multiple models; Obtaining the final assessment topic further includes: performing a difficulty assessment on the assessment topic, and sorting the assessment topic according to the assessment result to obtain a difficulty distribution; Determining the difficulty span based on the difficulty distribution; if the difficulty span exceeds a preset difficulty span, smoothing the difficulty of the assessment questions to obtain assessment questions with a balanced difficulty distribution; Using Bloom's cognitive hierarchy, we generate classification questions at different cognitive levels based on the final assessment questions, including memory, comprehension, application, analysis, evaluation, and creation. This evaluates the performance of the large language model at different cognitive levels, including the classification questions, from multiple dimensions, and obtains multi-dimensional evaluation results. Use several large language models to test the quality, accuracy, logic rationality, consistency and diversity of each final assessment question and obtain the test results; The multi-dimensional evaluation results and the detection results are combined to complete the capability evaluation of the large language model.
2. A system for implementing the large language model capability assessment method based on dynamic data assessment as described in claim 1, characterized in that: include: The knowledge point and main theme extraction module is used to obtain the topic input by the user and extract the core knowledge points and main theme content from the topic; An online search and knowledge elaboration module is used to perform online search based on the core knowledge points and main content using a pre-trained large language model to generate knowledge elaboration related to the topic; A question design module, for generating assessment questions based on the core knowledge points, main content and knowledge elaboration; A complexity control module is used to adjust and optimize the difficulty of the assessment questions to obtain the final assessment questions; A multi-dimensional assessment module, used to perform a multi-dimensional ability assessment on the final assessment topic and obtain an assessment result; The quality detection module is used to perform quality detection on the final evaluation questions and obtain detection results.
Citation Information
Patent Citations
Semantic model-based power grid dispatching adaptive evaluation question generation method and system
CN118939789A
Model evaluation method and device, electronic equipment and storage medium
CN119760376A