Dynamic evaluation method, device and equipment of large language model and medium

Through the dynamic evaluation framework, sample screening, enhancement and negative sample creation of large language models is solved, and the static evaluation method is difficult to comprehensively measure the model's capabilities in real interactive scenarios, achieving more accurate and flexible evaluation results.

CN120258144APending Publication Date: 2025-07-04THE THIRD RES INST OF MIN OF PUBLIC SECURITY +2
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510403498.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing large language model evaluation methods mainly rely on static evaluation, and it is difficult to comprehensively measure the model's ability in real interactive scenarios, cannot adapt to the diversity of different user needs, and lack effective measurements of the factuality, morality and security of generated content.

Method used

Using a dynamic evaluation framework, through sample screening, sample enhancement and negative sample creation, the original benchmark sample set is optimized and processed to generate more complex and diverse evaluation samples, including collaboration between sample screening agents, iterative decision-making agents and negative sample generation agents, and iterative optimization and verification is used for multi-agent system.

Benefits of technology

It provides a more detailed and comprehensive evaluation of the performance of large language models, improves the accuracy and robustness of the evaluation, and can better reflect the performance of the model in complex tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258144A_ABST
    Figure CN120258144A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a dynamic assessment method and device for a large language model, equipment and a medium. The method comprises the steps that an original reference sample set of the large language model to be assessed is acquired; performing sample dynamic optimization processing on the original reference sample set to obtain a dynamic reference sample set of the to-be-evaluated large language model, the sample dynamic optimization processing including sample screening, sample enhancement and negative sample creation; and based on at least one set large language model and the dynamic reference sample set, verifying and evaluating the to-be-evaluated large language model to obtain a verification and evaluation result of the to-be-evaluated large language model. Different from a traditional static method, the method is used for dynamically evaluating the to-be-evaluated large language model by adopting a dynamic evaluation framework to generate a more complex and novel evaluation sample, so that more detailed and comprehensive evaluation on the model performance is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a dynamic evaluation method, device, equipment and medium for large language models. Background Art

[0002] With the rapid development of deep learning and natural language processing technologies, large language models have demonstrated excellent capabilities in fields such as text generation, machine translation, code completion, and dialogue systems. However, due to the large scale and numerous parameters of these models, there are still many challenges in aspects such as the quality, accuracy, fairness, and security of the generated content. Therefore, how to effectively evaluate the performance of large language models has become an urgent problem for researchers and application developers.

[0003] Currently, the evaluation methods for large language models are mainly divided into two categories: static evaluation and dynamic evaluation. Static evaluation usually relies on pre-set benchmark test datasets and uses artificially defined metrics for scoring. Although static evaluation methods provide certain quantitative criteria, their limitation is that it is difficult to comprehensively measure the true capabilities of large language models. Due to relying on pre-set datasets and fixed metrics, static evaluation often lacks the investigation of the model's performance in real interaction scenarios and cannot adapt to the diversity of different user needs. In addition, such methods usually focus on short text matching or the accuracy of specific tasks, while ignoring the coherence, consistency, and adaptability of the model in long-term interactions. At the same time, static evaluation is difficult to effectively measure the factuality, morality, and security of the generated content by the model. Especially when it comes to misleading information, its performance is often unreliable. Therefore, simply relying on static evaluation is difficult to comprehensively reflect the comprehensive capabilities of large language models in real environments, and more dynamic and flexible evaluation means are urgently needed to supplement. To make up for the deficiencies of static evaluation methods, in recent years, researchers have proposed dynamic evaluation methods, emphasizing using dynamically transformed data to test the model's capabilities, or continuously monitoring and optimizing large language models in real application environments or interaction processes. The dynamic evaluation using dynamically transformed datasets has high controllability and repeatability, which can ensure the standardization of the evaluation process and adapt to the changes of different tasks and models. By continuously updating and expanding the datasets, the evaluation method can cover a wider range of language phenomena and diverse test scenarios, making the evaluation results more comprehensive and representative. In addition, this method allows researchers to conduct fine-grained analysis for specific capabilities, which helps to accurately locate the advantages and disadvantages of the model. However, existing dynamic testing methods usually require manual intervention, and the quality of the generated samples is poor. Summary of the Invention

[0004] An embodiment of the present invention provides a dynamic evaluation method, device, equipment and medium for large language models, which realizes generating more complex and diverse evaluation samples by sample screening, sample enhancement and negative sample creation, thereby providing a more detailed and comprehensive evaluation of the model performance.

[0005] In a first aspect, this embodiment provides a dynamic evaluation method for large language models, the method includes:

[0006] Obtain the original benchmark sample set of the large language model to be evaluated;

[0007] Perform sample dynamic optimization processing on the original benchmark sample set to obtain the dynamic benchmark sample set of the large language model to be evaluated, and the sample dynamic optimization processing includes sample screening, sample enhancement and negative sample creation;

[0008] Based on at least one set large language model and the dynamic benchmark sample set, perform verification evaluation on the large language model to be evaluated, and obtain the verification evaluation result of the large language model to be evaluated.

[0009] In a second aspect, this embodiment provides a dynamic evaluation device for large language models, the device includes:

[0010] An original benchmark acquisition module, configured to obtain the original benchmark sample set of the large language model to be evaluated;

[0011] An optimization processing module, configured to perform sample dynamic optimization processing on the original benchmark sample set to obtain the dynamic benchmark sample set of the large language model to be evaluated, and the sample dynamic optimization processing includes sample screening, sample enhancement and negative sample creation;

[0012] A model evaluation module, configured to perform verification evaluation on the large language model to be evaluated based on at least one set large language model and the dynamic benchmark sample set, and obtain the verification evaluation result of the large language model to be evaluated.

[0013] In a third aspect, this embodiment provides an electronic device, including:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the dynamic evaluation method for large language models according to any embodiment of the present invention.

[0017] Fourthly, this embodiment provides a computer-readable storage medium, and the computer program is executed by the at least one processor, so that the at least one processor can execute the dynamic evaluation method of the large language model according to any embodiment of the present invention.

[0018] The embodiment of the present invention provides a dynamic evaluation method, device, equipment and medium for a large language model. The method includes: first, obtaining the original benchmark sample set of the large language model to be evaluated; secondly, performing sample dynamic optimization processing on the original benchmark sample set to obtain the dynamic benchmark sample set of the large language model to be evaluated, and the sample dynamic optimization processing includes sample screening, sample enhancement and negative sample creation; finally, based on at least one set large language model and the dynamic benchmark sample set, verifying and evaluating the large language model to be evaluated to obtain the verification and evaluation result of the large language model to be evaluated. Different from the traditional static method, the above technical solution adopts a dynamic evaluation framework for dynamically evaluating the large language model to be evaluated, performs sample dynamic optimization processing on the original benchmark sample set, and the optimal samples are screened out in each stage. The ability of the model to be evaluated is evaluated by a dynamic and high-order model. The iterative dynamic evaluation framework uses sample screening, sample enhancement and negative sample creation to generate more complex and diverse evaluation samples, thereby providing a more detailed and comprehensive evaluation of the model performance and having good evaluation performance in complex tasks.

[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0021] Figure 1 It is a schematic flowchart of a dynamic evaluation method for a large language model provided in Embodiment 1 of the present invention;

[0022] Figure 2 It is a schematic flowchart of another dynamic evaluation method for a large language model provided in Embodiment 2 of the present invention;

[0023] Figure 3 It is a structural example diagram of a dynamic evaluation framework for a large language model in a certain application scenario provided in Embodiment 2 of the present invention;

[0024] Figure 4Schematic structural diagram of a dynamic evaluation device for a large language model provided in Embodiment 3 of the present invention;

[0025] Figure 5 Schematic structural diagram of an electronic device provided in Embodiment 4 of the present invention. Detailed implementation manners

[0026] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0028] It should be noted that the existing evaluation methods for large language models have the following problems: 1) Static testing: Static evaluation usually relies on a preset benchmark test dataset and uses artificially defined metrics for scoring. This method often lacks the investigation of the model's performance in real interaction scenarios and cannot adapt to the diversity of different user needs. 2) Data leakage: The large language model may have been exposed to the benchmark dataset during training, which may affect the effectiveness of the evaluation. 3) Manual intervention and lack of powerful iterative optimization: The existing dynamic testing methods usually require manual intervention, and at the same time, the sample transformation means are relatively single, resulting in poor quality of the generated samples.

[0029] Specifically, the evaluation of large language models currently involves multiple dimensions, covering aspects such as generation quality, reasoning ability, knowledge, and security. Mainstream large language model evaluation methods mainly rely on benchmark tests and fixed datasets. These benchmark tests provide standardized metrics for comparing model performance. For example, benchmark tests such as SuperGLUE, C-EVAl, GSM8K, and MMLU are commonly used to evaluate the performance of large language models in typical natural language understanding tasks. To supplement these general benchmark tests, domain-specific evaluation benchmarks have been developed, covering fields such as law, medicine, education, and finance. However, static benchmark tests have some limitations, especially the problem of data contamination, which gives models an unfair advantage during the evaluation process, leading to overly optimistic results and unable to accurately reflect their performance in novel or open-ended tasks.

[0030] To address the limitations of static evaluation, several dynamic evaluation strategies have been proposed, but automatically generating high-quality evaluation data and avoiding error samples remains an important challenge.

[0031] Embodiment 1

[0032] Figure 1 FIG. is a schematic flowchart of a dynamic evaluation method for a large language model provided in Embodiment 1 of the present invention. This method is applicable to the situation of dynamically evaluating a large language model. This method can be executed by a dynamic evaluation device for a large language model. The dynamic evaluation device for a large language model can be implemented in the form of hardware and / or software and is generally integrated in an electronic device.

[0033] As Figure 1 shown, a dynamic evaluation method for a large language model provided in Embodiment 1 specifically may include the following steps:

[0034] S101. Obtain the original benchmark sample set of the large language model to be evaluated.

[0035] It should be clear that the original benchmark refers to the standard test set or evaluation criteria used to evaluate the performance of a model, algorithm, or system in a certain field or task. These benchmarks are initially defined to measure the performance of a technology under given conditions and are usually used in the early stages of technology or research to establish a unified measurement standard. In this embodiment, the large language model whose capabilities are to be evaluated is denoted as the large language model to be evaluated. For example, the code ability, natural language processing ability, image processing ability, etc. of the large language model can be evaluated. The sample set composed of the original benchmarks corresponding to the large language model to be evaluated is denoted as the original benchmark sample set. In this embodiment, the original benchmark sample set can adopt the commonly used evaluation datasets of large language models, such as question answering, reasoning ability tests, code generation, etc.

[0036] S102. Perform sample dynamic optimization processing on the original benchmark sample set to obtain the dynamic benchmark sample set of the large language model to be evaluated.

[0037] Among them, the sample dynamic optimization processing includes sample screening, sample enhancement, and negative sample creation.

[0038] Considering the huge cost of building an evaluation benchmark from scratch, in this embodiment, an iterative dynamic evaluation framework is adopted, aiming to generate diverse evaluation benchmarks by leveraging the existing original benchmark sample set. This framework iteratively optimizes and validates samples through multi-agent collaboration to generate more complex evaluation instances.

[0039] In this embodiment, in order to obtain high-quality and diverse samples, the samples in the original benchmark sample set are sequentially subjected to sample screening, sample enhancement, and negative sample creation. The final sample set obtained is denoted as the dynamic benchmark sample set. Subsequently, the capabilities of the large language model to be evaluated are verified and evaluated based on the dynamic benchmark sample set. The iterative dynamic evaluation framework operates through a structured multi-stage process. In the initial stage, the sample screening agent selects appropriate seed samples from the original benchmark. Next, the iterative decision-making agent collaborates with the tool invocation agent to apply various enhancement techniques to these seed samples to generate a set of candidate variants. From this candidate pool, the best candidate samples are selected and further optimized through iterative optimization. Then, the sample screening agent evaluates the quality of the newly generated samples, while the negative sample generation agent creates corresponding negative samples.

[0040] Among them, the process of sample screening for the samples in the original benchmark sample set can be described as follows: There may be some impurities or bad question-and-answer pairs in the samples in the original benchmark sample set. Through this step of processing, high-quality samples can be selected and low-quality samples can be excluded. In this embodiment, a high-order large language model can be used as the sample screening agent to screen out the seed sample set within the basic capabilities range of the large language model to be evaluated from each sample in the original benchmark sample set.

[0041] Continuing with the above description, after screening out the seed sample set from the original benchmark sample set, each sample in the seed sample set can be enhanced. In this embodiment, an enhancement tool set is pre-constructed. The enhancement tool set can be specifically understood as a set containing various tools or resources, which can be used to complete specific tasks, such as tools, methods, or technologies. These tools and resources can help improve, enhance, or optimize the execution ability of a certain system, process, or task, making it more effective, efficient, and perfect. These tools can be deployed as independent modules. In each round of iteration, the large language model selects an appropriate transformation strategy according to the characteristics of the sample to generate multiple candidate samples, and evaluates the sample effects from multiple dimensions, selecting the sample with the best performance as the seed sample for the next round of iteration, or if the iteration limit is reached, as the final output, denoted as the best candidate sample set.

[0042] In this embodiment, considering that in tasks based on judgment, the large language model often aligns with human expectations, which may introduce biases in evaluating the accuracy of samples. To solve this problem, in this embodiment, a negative sample generation agent can be set up to execute this step. Negative samples are created for the best candidate sample set. Although the generated negative samples are similar to the correct answers, they deliberately contain logical or factual errors. These samples are used to evaluate the ability of the verification agent to distinguish between correct and wrong answers, thereby improving the precision and robustness of the verification process. It can be understood that the samples included in the best candidate sample set are positive samples. Combining with the created negative samples, they jointly form a sample set containing positive and negative samples, serving as the dynamic benchmark sample set for evaluating the large language model to be evaluated.

[0043] S103. Based on at least one set large language model and the dynamic benchmark sample set, verify and evaluate the large language model to be evaluated, and obtain the verification and evaluation result of the large language model to be evaluated.

[0044] In this embodiment, one large language model or multiple large language models are used to verify and evaluate the large language model to be evaluated based on a dynamic benchmark sample set, and the verification result obtained is recorded as the verification and evaluation result of the large language model to be evaluated. The model for verifying and evaluating the large language model to be evaluated based on the dynamic benchmark sample set can be regarded as a verification agent, and the verification agent decides whether to use multi-model or single-model verification according to the complexity of the samples. For samples with a low or medium difficulty type, the positive and negative samples in the dynamic benchmark sample set are identified based on the large language model to be evaluated, and the identification situation of the large language model to be evaluated is evaluated and verified by a set higher-order large language model, so as to obtain the verification and evaluation result of the large language model to be evaluated. For complex tasks, that is, samples with a high difficulty type, the positive and negative samples in the dynamic benchmark sample set are identified based on the large language model to be evaluated, and the identification situation of the large language model to be evaluated is evaluated and verified by two or more set large language models. A voting mechanism can be used to determine the evaluation and verification results of these set large language models, so as to obtain the verification and evaluation result of the large language model to be evaluated. Exemplarily, the positive and negative samples in the dynamic benchmark sample set can be identified based on the large language model to be evaluated, and the verification agent determines the value of the verification evaluation index based on the identification result, and further determines the verification evaluation result. In this embodiment, no specific limitation is imposed on the verification evaluation index. For example, it can be accuracy, recall rate, etc.

[0045] The above technical solution adopts a dynamic evaluation framework for dynamically evaluating the large language model to be evaluated, performs sample dynamic optimization processing on the original benchmark sample set, selects the optimal samples at each stage, evaluates the capabilities of the model to be evaluated through a dynamic higher-order model, and then generates more complex and diverse evaluation samples through sample enhancement and negative sample creation, thereby providing a more detailed and comprehensive evaluation of the model performance.

[0046] Embodiment 2

[0047] Figure 2 FIG. is a schematic flowchart of another dynamic evaluation method for a large language model provided in Embodiment 2 of the present invention. This embodiment is a further optimization of the above embodiment. In this embodiment, further limitations and optimizations are made on "performing sample dynamic optimization processing on the original benchmark sample set to obtain the dynamic benchmark sample set of the large language model to be evaluated", and on "verifying and evaluating the large language model to be evaluated based on at least one set large language model and the dynamic benchmark sample set to obtain the verification and evaluation result of the large language model to be evaluated".

[0048] As Figure 2 shown, Embodiment 2 of the present invention provides a dynamic evaluation method for a large language model, which specifically includes the following steps:

[0049] S201. Obtain the original benchmark sample set of the large language model to be evaluated.

[0050] S202. Perform sample screening on the original benchmark sample set to screen out the seed sample set from the original benchmark sample set.

[0051] In this embodiment, it can be considered that this step is executed by deploying a sample screening agent. The sample screening agent can adopt a high-order large language model. Specifically, by utilizing the powerful reasoning and understanding capabilities of the large language model, appropriate seed data is selected from the original benchmark sample set. The seed data screened out from the original benchmark sample set is denoted as the seed sample set. Taking the samples in the original benchmark sample set as inputs, each sample is evaluated, and only those samples within the basic capabilities of the large language model are selected as seed samples, and the screened seed samples form the seed sample set.

[0052] As a specific implementation method, the step of performing sample screening on the original benchmark sample set to screen out the seed sample set can be optimized, including:

[0053] Input each sample in the original benchmark sample set into the set first large language model to screen out the seed sample set within the basic capabilities of the large language model to be evaluated.

[0054] It should be noted that there may be some impurities or bad Q&A pairs in the samples in the original benchmark sample set. Through this step of processing, high-quality samples can be selected and low-quality samples can be excluded. In this embodiment, a high-order large language model can be used as the sample screening agent to screen out the seed sample set from the original benchmark sample set. This large language model is denoted as the first large language model. The first large language model can be specifically understood as being used to screen out the seed sample set within the basic capabilities of the large language model to be evaluated from each sample in the original benchmark sample set.

[0055] Exemplarily, the sample screening prompt can be set as: Screen out high-quality samples from the original benchmark sample set and define what high-quality samples are like. Input each sample in the original benchmark sample set and the sample screening prompt into the first large language model, so as to screen out the seed sample set within the basic capabilities of the large language model to be evaluated. A suitable high-order large language model can be selected as the basic model for sample screening. For the design of the sample screening prompt, two demonstrations can be used to improve the judgment accuracy.

[0056] The above technical solution concretizes the step of screening out the seed sample set from the original benchmark sample set, which can reduce unnecessary complexity and potential errors in the subsequent enhancement stage.

[0057] S203. Perform sample augmentation on the seed sample set to obtain the best candidate sample set.

[0058] In this embodiment, first, the large language model selects a sample transformation strategy according to the sample characteristics to transform the samples. Based on this, the samples are gradually augmented through multiple rounds of tool calls, integrating multiple augmentation strategies to ensure that the generated data has high diversity and complexity. However, it is difficult to accurately evaluate the augmentation effect by generating a single modified sample because of the lack of a comparison reference. To solve this problem, the iterative decision agent proposes a comparison method. In each round of iteration, two different augmentation tools, that is, sample transformation strategies, are selected by the large language model to generate candidate samples. Finally, for the sample set, the operation in each iteration is that the large language model selects two appropriate tools according to the characteristics of the samples, and the evaluator uses the large language model to guide the agent to select the best-performing sample as the seed sample for the next round of iteration, or as the final output if the iteration limit is reached, which is denoted as the best candidate sample set.

[0059] As a specific implementation, the steps of performing sample augmentation on the seed sample set to obtain the best candidate sample set can be optimized, including:

[0060] a1) For each sample in the seed sample set, based on the set second large language model, select a sample transformation strategy according to the characteristics of the sample, perform augmentation transformation on the sample, and obtain at least two candidate samples.

[0061] In this embodiment, the second large language model is used to select a corresponding sample transformation strategy according to the characteristics of the sample to achieve a better augmentation transformation effect. For the seed sample set, the operation of the second large language model in each iteration is to select two appropriate augmentation tools according to the characteristics of the samples, and gradually enhance the quality and complexity of the samples through multiple rounds of tool calls. The second large language model can be regarded as a tool call agent.

[0062] In this embodiment, an augmentation tool set is pre-constructed. The augmentation tool set can be specifically understood as a set containing various tools or resources, which can be used to complete specific tasks, such as tools, methods, or technologies. These tools and resources can help improve, enhance, or optimize the execution ability of a certain system, process, or task, making it more effective, efficient, and perfect. These augmentation tools can include external software, such as translation tools, or rule scripts for synonym replacement. Custom prompt functions are also used to make more complex modifications using the large language model. These tools can be deployed as independent modules. This step can use a high-order large language model to call the augmentation tools in the augmentation tool set.

[0063] Continuing with the above description, the large language model for invoking enhancement tools can serve as a tool invocation proxy to execute data enhancement strategies by dynamically invoking various enhancement tools. Currently, all sample enhancement tasks are invoked by the second large language model, and this set of enhancement tools serves the second large language model. The second large language model invokes enhancement tools from a pre-built set of enhancement tools to enhance samples respectively, obtaining at least two candidate samples. It can be understood that in this step, the large language model selects appropriate sample transformation strategies according to its characteristics, so as to realize the conversion from old samples to new samples.

[0064] Specifically, for each sample in the seed sample set, based on the second large language model according to the characteristics of the sample, two or more enhancement tools are selected from the pre-built set of enhancement tools, and the selected enhancement tools are invoked to enhance the sample respectively, so that two or more candidate samples will be obtained for each sample in the seed sample set. Exemplarily, assuming that the set of enhancement tools includes enhancement tool 1, enhancement tool 2, enhancement tool 3, enhancement tool 4, enhancement tool 5, and enhancement tool 6, etc., and the seed sample set includes sample 1, sample 2, and sample 3, etc., then enhancement tool 1 and enhancement tool 2 can be used to enhance sample 1 respectively to obtain candidate sample 1 and candidate sample 2, enhancement tool 3 and enhancement tool 4 can be used to enhance sample 2 respectively to obtain candidate sample 3 and candidate sample 4, and enhancement tool 5 and enhancement tool 6 can be used to enhance sample 3 respectively to obtain candidate sample 5 and candidate sample 6, etc. The above-mentioned enhancement tools are the sample transformation strategies.

[0065] b1) The candidate samples corresponding to each sample form a candidate sample set.

[0066] Specifically, the candidate samples corresponding to each sample are jointly formed into a sample set, denoted as the candidate sample set. Continuing with the above example for description, candidate sample 1, candidate sample 2, candidate sample 3, candidate sample 4, candidate sample 5, and candidate sample 6, etc. are formed into the candidate sample set.

[0067] c1) Based on the set third large language model, quality assessment is performed on the samples in the candidate sample set, and the best candidate sample set is screened out.

[0068] In this embodiment, it is necessary to deploy an iterative decision-making agent as the core component. The iterative decision-making agent is implemented using a set high-order large language model, which is denoted as the third large language model. The third large language model can be specifically understood as a model for evaluating the quality of samples. When evaluating the quality of samples, it can be evaluated from multiple dimensions such as correctness and complexity. Considering that it is difficult to accurately evaluate the enhancement effect by generating a single modified sample because of the lack of a comparison reference. To solve this problem, a comparison method is proposed. In each round of iteration, the set second large language model selects two different enhancement tools to generate candidate samples. For the seed sample set, the operation of the second large language model in each iteration is to select two tools according to the characteristics of the seed samples. The evaluator uses the third large language model to guide the evaluation process to select the best-performing sample as the seed sample for the next round of iteration, or if the iteration limit is reached, as the final output, denoted as the best candidate sample set. By using the third large language model to evaluate the quality of the samples in the candidate sample set, samples with better quality can be screened out to form the best candidate sample set.

[0069] The above technical solution concretizes the steps of sample enhancement for the seed sample set. It adopts an iterative enhancement and comparison selection mechanism. Through multiple rounds of data enhancement, this framework simulates the evolutionary process. By gradually increasing the sample complexity, and the competitive selection ensures that only the optimal samples are retained. This iterative optimization and screening produce an evaluation sample set with increasingly high quality.

[0070] S204. Create negative samples for the best candidate sample set to generate a dynamic benchmark sample set for the large language model to be evaluated.

[0071] Considering that in judgment-based tasks, large language models often align with human expectations, which may introduce biases when evaluating the accuracy of sample. To solve this problem, in this embodiment, a negative sample generation agent can be set up to execute this step. When creating negative samples for the best candidate sample set, the generated negative samples, although similar to the correct answers, deliberately contain logical or factual errors. These samples are used to evaluate the ability of the verification agent to distinguish between correct and wrong answers, thereby improving the precision and robustness of the verification process. It can be understood that the samples included in the best candidate sample set are positive samples. Combining with the created negative samples, they jointly form a sample set containing positive and negative samples, as the dynamic benchmark sample set for the large language model to be evaluated.

[0072] As a specific implementation, the step of creating negative samples for the best candidate sample set to generate a dynamic benchmark sample set for the large language model to be evaluated can be optimized, including:

[0073] a2) Using the set fourth language model, create negative samples for each sample in the best candidate sample set to generate erroneous negative samples corresponding to the samples.

[0074] In this embodiment, a set large language model is used to create negative samples for each sample in the best candidate sample set. The large language model used in this step is a model that is good at creating negative samples, and is recorded as the fourth largest language model. It should be noted that although the negative sample is similar to the positive sample answer, it intentionally contains logical or factual errors.

[0075] b2) The best candidate sample set and each negative sample constitute a dynamic benchmark sample set of the large language model to be evaluated.

[0076] In this embodiment, the samples included in the best candidate sample set are positive samples. The positive samples and the created negative samples constitute a sample set, which is recorded as a dynamic benchmark sample set and is used for subsequent verification and evaluation of the large language model to be evaluated.

[0077] The above technical solution specifies the steps of creating negative samples for the best candidate sample set. Taking into account that in judgment-based tasks, large language models tend to align with human expectations, which may introduce biases when evaluating sample accuracy, positive and negative samples are used to evaluate the verification agent's ability to distinguish between correct and incorrect answers, thereby improving the accuracy and robustness of the verification process.

[0078] S205. If the difficulty type of the samples in the dynamic benchmark sample set is medium or low, the set fifth largest language model and the dynamic benchmark sample set are used to evaluate the large language model to be evaluated, and a verification evaluation result of the large language model to be evaluated is obtained.

[0079] In this embodiment, step S205 and step S206 can be considered as steps performed by the verification agent, which is responsible for verifying the validity of the generated data. It evaluates positive samples and negative samples and verifies their correctness by comparing the results. The judgment of the verification agent is considered to be reliable. When the positive sample is correctly identified as the accurate answer and the negative sample is identified as the wrong answer, the recognition is considered to be correct. The verification agent can be implemented using a high-order large language model. The set fifth language model can be considered as an evaluation model used to evaluate the ability of the large language model to be evaluated. This process is equivalent to identifying the positive samples and negative samples in the dynamic benchmark sample set based on the large language model to be evaluated, and the set fifth language model is used to evaluate and verify the recognition of the large language model to be evaluated, thereby obtaining the verification evaluation result of the large language model to be evaluated. In this embodiment, a suitable high-order large language model is selected to verify samples with medium or low difficulty types.

[0080] S206. If the difficulty type of the samples in the dynamic benchmark sample set is high, at least two set large language models and the dynamic benchmark sample set are used to evaluate the large language model to be evaluated, and the verification evaluation result of the large language model to be evaluated is obtained.

[0081] In this embodiment, complex and highly difficult samples are processed through a multi-model voting mechanism. For complex tasks, the verification agent adopts a voting mechanism, involving multiple large language models to jointly evaluate samples with a high difficulty type, so as to reduce the bias that may be introduced by a single model. When evaluating the large language model to be evaluated, the large language model to be evaluated is evaluated based on two or more set large language models. These two or more set large language models can be regarded as evaluation models, and the evaluation models are used to evaluate the capabilities of the large language model. This process is equivalent to identifying positive and negative samples in the dynamic benchmark sample set based on the large language model to be evaluated, and the evaluation and verification of the identification situation of the large language model to be evaluated are carried out by two or more set large language models. A voting mechanism can be used to determine the evaluation and verification results of these set large language models, so as to obtain the verification evaluation result of the large language model to be evaluated.

[0082] It should be noted that in this embodiment, a dynamic and high-order large language model is used as an agent model to verify and evaluate the capabilities of the large language model to be evaluated. Different large language models are used to evaluate the capabilities of the model to be evaluated at each stage, and the optimal samples are selected through the processing of the large language models at each stage. For example, in the sample screening stage, a suitable large language model is selected as the sample screening agent model through model switching at this stage to evaluate the samples in the front-end original benchmark sample set, and the optimal samples at this stage are obtained. Similarly, in the sample enhancement stage, a suitable large language model is selected as the sample enhancement agent model through model switching at this stage to further enhance the screened samples, and the optimal samples at this stage are obtained. At the same time, in the selection stage of the transformation method, the characteristics of the input samples are also analyzed by the large language model, and then a suitable method is selected for processing. Similarly, in the negative sample generation stage, a suitable large language model is also selected as the negative sample generation agent model through model switching at this stage to generate negative samples for the enhanced samples, and the optimal samples at this stage are obtained, and finally the dynamic benchmark sample set is obtained. In addition, in the verification stage, one or more suitable large language models are also selected as the verification agent models through model switching at this stage to realize the verification and evaluation of the large language model to be evaluated based on the dynamic benchmark sample set, and the verification evaluation result of the large language model to be evaluated is obtained.

[0083] Continuing with the above description, the large language models used in each stage can be regarded as evaluation models, which are used to evaluate the large language model to be evaluated. In addition, the large language models used in each stage are not fixed. Appropriate large language models can be selected according to the actual situation and will be replaced with appropriate large language models as the actual situation changes.

[0084] The above technical solution concretizes the steps of optimizing the original benchmark sample set to obtain a dynamic benchmark sample set and the steps of verifying and evaluating the large language model to be evaluated, realizing a dynamic evaluation method for multi-agent collaboration, dynamically generating complex and diverse evaluation samples. This method enables the benchmark test to be flexibly updated and better reflects the evolution of the model's capabilities. By introducing various enhancement methods, samples are processed iteratively to create increasingly challenging evaluation scenarios. In addition, effective decision-making and verification methods are realized. The iterative dynamic evaluation framework adopts a comparative evaluation strategy, selects the best samples according to multiple criteria, and ensures the quality and reliability of the evaluation through the collaborative verification process of multiple large language models.

[0085] As an optional embodiment of the embodiment of the present invention, on the basis of the above embodiment, the method can be further optimized to include: based on the first large language model, perform problem complexity analysis on each sample in the seed sample set and the best candidate sample set, and identify the samples as samples of three difficulty types: high, medium, and low.

[0086] In this embodiment, this step can be executed through sample screening. Based on the first large language model, the difficulty of each newly generated sample is evaluated by analyzing the complexity of the problem and the required depth of knowledge, and the samples are identified as samples of three difficulty types: high, medium, and low.

[0087] The above technical solution adds the function of classifying the difficulty level of samples, and this classification helps to guide the selection of appropriate processing strategies in the subsequent evaluation stage.

[0088] Exemplarily, in order to more clearly describe the dynamic evaluation method of the large language model provided by the embodiment of the present invention, taking the actual application scenario of the dynamic evaluation of a certain large language model as an example, the structure of the dynamic evaluation framework for executing the large language model is described. Exemplarily, Figure 3 FIG. is a structural example diagram of the dynamic evaluation framework of the large language model in an application scenario provided by the second embodiment of the present invention, as Figure 3As shown, the dynamic evaluation framework of the large language model can specifically include: The original benchmark sample set (shown as the original benchmark in the figure) is input into the sample screening agent (i.e., the first large language model). The sample screening agent screens the samples in the original benchmark sample set to obtain the seed sample set. At the same time, the sample screening agent can calibrate the difficulty levels of the samples in the seed sample set, respectively calibrated as simple samples, medium samples, or difficult samples; The tool invocation agent (i.e., the second large language model) selects two appropriate enhancement tools according to the sample characteristics, and invokes the enhancement tools in the enhancement tool set to perform enhancement processing on the samples in the seed sample set to obtain the candidate sample set; The iterative decision-making agent (i.e., the third large language model) performs quality evaluation on the samples in the candidate sample set, selects the best candidate samples to form the best candidate sample set. Similarly, the generated best candidate sample set will also be calibrated by the sample screening agent for the problem complexity of the samples, respectively calibrated as samples with low, medium, or high difficulty types; The best candidate sample set obtained through enhancement processing is input into the negative sample generation agent (i.e., the fourth large language model). The negative sample generation agent generates corresponding negative samples for the samples in the best candidate sample set, and takes the sample set containing positive samples and negative samples as the final dynamic benchmark sample set; The verification agent (i.e., the fifth large language model or multiple large language models) performs verification evaluation on the large language model to be evaluated based on the dynamic benchmark sample set, and obtains the verification evaluation result of the large language model to be evaluated; Dynamic benchmark: A benchmark standard that is dynamically adjusted and updated according to real-time data or changing environmental conditions. This benchmark standard is different from the traditional static benchmark, and it will change with time, conditions, or the system, and is used to measure and compare the performance of systems, processes, or performances. Compared with traditional methods, the iterative dynamic evaluation framework more effectively reveals the limitations of the model in complex tasks, providing a more comprehensive basis for the evaluation, optimization, and application of large language models.

[0089] Embodiment 3

[0090] Figure 4 As shown in the figure, it is a schematic structural diagram of a dynamic evaluation device for a large language model provided by Embodiment 3 of the present invention. This device is applicable to the situation of dynamically evaluating a large language model. The dynamic evaluation device for this large language model can be implemented in the form of hardware and / or software, and is generally integrated in an electronic device. As Figure 4 shown, the device includes: an original benchmark acquisition module 31, an optimization processing module 32, and a verification evaluation module 33, where

[0091] The original benchmark acquisition module 31 is used to acquire the original benchmark sample set of the large language model to be evaluated;

[0092] The optimization processing module 32 is used to perform sample dynamic optimization processing on the original benchmark sample set to obtain a dynamic benchmark sample set for the large language model to be evaluated. The sample dynamic optimization processing includes sample screening, sample enhancement, and negative sample creation;

[0093] The verification and evaluation module 33 is used to perform verification and evaluation on the large language model to be evaluated based on at least one set large language model and the dynamic benchmark sample set, and obtain the verification and evaluation result of the large language model to be evaluated.

[0094] The above technical solution adopts a dynamic evaluation framework for dynamically evaluating the large language model to be evaluated, and performs sample dynamic optimization processing on the original benchmark sample set. The optimal samples are screened out in each stage, and the ability of the model to be evaluated is evaluated by a dynamic and high-order model. The iterative dynamic evaluation framework uses sample screening, sample enhancement, and negative sample creation to generate more complex and diverse evaluation samples, thereby providing a more detailed and comprehensive evaluation of the model performance.

[0095] Optionally, the optimization processing module 32 includes:

[0096] The sample screening unit is used to screen the samples in the original benchmark sample set and screen out the seed sample set from the original benchmark sample set;

[0097] The sample enhancement unit is used to enhance the samples in the seed sample set to obtain the best candidate sample set;

[0098] The negative sample creation unit is used to create negative samples for the best candidate sample set to generate a dynamic benchmark sample set for the large language model to be evaluated.

[0099] Optionally, the sample screening unit is specifically used for:

[0100] Input each sample in the original benchmark sample set into the set first large language model, and screen out the seed sample set within the basic ability range of the large language model to be evaluated.

[0101] Optionally, the sample enhancement unit is specifically used for:

[0102] For each sample in the seed sample set, based on the set second large language model, select a sample transformation strategy according to the characteristics of the sample, perform enhancement transformation on the sample, and obtain at least two candidate samples;

[0103] The candidate samples corresponding to each sample form a candidate sample set;

[0104] Based on the set third large language model, perform quality evaluation on each sample in the candidate sample set, and screen out the best candidate sample set.

[0105] Optionally, the negative sample creation unit is specifically used for:

[0106] Using the set fourth large language model, negative samples are created for each sample in the optimal candidate sample set to generate incorrect negative samples corresponding to the samples.

[0107] The optimal candidate sample set and each negative sample are used to form a dynamic benchmark sample set for the large language model to be evaluated.

[0108] Optionally, the verification and evaluation module 33 is specifically configured to:

[0109] If the difficulty type of the samples in the dynamic benchmark sample set is medium or low, the set fifth large language model and the dynamic benchmark sample set are used to evaluate the large language model to be evaluated, and the verification and evaluation result of the large language model to be evaluated is obtained.

[0110] If the difficulty type of the samples in the dynamic benchmark sample set is high, at least two set large language models and the dynamic benchmark sample set are used to evaluate the large language model to be evaluated, and the verification and evaluation result of the large language model to be evaluated is obtained.

[0111] Optionally, the device further includes a sample level division module for:

[0112] Based on the first large language model, problem complexity analysis is performed on each sample in the seed sample set and the optimal candidate sample set, and the samples are identified as samples of three difficulty types: high, medium, and low.

[0113] The dynamic evaluation device of the large language model provided by the embodiments of the present invention can execute the dynamic evaluation method of the large language model provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0114] Embodiment 4

[0115] Figure 5 It is a schematic structural diagram of an electronic device provided by Embodiment 4 of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0116] As Figure 5As shown, the electronic device 40 includes at least one processor 41 and a memory communicatively connected to the at least one processor 41, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc. The memory stores a computer program executable by the at least one processor. The processor 41 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 42 or the computer program loaded from the storage unit 48 into the random access memory (RAM) 43. In the RAM 43, various programs and data required for the operation of the electronic device 40 can also be stored. The processor 41, the ROM 42, and the RAM 43 are connected to each other via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.

[0117] Multiple components in the electronic device 40 are connected to the I / O interface 45, including: an input unit 46, such as a keyboard, a mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a magnetic disk, an optical disc, etc.; and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the electronic device 40 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0118] The processor 41 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 41 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 41 executes the various methods and processes described above, such as the dynamic evaluation method of a large language model.

[0119] In some embodiments, the dynamic evaluation method of a large language model can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 48. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 40 via the ROM 42 and / or the communication unit 49. When the computer program is loaded into the RAM 43 and executed by the processor 41, one or more steps of the dynamic evaluation method of the large language model described above can be executed. Alternatively, in other embodiments, the processor 41 can be configured to execute the dynamic evaluation method of the large language model by any other appropriate means (e.g., by means of firmware).

[0120] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0121] The computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0122] In the context of the present invention, a computer-readable storage medium may be a tangible medium that can contain, or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0123] To provide interaction with a user, the systems and techniques described herein can be implemented on a vehicle having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the vehicle. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0124] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0125] The computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0126] An embodiment of the present invention also provides a computer program product, including a computer program which, when executed by a processor, implements the dynamic evaluation method of the large language model provided in any embodiment of the present invention.

[0127] In the process of implementing the computer program product, computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0128] It should be understood that various forms of the processes shown above may be used, steps may be reordered, added, or deleted. For example, the steps recited in the present invention may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0129] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A dynamic evaluation method for large language models, characterized in that, including: obtaining an original benchmark sample set of the large language model to be evaluated; performing sample dynamic optimization processing on the original benchmark sample set to obtain a dynamic benchmark sample set of the large language model to be evaluated, where the sample dynamic optimization processing includes sample screening, sample enhancement, and negative sample creation; based on at least one set large language model and the dynamic benchmark sample set, performing verification evaluation on the large language model to be evaluated to obtain a verification evaluation result of the large language model to be evaluated.

2. The method according to claim 1, wherein The performing sample dynamic optimization processing on the original benchmark sample set to obtain a dynamic benchmark sample set of the large language model to be evaluated includes: performing sample screening on the original benchmark sample set to screen out a seed sample set from the original benchmark sample set; performing sample enhancement on the seed sample set to obtain an optimal candidate sample set; performing negative sample creation on the optimal candidate sample set to generate a dynamic benchmark sample set of the large language model to be evaluated.

3. The method according to claim 2, wherein The performing sample screening on the original benchmark sample set to screen out a seed sample set from the original benchmark sample set includes: inputting each sample in the original benchmark sample set into a set first large language model to screen out a seed sample set within the basic capabilities of the large language model to be evaluated.

4. The method according to claim 2, wherein The performing sample enhancement on the seed sample set to obtain an optimal candidate sample set includes: for each sample in the seed sample set, using a set second large language model to select a sample transformation strategy according to the characteristics of the sample, performing enhancement transformation on the sample to obtain at least two candidate samples; forming a candidate sample set with the candidate samples corresponding to each sample; based on a set third large language model, performing quality evaluation on each sample in the candidate sample set to screen out an optimal candidate sample set.

5. The method according to claim 2, wherein The performing negative sample creation on the optimal candidate sample set to generate a dynamic benchmark sample set of the large language model to be evaluated includes: using a set fourth large language model to perform negative sample creation on each sample in the optimal candidate sample set to generate an incorrect negative sample corresponding to the sample; forming the optimal candidate sample set and each of the negative samples into a dynamic benchmark sample set of the large language model to be evaluated.

6. The method according to claim 1, characterized in that The performing verification evaluation on the large language model to be evaluated based on at least one set large language model and the dynamic benchmark sample set to obtain a verification evaluation result of the large language model to be evaluated includes: if the difficulty type of the samples in the dynamic benchmark sample set is medium or low, using a set fifth large language model and the dynamic benchmark sample set to evaluate the large language model to be evaluated to obtain a verification evaluation result of the large language model to be evaluated; if the difficulty type of the samples in the dynamic benchmark sample set is high, using at least two set large language models and the dynamic benchmark sample set to evaluate the large language model to be evaluated to obtain a verification evaluation result of the large language model to be evaluated.

7. The method according to claim 3, characterized in that, also including: Based on the first large language model, perform problem complexity analysis on each sample in the seed sample set and the best candidate sample set, and identify the samples as samples of three difficulty types: high, medium, and low.

8. A dynamic evaluation device for a large language model, characterized in that, Including: An original benchmark acquisition module for acquiring an original benchmark sample set of the large language model to be evaluated; An optimization processing module for performing sample dynamic optimization processing on the original benchmark sample set to obtain a dynamic benchmark sample set of the large language model to be evaluated, where the sample dynamic optimization processing includes sample screening, sample enhancement, and negative sample creation; A verification and evaluation module for verifying and evaluating the large language model to be evaluated based on at least one set large language model and the dynamic benchmark sample set, and obtaining a verification and evaluation result of the large language model to be evaluated.

9. An electronic device, characterized in that, Including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the dynamic evaluation method of the large language model according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to implement the dynamic evaluation method of the large language model according to any one of claims 1-7 when executed.

Citation Information

Cited By

  • Explanatability fusion and recovery method after large language model training based on interpretability

    CN121936572A

  • Explainability-based large language model post-training capability fusion and recovery method

    CN121936572B