Assessment method, system and equipment of large language model and storage medium

By building a multi-dimensional evaluation index system and dynamic weight allocation, and using a third-party large language model for evaluation, the problems of low evaluation efficiency and lack of targetedness in the existing technology are solved, and the rapid, precise optimization and performance improvement of large language models are achieved.

CN120336464APending Publication Date: 2025-07-18XIAMEN YUANTING INFORMATION TECH CO LTD
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510378145.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The evaluation methods of existing large language models are inefficient and lack targeted, and cannot achieve a closed loop of evaluation and optimization, resulting in insufficient feedback during the model iteration process and affecting performance improvement.

Method used

Build a multi-dimensional evaluation index system, including contextual relevance, term consistency, language fluency and expression accuracy, dynamically allocate weights, and use third-party large language models to perform multi-dimensional scoring, generate structured feedback suggestions, and establish an evaluation-feedback optimization closed loop.

Benefits of technology

It realizes comprehensive, objective and rapid evaluation of large language models, provides accurate optimization guidance, improves the iteration efficiency and performance of the model, and supports multi-task and multi-field evaluation needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336464A_ABST
    Figure CN120336464A_ABST
Patent Text Reader

Abstract

The invention provides an evaluation method, system and device for a large language model and a storage medium, and the method comprises the steps: constructing a sample data set which comprises input data and corresponding reference answers, inputting the input data into an evaluated model, and generating an output result; a multi-dimensional evaluation index system is designed according to task requirements, a dynamic weight is allocated to each evaluation index, each evaluation index is provided with scoring standard description, and the evaluation indexes comprise at least two items of context correlation, term consistency, language fluency and expression accuracy; combining the input data, the reference answer, the output result and the scoring standard description into a standardized input instruction, and calling an evaluation model to perform multi-dimensional scoring on the standardized input instruction to generate an evaluation result; and analyzing the evaluation result according to a preset threshold value, and generating a structured feedback suggestion containing an improvement direction. According to the method, rapid optimization and iteration of the large language model can be effectively supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of large language model evaluation, and specifically relates to an evaluation method, system, device, and storage medium for large language models. Background Art

[0002] With the rapid development of deep learning and big data technologies, natural language processing technologies based on large language models have made remarkable progress and demonstrated powerful performance in many fields such as machine translation, text generation, and question answering systems. However, comprehensively and objectively evaluating the performance of these large models remains an urgent challenge to be solved.

[0003] The current evaluation methods mainly have the following deficiencies:

[0004] First, although manual evaluation is regarded as the gold standard, it is inefficient and difficult to provide timely feedback during the rapid iteration of models, seriously restricting the speed of model optimization.

[0005] Second, the automatic evaluation metrics are limited. Preset metrics such as BLEU and ROUGE cannot comprehensively reflect the performance of the model in different tasks and scenarios, and only output the total score, unable to accurately locate specific errors, resulting in a lack of pertinence in model optimization.

[0006] Third, the existing methods fail to establish an "evaluation - feedback optimization" closed - loop. The evaluation and optimization processes are disjointed, forming an iterative breakpoint, affecting the model's ability to continuously improve.

[0007] In view of this, the present invention proposes an evaluation method, system, device, and storage medium for large language models, which can automatically, objectively, and comprehensively evaluate the performance of large models. Summary of the Invention

[0008] To solve the problems of the existing evaluation methods lacking pertinence and failing to achieve an evaluation - feedback optimization closed - loop, etc., the present invention provides an evaluation method, system, device, and storage medium for large language models to solve the above - mentioned technical defect problems.

[0009] In a first aspect, the present invention proposes an evaluation method for large language models, and the method includes the following steps:

[0010] S1. Construct a sample data set, where the sample data set includes input data and its corresponding reference answers, and input the input data into the model to be evaluated to generate an output result;

[0011] S2. Design a multi - dimensional evaluation index system according to task requirements, and assign dynamic weights to each evaluation index. Each evaluation index is accompanied by a description of the scoring criteria. The evaluation indexes include at least two of context relevance, term consistency, language fluency, and expression accuracy;

[0012] S3. Combine the input data, reference answers, output results in step S1 and the description of the scoring criteria in step S2 into a standardized input instruction, and call the evaluation model to perform multi-dimensional scoring on the standardized input instruction to generate an evaluation result;

[0013] S4. Analyze the evaluation result according to a preset threshold to generate a structured feedback suggestion including improvement directions.

[0014] Preferably, in step S2, design a multi-dimensional evaluation index system according to the task requirements and assign dynamic weights to each evaluation index, including:

[0015] In the machine translation task, the weight of expression accuracy is higher than that of language fluency;

[0016] In the dialogue generation task, the weight of context relevance is higher than that of term consistency;

[0017] In the text summarization task, the weight of language fluency is higher than that of term consistency.

[0018] Preferably, in step S4, analyze the evaluation result according to a preset threshold to generate a structured feedback suggestion including improvement directions, including:

[0019] If the context relevance score is lower than the first threshold, it is recommended to increase the dialogue history context attention mechanism;

[0020] If the term consistency score is lower than the second threshold, it is recommended to inject a domain term dictionary into the training data or strengthen the alignment fine-tuning strategy;

[0021] If the language fluency score is lower than the third threshold, it is recommended to adjust the text generation strategy;

[0022] If the expression accuracy score is lower than the fourth threshold, it is recommended to increase a domain-specific dataset for retraining.

[0023] Preferably, in step S1, construct a sample dataset including obtaining sample data from public datasets and performing preprocessing operations such as cleaning, deduplication, and annotation on the sample data.

[0024] Preferably, in step S1, the sample dataset includes data in the fields of science and technology, medicine, finance, and entertainment, and the difficulty levels of the sample dataset include simple sentences, complex sentences, and ambiguous sentences.

[0025] Preferably, in step S3, it further includes:

[0026] Select a third-party large language model different from the architecture of the model to be evaluated and the training data as the evaluation model;

[0027] Adopt a method of combining instructions and data described in natural language, and combine the input data, reference answers, output results, and scoring criteria descriptions according to a preset format to obtain standardized input instructions.

[0028] Preferably, in step S3, call an evaluation model to perform multi-dimensional scoring on the standardized input instructions to generate an evaluation result, including:

[0029] Each standardized input instruction is evaluated multiple times. Each time during the evaluation, different evaluation model parameter settings are used. Calculate the preliminary average value for the multiple scores obtained, and after excluding the abnormal scores that deviate from the preliminary average value by more than a preset range, recalculate the final average value of the remaining scores, and use the final average value as the evaluation result.

[0030] In a second aspect, the present invention proposes an evaluation system for a large language model, which includes:

[0031] A data acquisition module configured to construct a sample data set. The sample data set includes input data and its corresponding reference answers, and input the input data into the model to be evaluated to generate an output result;

[0032] An evaluation index design module configured to design a multi-dimensional evaluation index system according to task requirements, and assign dynamic weights to each evaluation index. Each evaluation index is attached with a scoring criteria description. The evaluation indexes include at least two of context relevance, term consistency, language fluency, and expression accuracy;

[0033] An evaluation result generation module configured to combine the input data, reference answers, output results of the data acquisition module, and the scoring criteria description of the evaluation index design module into standardized input instructions, call an evaluation model to perform multi-dimensional scoring on the standardized input instructions, and generate an evaluation result;

[0034] A feedback and suggestion generation module configured to analyze the evaluation result according to a preset threshold and generate a structured feedback and suggestion including improvement directions.

[0035] In a third aspect, the present invention proposes a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the evaluation method for a large language model as described in any one of the above are implemented.

[0036] In a fourth aspect, the present invention proposes a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the evaluation method for a large language model as described in any one of the above are implemented.

[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0038] (1) Comprehensive evaluation: By constructing a sample dataset covering multiple fields and difficulty levels, and designing an evaluation index system covering multiple dimensions such as context relevance, term consistency, language fluency, and expression accuracy, the present invention can comprehensively and meticulously evaluate the performance of large language models in different tasks and scenarios, avoiding the one-sidedness of single-index evaluation, and providing a comprehensive guiding direction for model optimization.

[0039] (2) Dynamic weight allocation: Dynamically adjust the weights of each evaluation index according to different task requirements, making the evaluation more in line with the actual application scenario. For example, highlighting expression accuracy in the machine translation task, emphasizing context relevance in the dialogue generation task, and focusing on language fluency in the text summarization task. This flexible weight allocation mechanism improves the pertinence and effectiveness of the evaluation.

[0040] (3) Objectivity and consistency: Using a third-party large language model as the evaluation model and adopting standardized input instructions for multi-dimensional scoring reduce the subjectivity and inconsistency of manual evaluation. The evaluation model automatically evaluates according to the preset indicators and standards, ensuring the objectivity and stability of the evaluation results, and facilitating the performance comparison between different models or different versions of the same model.

[0041] (4) Efficiency and scalability: The present invention realizes the automation of the evaluation process, without the need for a large amount of manual participation, greatly improving the evaluation efficiency, being able to quickly process large-scale datasets, and meeting the needs of rapid model iteration. At the same time, this method has good scalability, supports the evaluation of multiple tasks and multiple fields, and is applicable to various natural language processing tasks such as machine translation, text summarization, and dialogue generation.

[0042] (5) Precise feedback and optimization guidance: By dynamically analyzing the evaluation results through preset thresholds, the weak links of the model can be accurately located, and structured feedback suggestions including specific improvement directions and optimization priorities can be generated. This detailed feedback information provides a clear direction and actionable suggestions for the subsequent optimization of the model, helping developers to improve the model targeted and enhance its performance.

[0043] (6) Model optimization closed-loop: The present invention establishes a "evaluation - feedback optimization" closed-loop mechanism, directly associating the optimization suggestions with the model training link, and realizing the dynamic iterative optimization of the model. The optimized model can be reconnected to the evaluation process until the score reaches the preset target, thereby promoting the continuous improvement of the model performance and accelerating the R & D and application process of large language models. Description of the Drawings

[0044] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, purposes, and advantages of the present application will become more obvious:

[0045] Figure 1 is a schematic flowchart of the evaluation method for the large language model according to the present invention;

[0046] Figure 2 is a flowchart for optimizing the large language model based on multi-round evaluation according to the present invention;

[0047] Figure 3 is a structural diagram of the evaluation system for the large language model according to the present invention;

[0048] Figure 4 is a schematic structural diagram of the computer system of the electronic device suitable for implementing the embodiments of the present invention. Detailed implementation manners

[0049] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than limiting the invention. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.

[0050] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and embodiments.

[0051] Figure 1 shows a schematic flowchart of the evaluation method for the large language model of the present invention. As Figure 1 shown, the present invention proposes an evaluation method for a large language model, and the method includes the following steps:

[0052] S1. Construct a sample data set, where the sample data set includes input data and its corresponding reference answers, and input the input data into the model to be evaluated to generate an output result.

[0053] In this embodiment, first, clarify the type of task to be evaluated, such as: machine translation, text summarization, dialogue generation, etc. Secondly, obtain a large amount of sample data covering the required tasks from public data sets or own data, and construct a sample data set, which includes input data, reference answers, etc. And perform preprocessing operations such as cleaning, deduplication, and annotation on the sample data to ensure data quality.

[0054] The sample data set covers different topics and fields, such as: data in the fields of technology, medicine, finance, and entertainment; the sample data set includes samples at different difficulty levels, such as: simple sentences, complex sentences, and ambiguous sentences. The sample data covers a diversity from simple to complex, and can comprehensively evaluate the model's capabilities.

[0055] Continue to refer to Figure 1, the evaluation method for the large language model provided by the present invention further includes the following steps:

[0056] S2. Design a multi-dimensional evaluation index system according to the task requirements, assign dynamic weights to each evaluation index, clearly define each evaluation index, and attach a description of the scoring criteria. The evaluation indices include at least two of context relevance, term consistency, language fluency, and expression accuracy. Write a detailed description of the scoring criteria for each index to facilitate the evaluation model's understanding of the evaluation requirements. The description of the scoring criteria includes a detailed explanation of the scoring levels, such as the specific performance corresponding to 1-5 points.

[0057] In this embodiment, according to the task requirements and importance, design a multi-dimensional evaluation index system and assign dynamic weights to each evaluation index, including:

[0058] In the machine translation task, the weight of expression accuracy is higher than the weight of language fluency;

[0059] In the dialogue generation task, the weight of context relevance is higher than the weight of term consistency;

[0060] In the text summarization task, the weight of language fluency is higher than the weight of term consistency.

[0061] Continue to refer to Figure 1 , the evaluation method for the large language model provided by the present invention further includes the following steps:

[0062] S3. Combine the input data, reference answer, output result of step S1, and the description of the scoring criteria in step S2 into a standardized input instruction, and call the evaluation model to perform multi-dimensional scoring on the standardized input instruction to generate an evaluation result.

[0063] In this embodiment, it further includes: selecting a third-party large language model different from the architecture and training data of the model to be evaluated as the evaluation model. The evaluation model should have strong language understanding and generation capabilities and be able to accurately understand the evaluation indices and task requirements. Preferably, use the latest open-source large language model as the evaluation model. It should be understood that what is emphasized here is that the data used by the model to be evaluated and the third-party large language model should be different during the pre-training or fine-tuning training stage. The core purpose is to ensure the fairness and objectivity of the evaluation and avoid biases and self-reinforcing effects caused by data similarity. Specifically, select a third-party large language model developed by a different company or team and having no direct association with the model to be evaluated as the evaluation tool to ensure the independence and credibility of the evaluation result. The sample data set mentioned above is an independent validation set, which is used to ensure the standardization and accuracy of the evaluation process and is different from the training data of the model to be evaluated.

[0064] Adopt a method of combining instructions described in natural language with data, and combine the input data, reference answers, output results, and scoring criteria descriptions according to a preset format to obtain standardized input instructions. Adopting a unified input format facilitates the evaluation of model processing. For example, organize the input into question-and-answer pairs or texts containing specific instructions. Provide clear evaluation instructions to the evaluation model, such as: "Please score the output of the model according to the following indicators and give reasons."

[0065] Preferably, in step S3, call the API or interface of the evaluation model to perform multi-dimensional scoring on the standardized input instructions to generate an evaluation result, where the evaluation result includes the scores of each indicator. Specifically, it includes the following steps:

[0066] Evaluate each standardized input instruction multiple times. Each time during the evaluation, use different evaluation model parameter settings. Calculate the preliminary average value of the multiple scores obtained, and after excluding the abnormal scores that deviate from the preliminary average value by more than the preset range, recalculate the final average value of the remaining scores, and use the final average value as the evaluation result. Taking the average value can improve the reliability of the evaluation.

[0067] S4. Analyze the evaluation result according to the preset threshold to generate a structured feedback suggestion including the improvement direction, including:

[0068] If the context relevance score is lower than the first threshold, it is recommended to add a dialogue history context attention mechanism;

[0069] If the term consistency score is lower than the second threshold, it is recommended to inject a domain term dictionary into the training data or strengthen the alignment fine-tuning strategy;

[0070] If the language fluency score is lower than the third threshold, it is recommended to adjust the text generation strategy;

[0071] If the expression accuracy score is lower than the fourth threshold, it is recommended to add a domain-specific dataset for retraining.

[0072] That is, for each evaluation dimension, determine which dimensions have scores lower than the set threshold. For the identified problems, generate specific optimization suggestions and consider different problem types.

[0073] Preferably, integrate the scoring and feedback information into the final output format to ensure clarity and readability. Collect and analyze the evaluation results output by the evaluation model, identify low-scoring items and potential problems. Use statistical methods (such as mean, standard deviation, etc.) to perform descriptive analysis on the results. Return the optimization suggestions to the user for their reference or direct application.

[0074] Example 1

[0075] The machine translation model is evaluated using the evaluation method of the present invention. The overall method process is divided into 5 steps: data preparation, evaluation index setting, selection of evaluation model, evaluation process, and result analysis. The detailed steps are as follows:

[0076] Step 1: Data Preparation

[0077] Determination of evaluation task: Evaluate the Chinese-to-English machine translation model.

[0078] Dataset construction: (1) Data collection: Select 1000 Chinese sentences and their corresponding English reference translations from publicly available machine translation datasets (such as WMT, UN parallel corpus). (2) Data preprocessing: Clean the data, remove sentences containing errors or incompleteness, and ensure the diversity of sentence length and complexity. (3) Multi-domain and diversity guarantee: Domain coverage: The data covers texts in different domains such as news, technology, literature, and dialogue. Difficulty level: Include daily language, technical terms, long sentences, sentences with ambiguity, etc.

[0079] Step 2: Evaluation Index Setting

[0080] Design of index system:

[0081] Expression accuracy (Accuracy): The degree to which the translation is semantically consistent with the source text.

[0082] Language fluency (Fluency): Whether the translated text conforms to English grammar and expression habits.

[0083] Terminology consistency: Whether the translation of technical terms is accurate and consistent.

[0084] Setting of index weights: The weight of expression accuracy is 0.5; the weight of language fluency is 0.3; the weight of terminology consistency is 0.2.

[0085] Writing of evaluation index and scoring standard description:

[0086] (1) Scoring standard for expression accuracy:

[0087] 5 points: Completely faithful to the original text, accurately conveying all information.

[0088] 4 points: Basically faithful, with minor deviations in details.

[0089] 3 points: There are some mistranslations or omissions, affecting understanding.

[0090] 2 points: Most of the content is inaccurate, seriously affecting understanding.

[0091] 1 point: Completely inconsistent with the original meaning.

[0092] (2) Language fluency scoring criteria:

[0093] 5 points: The expression is natural and fluent without language errors.

[0094] 4 points: Basically fluent, with occasional minor grammar or expression problems.

[0095] 3 points: There are obvious grammar or expression errors, but it does not affect the overall understanding.

[0096] 2 points: A large number of language errors affect the understanding.

[0097] 1 point: Unintelligible.

[0098] (3) Terminology consistency scoring criteria:

[0099] 5 points: Professional terms are translated accurately and consistently.

[0100] 4 points: A small number of terms are translated inconsistently or inaccurately.

[0101] 3 points: There are multiple inconsistencies or inaccuracies in the translation of terms.

[0102] 2 points: Most of the translations of terms are incorrect.

[0103] 1 point: The translations of terms are completely incorrect.

[0104] Step 3: Select an evaluation model

[0105] Requirements for the evaluation model: Select a large language model with Chinese and English understanding capabilities, and it is different from the machine translation model to be evaluated.

[0106] Selection of the evaluation model: Use an open multilingual large language model, such as using Qwen72B as the evaluation model, etc.

[0107] Step 4: Evaluation process

[0108] Figure 2 Shows the optimization flow chart of the large language model based on multi-round evaluation of the present invention, as Figure 2 shown, the detailed steps are as follows:

[0109] Organize the evaluation data into input instructions: Please evaluate the provided translation according to the evaluation indicators and scoring criteria, and give a score of 1-5 and corresponding comments for each indicator.

[0110] Evaluation content: (1) Original text: [Chinese sentence]; (2) Machine translation result: [Translation of the model to be evaluated]; (3) Reference translation: [Reference translation].

[0111] Please evaluate the machine translation result according to the above evaluation content information.

[0112] Evaluation Execution: Use a programming interface (such as calling an evaluation model API through a Python script) to implement the model calling operation, and pass the prepared input to the evaluation model. Multiple Evaluations: For each sample data, perform multiple model calling operations and calculate the average score of these multiple evaluation results.

[0113] Step 5: Real-time Feedback Analysis

[0114] Based on the evaluation results, the system analyzes the translation quality and generates feedback information.

[0115] Example Feedback Logic:

[0116] if accuracy_score < threshold:

[0117] feedback = "The translation quality is low. It is recommended to use an expression closer to the original text."

[0118] if terminology_consistency < threshold:

[0119] feedback += "Pay attention to the consistency of terminology."

[0120] Return the generated feedback information and the score to the user or the system together.

[0121] Example Output Format:

[0122] {

[0123] "score": 3.5,

[0124] "feedback": "The translation quality is low. It is recommended to use an expression closer to the original text. Pay attention to the consistency of terminology."

[0125] }

[0126] Step 6: Model Optimization Suggestions

[0127] After the evaluation, the system analyzes the scoring results and detects low-scoring items. For example, if it is found that the accuracy index score is low, the system starts the process of generating optimization suggestions and proposes specific model optimization measures based on the analyzed low-scoring items.

[0128] Example Logic:

[0129] if accuracy_score < threshold:

[0130] optimization_suggestion = "Increase the training data in related fields to improve the model's understanding of professional terms."

[0131] elif fluency_score < threshold:

[0132] optimization_suggestion = "Adjust the decoding strategy to improve the fluency of the generated text."

[0133] Return the generated optimization suggestion and the evaluation result to the user together.

[0134] Example output format:

[0135] {

[0136] "score": 3.2,

[0137] "optimization_suggestion": "Increase the training data in related fields to improve the model's understanding of professional terms."

[0138] }

[0139] In summary, the purpose of the present invention is to provide a general evaluation method based on a large model. The core lies in using a large model with powerful understanding and generation capabilities as a third-party evaluator to evaluate and dynamically feedback on the large model. By establishing an "evaluation - feedback optimization" closed loop, the present invention can achieve a comprehensive, objective, and automated multi-dimensional evaluation of the performance of the large model, and real-time feedback the evaluation results to the optimization process of the model, forming a dynamic iterative mechanism of "evaluation - feedback optimization".

[0140] Further referring to Figure 3 , as an implementation of the above method, in the second aspect, the present application provides an embodiment of an evaluation system 300 for a large language model, which can be specifically applied to various electronic devices. The system 300 includes the following modules:

[0141] A data acquisition module 310, configured to construct a sample data set, where the sample data set includes input data and its corresponding reference answers, and input the input data into the model to be evaluated to generate an output result;

[0142] An evaluation index design module 320, configured to design a multi-dimensional evaluation index system according to task requirements and assign dynamic weights to each evaluation index. Each evaluation index is accompanied by a description of the scoring criteria. The evaluation indexes include at least two of context relevance, term consistency, language fluency, and expression accuracy;

[0143] The evaluation result generation module 330 is configured to combine the input data of the data acquisition module, the reference answer, the output result, and the scoring standard description of the evaluation index design module into a standardized input instruction, call an evaluation model to perform multi-dimensional scoring on the standardized input instruction, and generate an evaluation result;

[0144] The feedback suggestion generation module 340 is configured to analyze the evaluation result according to a preset threshold and generate a structured feedback suggestion including an improvement direction.

[0145] In a third aspect, the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the evaluation method of the large language model as described in any one of the above are implemented.

[0146] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the evaluation method of the large language model as described in any one of the above are implemented.

[0147] Next, reference is made to Figure 4 which shows a schematic structural diagram of a computer system 400 of a terminal device or a server suitable for implementing the embodiments of the present application. Figure 4 The shown terminal device or server is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0148] As Figure 4 shown, the computer system 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage section 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the computer system 400 are also stored. The CPU 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0149] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 410 as needed so that a computer program read therefrom is installed in the storage section 408 as needed.

[0150] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 409 and / or installed from the removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, the above functions defined in the methods of the present application are performed. It should be noted that the computer-readable medium described in the present application can be a computer-readable signal medium, a computer-readable medium, or any combination of the two. The computer-readable medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0151] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as C language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).

[0152] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0153] The above description is only a preferred embodiment of this application and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in this application.

Claims

1. An evaluation method for large language models, characterized in that, It includes the following steps: S1. Construct a sample dataset, where the sample dataset includes input data and its corresponding reference answers, and input the input data into the model to be evaluated to generate an output result; S2. Design a multi-dimensional evaluation index system according to task requirements, and assign dynamic weights to each evaluation index. Each evaluation index is accompanied by a description of the scoring criteria. The evaluation indices include at least two of context relevance, term consistency, language fluency, and expression accuracy; S3. Combine the input data, reference answer, output result in step S1 and the scoring criteria description in step S2 into a standardized input instruction, and call an evaluation model to perform multi-dimensional scoring on the standardized input instruction to generate an evaluation result; S4. Analyze the evaluation result according to a preset threshold to generate a structured feedback suggestion including improvement directions.

2. The evaluation method of the large language model according to claim 1, wherein In step S2, designing a multi-dimensional evaluation index system according to task requirements and assigning dynamic weights to each evaluation index includes: In a machine translation task, the weight of expression accuracy is higher than the weight of language fluency; In a dialogue generation task, the weight of context relevance is higher than the weight of term consistency; In a text summarization task, the weight of language fluency is higher than the weight of term consistency.

3. The evaluation method of the large language model according to claim 1, characterized in that In step S4, analyzing the evaluation result according to a preset threshold to generate a structured feedback suggestion including improvement directions includes: If the context relevance score is lower than the first threshold, it is recommended to increase the dialogue history context attention mechanism; If the term consistency score is lower than the second threshold, it is recommended to inject a domain term dictionary into the training data or strengthen the alignment fine-tuning strategy; If the language fluency score is lower than the third threshold, it is recommended to adjust the text generation strategy; If the expression accuracy score is lower than the fourth threshold, it is recommended to add a domain-specific dataset for retraining.

4. The evaluation method of the large language model according to claim 1, wherein In step S1, constructing the sample dataset includes obtaining sample data from a public dataset and performing preprocessing operations such as cleaning, deduplication, and annotation on the sample data.

5. The evaluation method of the large language model according to claim 1, characterized in that In step S1, the sample dataset includes data in the fields of technology, medicine, finance, and entertainment. The difficulty levels of the sample dataset include simple sentences, complex sentences, and ambiguous sentences.

6. The evaluation method of the large language model according to claim 1, characterized in that, In step S3, it further includes: Selecting a third-party large language model different from the architecture and training data of the model to be evaluated as the evaluation model; Combining the input data, reference answer, output result, and scoring criteria description in a preset format by using a method that combines natural language description instructions with data to obtain a standardized input instruction.

7. The evaluation method of the large language model according to claim 1, characterized in that In step S3, calling the evaluation model to perform multi-dimensional scoring on the standardized input instruction to generate an evaluation result includes: Performing multiple evaluations on each standardized input instruction. Each time an evaluation is performed, different evaluation model parameter settings are used. Calculate the preliminary average value of the multiple scores obtained, and after excluding abnormal scores that deviate from the preliminary average value by more than a preset range, recalculate the final average value of the remaining scores, and use the final average value as the evaluation result.

8. An evaluation system for a large language model, characterized in that, The system includes: A data acquisition module, configured to construct a sample data set, where the sample data set includes input data and its corresponding reference answers, and input the input data into an evaluated model to generate an output result; An evaluation index design module, configured to design a multi-dimensional evaluation index system according to task requirements, and assign dynamic weights to each evaluation index. Each evaluation index is accompanied by a description of the scoring criteria. The evaluation index includes at least two of context relevance, term consistency, language fluency, and expression accuracy; An evaluation result generation module, configured to combine the input data, reference answers, output results of the data acquisition module, and the description of the scoring criteria of the evaluation index design module into a standardized input instruction, call an evaluation model to perform multi-dimensional scoring on the standardized input instruction, and generate an evaluation result; A feedback suggestion generation module, configured to analyze the evaluation result according to a preset threshold and generate a structured feedback suggestion including improvement directions.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the evaluation method of the large language model according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the evaluation method of the large language model according to any one of claims 1 to 7.

Citation Information

Cited By

  • Data generation method and device, equipment and storage medium

    CN120763308A

  • Data generation method and apparatus, device, and storage medium

    CN120763308B

  • Mirror operation description evaluation method and device, electronic equipment and storage medium

    CN120766188A

  • A camera movement description evaluation method and device, electronic equipment and storage medium

    CN120766188B

  • Model evaluation method and device for large language model, medium and equipment

    CN121008999A