Large model evaluation automatic scoring system and algorithm

By designing an automated scoring system for large model evaluation, the problem that large model evaluation relies on manual evaluation in the existing technology is solved, efficient and objective evaluation results are achieved, multi-modal and multi-scene evaluation is supported, and detailed visual reports are provided, which promotes the fairness and universality of the model.

CN119938519AInactive Publication Date: 2025-05-06BEIJING ZHONGKE JINCAI TECH

Patent Information

Application Number
CN202411871077.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the evaluation and comparison of large models mainly relies on manual evaluation, which makes it time-consuming and labor-intensive, and is highly subjective, making it difficult to ensure the consistency and accuracy of the evaluation results.

Method used

An automated scoring system for large-model evaluation was designed, including a data management unit, an evaluation task management unit, an evaluation execution unit, a scoring and calculation unit and a result display unit. Through the automated evaluation process, data upload, storage, cleaning and formatting, task creation and management, model call and evaluation results are realized, and the evaluation results are finally displayed intuitively.

Benefits of technology

It significantly improves the evaluation efficiency, reduces manual participation, ensures the objectivity and consistency of the evaluation results, supports multimodal and multi-scenario evaluation, provides detailed visual reports, and promotes the fairness and universality of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938519A_ABST
    Figure CN119938519A_ABST
Patent Text Reader

Abstract

The invention discloses a large model evaluation automatic scoring system and algorithm. The system comprises an evaluation system. The evaluation system comprises a data management unit, an evaluation task management unit, an evaluation execution unit, an evaluation calculation unit and a result display unit. The data management unit is responsible for managing and processing a data set required by evaluation; the functions of uploading, storing, cleaning and formatting data are realized; the evaluation task management unit is responsible for creating and managing evaluation tasks and selecting an evaluation data set, a model and a Prompt template; the evaluation execution unit is used for executing the evaluation task, calling a corresponding model and a corresponding data set for evaluation and collecting an evaluation result; the scoring and calculating unit is responsible for scoring and calculating the evaluation result to generate a final evaluation score; and the result display unit is responsible for visually displaying the evaluation result to the user. Through an automatic scoring algorithm, manual participation is remarkably reduced, the evaluation efficiency is improved, and large-scale models and data sets can be rapidly processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large model evaluation automatic scoring, and in particular to a large model evaluation automatic scoring system and algorithm. Background Art

[0002] With the development of artificial intelligence technology, large models have made significant progress in the field of natural language processing. However, in practical applications, how to accurately and comprehensively evaluate and compare these large models has become an important issue. Traditional evaluation methods mainly rely on manual evaluation, which is not only time-consuming and labor-intensive, but also highly subjective and difficult to ensure the consistency and accuracy of the evaluation results. Therefore, we need an automated evaluation method to improve evaluation efficiency and reduce the impact of human factors. Summary of the invention

[0003] The purpose of the present invention is to provide a large model evaluation automatic scoring system and algorithm, so as to solve the above-mentioned problems existing in the prior art.

[0004] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0005] A large model evaluation automatic scoring system, comprising: an evaluation system; the evaluation system comprises: a data management unit, an evaluation task management unit, an evaluation execution unit, an evaluation calculation unit and a result display unit;

[0006] The data management unit is responsible for managing and processing the data sets required for evaluation; it has the functions of uploading, storing, cleaning and formatting data;

[0007] The evaluation task management unit is responsible for creating and managing evaluation tasks, selecting evaluation data sets, models, and prompt templates;

[0008] The evaluation execution unit executes the evaluation tasks, calls the corresponding models and data sets for evaluation, and collects the evaluation results;

[0009] The scoring and calculation unit is responsible for scoring and calculating the evaluation results to generate the final evaluation score;

[0010] The result display unit is responsible for displaying the evaluation results intuitively to the user.

[0011] In a specific embodiment, the data management unit includes:

[0012] Data upload module: allows users to upload new evaluation data sets;

[0013] Data cleaning module: allows users to label uploaded unformatted data and / or use automatic cleaning and formatting functions;

[0014] Data storage module: Store the cleaned data in the system for subsequent evaluation.

[0015] In a specific embodiment, the evaluation task management unit includes:

[0016] Task creation module: allows users to create new assessment tasks;

[0017] Dataset selection module: selects stored datasets for evaluation;

[0018] Prompt template management module: manage and select prompt templates required for evaluation;

[0019] Model selection module: select the list of models to be evaluated.

[0020] In a specific embodiment, the evaluation execution unit includes:

[0021] Evaluation scheduling module: schedules the execution of evaluation tasks;

[0022] Model calling module: calls the model to be evaluated for testing;

[0023] Data processing module: processes the data generated during the evaluation process.

[0024] In one specific embodiment, the scoring and calculation unit includes:

[0025] Scoring module: Scores the evaluation results according to the set scoring dimensions and weights.

[0026] Score calculation module: calculates the final score of each model.

[0027] In a specific embodiment, the result display unit includes:

[0028] Visualization module: used to generate radar charts, heat maps and trend chart visualization results;

[0029] Report generation module: used to generate detailed evaluation reports.

[0030] In a specific embodiment, the following steps are also included:

[0031] Collect the data sets that need to be evaluated through the data management unit and design the evaluation dimensions;

[0032] Create and manage evaluation tasks through the evaluation task management unit, and select evaluation data sets, evaluation models, and prompt templates;

[0033] Evaluate the evaluation tasks through the evaluation execution unit and collect the evaluation results;

[0034] The evaluation results are scored and calculated by the scoring and calculation unit to generate the final evaluation score; the evaluation is conducted from the perspective of evaluation task quality indicators and performance indicators;

[0035] Finally, the result display unit displays the evaluation scores intuitively to the user.

[0036] In a specific embodiment, the evaluation dimensions include: accuracy, colloquialism, interactivity, completeness, fluency and comprehension analysis.

[0037] In a specific embodiment, the quality indicators include: accuracy, recall, precision, F1-Score and scoring score;

[0038] Among them, the calculation method of accuracy is: Accuracy = (number of correct predictions) / (total number of predictions)

[0039]

[0040] Recall rate calculation method: Recall rate = (number of correctly identified positive samples) / (actual number of positive samples)

[0041]

[0042] Calculation method of precision rate: Precision rate = (number of positive samples correctly identified) / (number of samples predicted to be positive)

[0043]

[0044] F1-Score calculation method: F1-Score = 2*(precision*recall) / (precision+recall)

[0045]

[0046] Among them, TP is true positive, TN is true negative, FP is false positive, and FN is false negative;

[0047] Calculation method of scoring:

[0048] (Accuracy score + oral score + interactivity score + completeness score + fluency score + understanding and analytical score) / 6.

[0049] In a specific embodiment, the performance indicators include: first token latency, complete answer latency, large model throughput, and large model QPS capability;

[0050] in,

[0051] First Token delay calculation method: First Token delay = first Token received time - request initiation time

[0052] First Token Delay = T first_token_received -T request_sent

[0053] Calculate the normalized score of the first token latency:

[0054]

[0055] T min and T max They are the minimum and maximum values ​​of the first token delay of all models respectively;

[0056] Calculation method of complete answer delay: Complete answer delay = answer completion time - request initiation time

[0057] Complete answer delay = T response_completed -T request_sent

[0058] Calculate the standardized score for complete response latency:

[0059]

[0060] Among them, Tresponse_completed is the complete response delay of the model, Tmin and Tmax are the minimum and maximum complete response delays of all models respectively;

[0061] How to calculate the throughput of a large model: Throughput = number of requests processed per unit time

[0062]

[0063] Calculate the normalized score for large model throughput:

[0064]

[0065] Q throughput is the throughput of the model, Q min and Q max are the minimum and maximum throughputs of all models, respectively;

[0066] Calculation method of large model QPS capacity: QPS = number of requests processed per unit time

[0067]

[0068] Calculate the normalized score of the large model's QPS capability:

[0069]

[0070] Among them, Q QPSis the QPS capability of the model, Qmin and Qmax are the minimum and maximum QPS capabilities of all models, respectively.

[0071] The beneficial effects of the present invention are as follows: the present invention discloses a large model evaluation automatic scoring system and algorithm, including: a data management unit, an evaluation task management unit, an evaluation execution unit, an evaluation calculation unit and a result display unit; the data management unit is responsible for managing and processing the data set required for the evaluation; it has the functions of uploading, storing, cleaning and formatting data; the evaluation task management unit is responsible for creating and managing evaluation tasks, selecting evaluation data sets, models and prompt templates; the evaluation execution unit executes evaluation tasks, calls corresponding models and data sets for evaluation, and collects evaluation results; the scoring and calculation unit is responsible for scoring and calculating the evaluation results and generating the final evaluation score; the result display unit is responsible for intuitively displaying the evaluation results to the user. The present invention has the following specific features:

[0072] 1. Improve evaluation efficiency:

[0073] Automated evaluation: Through automated scoring algorithms, manual involvement is significantly reduced, evaluation efficiency is improved, and large-scale models and data sets can be processed quickly.

[0074] Real-time feedback: The evaluation system can generate evaluation results and reports in real time, shortening the evaluation cycle and speeding up model iteration.

[0075] 2. Ensure the objectivity and consistency of evaluation results:

[0076] Standardized evaluation: Use unified evaluation standards and processes to ensure the objectivity and consistency of evaluation results and avoid the influence of human factors.

[0077] Multi-dimensional evaluation: Design multi-dimensional evaluation indicators, including quality indicators and performance indicators, to comprehensively evaluate all aspects of the model's performance.

[0078] 3. Support multi-modal and multi-scenario evaluation:

[0079] Multimodal support: Supports the evaluation of multimodal data such as text, images, and audio, and adapts to different types of AI models.

[0080] Multi-scenario application: The evaluation system can design evaluation tasks based on specific application scenarios and evaluate the performance of the model in actual business.

[0081] 4. Provide visual evaluation results:

[0082] Intuitive display: Visual tools such as radar charts, heat maps, and trend charts are used to intuitively display evaluation results, helping users quickly understand model performance.

[0083] Detailed report: Generate a detailed evaluation report, including specific data and analysis of each evaluation indicator, to provide a basis for model optimization.

[0084] 5. Promote model fairness and universality:

[0085] Fairness evaluation: The evaluation system can evaluate the performance of the model in different populations, languages, and cultural backgrounds to ensure the fairness of the model.

[0086] Universality verification: Through diverse data sets and evaluation tasks, the universality of the model in various application scenarios is verified to enhance the model's ability to be widely applied. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] Figure 1 It is a flow chart of a large model evaluation automated scoring algorithm of the present invention;

[0088] Figure 2 It is a system block diagram of a large model evaluation automatic scoring system of the present invention;

[0089] Figure 3 It is a complete evaluation pipeline diagram of the present invention;

[0090] Figure 4 is the evaluation flow chart of the present invention;

[0091] Figure 5 This is a complete example diagram of the evaluation prompt of the present invention. DETAILED DESCRIPTION

[0092] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation methods described herein are only used to explain the present invention and are not used to limit the present invention.

[0093] Reference Figure 1 , Figure 2 , Figure 3 , Figure 4 ,and Figure 5 A large model evaluation automatic scoring system shown includes: an evaluation system, the evaluation system includes: a data management unit, an evaluation task management unit, an evaluation execution unit, an evaluation calculation unit and a result display unit;

[0094] The data management unit is responsible for managing and processing the data sets required for evaluation; it has the functions of uploading, storing, cleaning and formatting data;

[0095] The evaluation task management unit is responsible for creating and managing evaluation tasks, selecting evaluation data sets, models, and prompt templates;

[0096] The evaluation execution unit executes the evaluation tasks, calls the corresponding models and data sets for evaluation, and collects the evaluation results;

[0097] The scoring and calculation unit is responsible for scoring and calculating the evaluation results to generate the final evaluation score;

[0098] The result display unit is responsible for displaying the evaluation results intuitively to the user.

[0099] In a specific embodiment, the data management unit includes:

[0100] Data upload module: allows users to upload new evaluation data sets;

[0101] Data cleaning module: allows users to label uploaded unformatted data and / or use automatic cleaning and formatting functions;

[0102] Data storage module: Store the cleaned data in the system for subsequent evaluation.

[0103] In a specific embodiment, the evaluation task management unit includes:

[0104] Task creation module: allows users to create new assessment tasks;

[0105] Dataset selection module: selects stored datasets for evaluation;

[0106] Prompt template management module: manage and select prompt templates required for evaluation;

[0107] Model selection module: select the list of models to be evaluated.

[0108] In a specific embodiment, the evaluation execution unit includes:

[0109] Evaluation scheduling module: schedules the execution of evaluation tasks;

[0110] Model calling module: calls the model to be evaluated for testing;

[0111] Data processing module: processes the data generated during the evaluation process.

[0112] In one specific embodiment, the scoring and calculation unit includes:

[0113] Scoring module: Scores the evaluation results according to the set scoring dimensions and weights.

[0114] Score calculation module: calculates the final score of each model.

[0115] In a specific embodiment, the result display unit includes:

[0116] Visualization module: used to generate radar charts, heat maps and trend chart visualization results;

[0117] Report generation module: used to generate detailed evaluation reports.

[0118] A large model evaluation automatic scoring algorithm based on the same concept includes the following steps:

[0119] S100. Collect the data set to be evaluated through the data management unit and design the evaluation dimensions;

[0120] S200, create and manage evaluation tasks through the evaluation task management unit, and select evaluation data sets, evaluation models and prompt templates;

[0121] S300, evaluating the evaluation task through the evaluation execution unit, and collecting the evaluation results;

[0122] S400, scoring and calculating the evaluation results through the scoring and calculation unit to generate a final evaluation score; evaluating from the perspective of evaluation task quality indicators and performance indicators;

[0123] In this embodiment, quality indicators are mainly used to evaluate the accuracy and effectiveness of the model on a specific task. These indicators are key factors in measuring model performance because they directly reflect the performance of the model when processing actual tasks.

[0124] Performance indicators are mainly used to evaluate the efficiency and ability of the model in processing tasks. These indicators directly affect the usability and user experience of the model in practical applications.

[0125] S500. Finally, the result display unit displays the evaluation scores to the user intuitively.

[0126] In a specific embodiment, the evaluation dimensions include: accuracy, colloquialism, interactivity, completeness, fluency and comprehension analysis.

[0127] In a specific embodiment, the quality indicators include: accuracy, recall, precision, F1-Score (objective data set) and scoring score (subjective data set);

[0128] Accuracy:

[0129] Reason: Accuracy is one of the most basic performance indicators, measuring the correctness of the model's overall predictions. For most application scenarios, accuracy is a very intuitive metric.

[0130] Logic: A high accuracy means that the model is able to make correct decisions in most cases.

[0131] Among them, the calculation method of accuracy is: Accuracy = (number of correct predictions) / (total number of predictions)

[0132]

[0133] Recall:

[0134] Reason: In some applications, such as medical diagnosis or security monitoring, failure to identify the positive class (e.g., patient, threat) may have serious consequences, so recall is very important.

[0135] Logic: A high recall rate means that the model can identify more positive samples and reduce false negatives.

[0136] Recall rate calculation method: Recall rate = (number of correctly identified positive samples) / (actual number of positive samples)

[0137]

[0138] Precision:

[0139] Reason: In some applications, such as spam filtering or recommendation systems, false positives can lead to negative experiences or waste of resources, so accuracy is critical.

[0140] Logic: A high precision rate means that the model's positive predictions are more accurate and reduce false positives.

[0141] Calculation method of precision rate: Precision rate = (number of positive samples correctly identified) / (number of samples predicted to be positive)

[0142]

[0143] F1-Score:

[0144] Reason: There is often a trade-off between accuracy and recall. F1-Score, as their harmonic mean, can provide a comprehensive evaluation.

[0145] Logic: F1-Score balances precision and recall and is suitable for application scenarios that need to consider both.

[0146] F1-Score calculation method: F1-Score = 2*(precision*recall) / (precision+recall)

[0147]

[0148] Among them, TP (True Positive) is true positive, TN (True Negative) is true negative, FP (False Positive) is false positive, and FN (False Negative) is false negative.

[0149] Subjective Score:

[0150] Reason: For some tasks with strong subjectivity (such as text generation and dialogue systems), traditional quality indicators are difficult to fully reflect the performance of the model and need to be evaluated through manual or advanced model scoring.

[0151] Logic: The scoring can reflect the performance of the model on subjective tasks and provide a more comprehensive evaluation perspective.

[0152] The dimensions include:

[0153] Accuracy: Accuracy is used to evaluate the relevance and accuracy of the answers generated by the model and the content of the user's question. The judgment of accuracy is somewhat subjective, so it is evaluated from a scoring perspective. (Out of 5 points)

[0154] Colloquialism: Colloquialism is used to evaluate the naturalness and fluency of the answers generated by the model in actual conversation scenarios. The judgment of colloquialism is somewhat subjective, so it is evaluated from a scoring perspective (out of 5 points)

[0155] Interactivity: Interactivity is used to evaluate whether the answers generated by the model can guide users to interact and enhance user participation. The judgment of interactivity is subjective, so it is evaluated from a scoring perspective.

[0156] (Out of 5 points)

[0157] Completeness: Completeness is used to evaluate whether the answers generated by the model can fully cover the key information of the user's question. The judgment of completeness is subjective, so it is evaluated from a scoring perspective. (Full score 5 points)

[0158] Fluency: Fluency is used to evaluate the fluency of the answers generated by the model in terms of grammar and expression. The judgment of fluency is somewhat subjective, so it is evaluated from a scoring perspective. (Out of 5 points)

[0159] Understanding and analytical ability: Understanding and analytical ability is used to evaluate the model's understanding and analytical ability of user questions. The judgment of understanding and analytical ability is subjective, so it is evaluated from a scoring perspective. (Full score 5 points)

[0160] Calculation method of scoring:

[0161] (Accuracy score + oral score + interactivity score + completeness score + fluency score + understanding and analytical score) / 6.

[0162] In a specific embodiment, the performance indicators include: first token latency, complete answer latency, large model throughput, and large model QPS capability;

[0163] i. First Token Latency: The latency from request initiation to receipt of the first response Token;

[0164] ii. Complete response delay: the delay from request initiation to response completion;

[0165] iii. Large model throughput: the maximum number of requests that a large model can accept at a time;

[0166] iv. Large model QPS capability: Request Per Second, the maximum number of tokens supported by a single request;

[0167] First Token Delay:

[0168] Reason: In conversational systems or real-time processing applications, response speed is crucial, and the first token delay directly affects the user experience.

[0169] Logic: A shorter first-token latency means that the model can respond quickly, improving user satisfaction.

[0170] in,

[0171] First Token delay calculation method: First Token delay = first Token received time - request initiation time

[0172] First Token Delay = T first_token_received -T request_sent

[0173] Calculate the normalized score of the first token latency (quantized to between 0 and 1):

[0174]

[0175] T min and T max They are the minimum and maximum values ​​of the first token delay of all models respectively;

[0176] Full response delay:

[0177] Reason: For tasks that require a complete answer, such as question answering systems or generative tasks, the latency to complete the answer is a key performance indicator.

[0178] Logic: Shorter complete answer latency improves overall system efficiency and user experience.

[0179] Calculation method of complete answer delay: Complete answer delay = answer completion time - request initiation time

[0180] Complete answer delay = T response_completed -T request_sent

[0181] Calculate the normalized score (quantized between 0-1) for the complete answer latency:

[0182]

[0183] Among them, Tresponse_completed is the complete response delay of the model, Tmin and Tmax are the minimum and maximum complete response delays of all models respectively;

[0184] Large model throughput:

[0185] Reason: In high-concurrency scenarios (such as online services and batch processing), the throughput of the model determines the processing capacity of the system.

[0186] Logic: High throughput means that the model can process more requests per unit time, improving the service capacity of the system.

[0187] How to calculate the throughput of a large model: Throughput = number of requests processed per unit time

[0188]

[0189] Calculate the normalized score (quantized between 0-1) for the throughput of the large model:

[0190]

[0191] Q throughput is the throughput of the model, Q min and Q max are the minimum and maximum throughputs of all models, respectively;

[0192] Large model QPS capability (RequestPerSecond):

[0193] Reason: QPS is a key indicator to measure the processing capacity of the system under high load, especially in applications that require real-time response.

[0194] Logic: High QPS capability means that the model can process more requests per unit time, improving the response speed and stability of the system.

[0195] Calculation method of large model QPS capacity: QPS = number of requests processed per unit time

[0196]

[0197] Calculate the standardized score of the large model's QPS capability (quantized between 0 and 1):

[0198]

[0199] Among them, Q QPS is the QPS capability of the model, Qmin and Qmax are the minimum and maximum QPS capabilities of all models, respectively.

[0200] Evaluation Dataset

[0201] a) Benchmark test set: a public open source dataset that tests the general capabilities of large models

[0202] b) Scenario datasets: questions and tasks designed based on specific evaluation requirements to test the capabilities of large models in certain actual businesses.

[0203] c) Adversarial test set: used to evaluate the performance of the model in extreme or boundary cases.

[0204] Prompt Project

[0205] a) Prompt for large model testing: When testing scenario datasets that do not have standard answers, it is necessary to assemble a prompt for the test. Prompts are constructed for different dataset types, but in the same dataset, different large models tested use the same prompt.

[0206] b) Judgment prompt: For subjective scene datasets, the datasets are relatively large and there are many models to be tested. It is unrealistic to rely on manual review and scoring. In addition, different people will inevitably have different understandings of the standards, and the objectivity of the scores cannot be guaranteed. Therefore, we use a more powerful super model for judgment. This requires the use of a judgment prompt.

[0207] i. Judgement prompt example: see Figure 5 Example of a judgement prompt

[0208] Testing Process

[0209] a) Upload the test dataset:

[0210] The test system has built-in relevant data sets mentioned above. If you need to test a new data set, you can upload it on the front end. If it is a standard data set, you can test it directly after uploading.

[0211] If it is subjective and non-standard data, it will be automatically extracted, QA pairs, data cleaning and other steps, and organized into a data set format.

[0212] b) Select the test data set and create a test task

[0213] Users select a dataset for testing from built-in or uploaded datasets.

[0214] Create a new test task, specify the test task name, description, etc.

[0215] c) Select prompt template

[0216] The system provides preset Prompt templates for selection.

[0217] If the existing templates do not meet your needs, you can create a new Prompt template and add it to the test task.

[0218] d) Select the model set:

[0219] The model set represents a list of models to be tested. All models in the list are tested under the test set and complete the subsequent scoring process.

[0220] e) Judge model scoring (if the selected dataset is a scenario dataset without a standard answer)

[0221] (Scene datasets require a referee model to score)

[0222] 1. Assume that the test set Test_dataset has 100 test data.

[0223] 2. The total score of a single piece of data is sum_score

[0224] Cockpit: 25 points, based on five indicators (readability, rationality, coherence, anthropomorphism, and relevance), each with a score of 5 points

[0225] E-commerce: 30 points, based on six indicators (accuracy, colloquialism, interactivity, completeness, fluency, and comprehension and analysis), with each indicator worth 5 points

[0226] 3. The super model scores the test model mode l and the benchmark model GPT3.5 respectively, denoted as A and B

[0227] 4. Final score of the test model = (A0+A1+...+A99) / (100*sum_score)

[0228] 5. Final score of the baseline model = (B0+B1+...+B99) / (100*sum_score)

[0229] f) Calculation of final model score

[0230] Weight setting: Through expert discussion and data analysis, the weight of each scoring dimension (such as accuracy, generation quality, task completion, etc.) is set.

[0231] Rating calculation: Each rating dimension is scored independently, and the comprehensive score is calculated based on the set weights.

[0232] After completing the test of all data sets, the scores of the test model on each data set will be normalized to a score between 0 and 1. Different weights will be assigned to the scores of each test set. The weights are multiplied by the scores of the data sets. The final score is the total score of the model.

[0233] g) Visualization of evaluation results:

[0234] After the test is completed, an intuitive visual interface is automatically output to display the evaluation results: a) Radar chart: Displays the scores of each evaluation dimension. b) Heat map: Displays the performance of the model on different task types. c) Trend chart: Tracks the changes in model performance over time;

[0235] The scoring process is as follows:

[0236] 1. Calculation of single score:

[0237] The score of each evaluation dimension (such as accuracy, recall, precision, F1-Score, first token latency, complete answer latency, large model throughput, large model QPS capability, etc.) is calculated according to the above calculation method.

[0238] 2. Weight setting:

[0239] Through expert discussion and data analysis, the weight of each scoring dimension (such as accuracy, generation quality, task completion, etc.) is set.

[0240] The weight wiwi of each dimension satisfies the following conditions:

[0241]

[0242] Where n is the number of evaluation dimensions.

[0243] 3. Comprehensive score calculation:

[0244] oEach scoring dimension is scored independently and a comprehensive score is calculated based on the set weights.

[0245] oComprehensive scoring formula:

[0246]

[0247] Among them, Xnormalized,i is the standardized score of the i-th dimension, and wi is the weight of the corresponding dimension.

[0248] 4. Dataset scoring calculation:

[0249] For each dataset, calculate the final score of the model on that dataset.

[0250] Dataset scoring formula:

[0251]

[0252] Among them, the single data score j is the score of the jj-th data, and m is the total number of data in the data set.

[0253] 5. Total score calculation:

[0254] The scores of the model on different data sets are weighted averaged to obtain the total score of the model.

[0255] Total score calculation formula:

[0256]

[0257] Among them, the dataset score k is the score of the kth dataset, wk is the weight of the corresponding dataset, and p is the number of datasets.

[0258] h. Visualization of evaluation results:

[0259] A. After the test is completed, the system provides an intuitive visual interface to display the evaluation results:

[0260] 1. Radar chart: displays the scores of each evaluation dimension.

[0261] 2. Heatmap: Shows the performance of the model on different task types.

[0262] 3. Trend graphs: Track model performance over time.

[0263] The beneficial effects of the present invention are as follows: the present invention discloses an automatic scoring algorithm for large model evaluation, including: a data management unit, an evaluation task management unit, an evaluation execution unit, an evaluation calculation unit and a result display unit; the data management unit is responsible for managing and processing the data set required for the evaluation; it has the functions of uploading, storing, cleaning and formatting data; the evaluation task management unit is responsible for creating and managing evaluation tasks, selecting evaluation data sets, models and prompt templates; the evaluation execution unit executes evaluation tasks, calls corresponding models and data sets for evaluation, and collects evaluation results; the scoring and calculation unit is responsible for scoring and calculating the evaluation results to generate the final evaluation score; the result display unit is responsible for intuitively displaying the evaluation results to the user. The present invention has the following specific advantages:

[0264] 1. Improve evaluation efficiency:

[0265] Automated evaluation: Through automated scoring algorithms, manual involvement is significantly reduced, evaluation efficiency is improved, and large-scale models and data sets can be processed quickly.

[0266] Real-time feedback: The evaluation system can generate evaluation results and reports in real time, shortening the evaluation cycle and speeding up model iteration.

[0267] 2. Ensure the objectivity and consistency of evaluation results:

[0268] Standardized evaluation: Use unified evaluation standards and processes to ensure the objectivity and consistency of evaluation results and avoid the influence of human factors.

[0269] Multi-dimensional evaluation: Design multi-dimensional evaluation indicators, including quality indicators and performance indicators, to comprehensively evaluate all aspects of the model's performance.

[0270] 3. Support multi-modal and multi-scenario evaluation:

[0271] Multimodal support: Supports the evaluation of multimodal data such as text, images, and audio, and adapts to different types of AI models.

[0272] Multi-scenario application: The evaluation system can design evaluation tasks based on specific application scenarios and evaluate the performance of the model in actual business.

[0273] 4. Provide visual evaluation results:

[0274] Intuitive display: Visual tools such as radar charts, heat maps, and trend charts are used to intuitively display evaluation results, helping users quickly understand model performance.

[0275] Detailed report: Generate a detailed evaluation report, including specific data and analysis of each evaluation indicator, to provide a basis for model optimization.

[0276] 5. Promote model fairness and universality:

[0277] Fairness evaluation: The evaluation system can evaluate the performance of the model in different populations, languages, and cultural backgrounds to ensure the fairness of the model.

[0278] Universality verification: Through diverse data sets and evaluation tasks, the universality of the model in various application scenarios is verified to enhance the model's ability to be widely applied.

[0279] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be considered as the scope of protection of the present invention.

Claims

1. A large model evaluation automatic scoring system, characterized in that: Including evaluation system; The evaluation system includes: a data management unit, an evaluation task management unit, an evaluation execution unit, an evaluation calculation unit and a result display unit; The data management unit is responsible for managing and processing the data sets required for the evaluation; it has the functions of uploading, storing, cleaning and formatting data; The evaluation task management unit is responsible for creating and managing evaluation tasks, and selecting evaluation data sets, models, and prompt templates; The evaluation execution unit executes the evaluation task, calls the corresponding model and the data set for evaluation, and collects the evaluation results; The scoring and calculation unit is responsible for scoring and calculating the evaluation results to generate a final evaluation score; The result display unit is responsible for displaying the evaluation results to the user intuitively.

2. The large model evaluation automatic scoring system according to claim 1 is characterized in that: The data management unit comprises: Data upload module: allows users to upload new evaluation data sets; Data cleaning module: allows users to label uploaded unformatted data and / or use automatic cleaning and formatting functions; Data storage module: Store the cleaned data in the system for subsequent evaluation.

3. The large model evaluation automatic scoring system according to claim 1 is characterized in that: The evaluation task management unit includes: Task creation module: allows users to create new assessment tasks; Dataset selection module: selects stored datasets for evaluation; Prompt template management module: manage and select prompt templates required for evaluation; Model selection module: select the list of models to be evaluated.

4. The large model evaluation automatic scoring system according to claim 1 is characterized in that: The evaluation execution unit comprises: Evaluation scheduling module: schedules the execution of evaluation tasks; Model calling module: calls the model to be evaluated for testing; Data processing module: processes the data generated during the evaluation process.

5. The large model evaluation automatic scoring system according to claim 1 is characterized in that: The scoring and calculation unit comprises: Scoring module: Scores the evaluation results according to the set scoring dimensions and weights. Score calculation module: calculates the final score of each model.

6. The large model evaluation automatic scoring system according to claim 1 is characterized in that: The result display unit comprises: Visualization module: used to generate radar charts, heat maps and trend chart visualization results; Report generation module: used to generate detailed evaluation reports.

7. An automated scoring algorithm for large model evaluation, characterized in that: The scoring steps include: Collect the data sets that need to be evaluated through the data management unit and design the evaluation dimensions; Create and manage evaluation tasks through the evaluation task management unit, and select evaluation data sets, evaluation models, and prompt templates; Evaluate the evaluation task through an evaluation execution unit and collect evaluation results; Scoring and calculating the evaluation results through a scoring and calculation unit to generate a final evaluation score; evaluating from the perspective of the evaluation task quality index and performance index; Finally, the result display unit displays the evaluation score intuitively to the user.

8. The large model evaluation automatic scoring algorithm according to claim 7 is characterized in that: The evaluation dimensions include: accuracy, colloquialism, interactivity, completeness, fluency and comprehension and analysis.

9. The large model evaluation automatic scoring algorithm according to claim 8 is characterized in that: The quality indicators include: accuracy, recall, precision, F1-Score and scoring score; The accuracy rate is calculated as follows: Accuracy rate = (number of correct predictions) / (total number of predictions) The recall rate calculation method is: Recall rate = (number of correctly identified positive samples) / (actual number of positive samples) The calculation method of the precision rate is: precision rate = (number of correctly identified positive samples) / (number of samples predicted to be positive) The calculation method of the F1-Score is: F1-Score = 2*(precision*recall) / (precision+recall) Among them, TP is true positive, TN is true negative, FP is false positive, and FN is false negative; The scoring score is calculated as follows: (Accuracy score + oral score + interactivity score + completeness score + fluency score + understanding and analytical score) / 6.

10. The large model evaluation automatic scoring algorithm according to claim 9 is characterized in that: The performance indicators include: first token latency, complete answer latency, large model throughput, and large model QPS capability; in, The calculation method of the first Token delay is: First Token delay = first Token received time - request initiation time First Token Delay = T first token received -T request_sent Calculate the normalized score of the first Token latency: T min and T max They are the minimum and maximum values ​​of the first token delay of all models respectively; The calculation method of the complete answer delay is: Complete answer delay = answer completion time - request initiation time Complete answer delay = T response_completed -T request_sent Calculate the normalized score for the complete answer latency: Among them, Tresponse_completed is the complete response delay of the model, Tmin and Tmax are the minimum and maximum complete response delays of all models respectively; The calculation method of the throughput of the large model is: Throughput = number of requests processed per unit time Calculate the normalized score for the throughput of the large model: Q throughput is the throughput of the model, Q min and Q max are the minimum and maximum throughputs of all models, respectively; The calculation method of the QPS capacity of the large model is: QPS = the number of requests processed per unit time Calculate the normalized score of the large model's QPS capability: Among them, Q QPS is the QPS capability of the model, Qmin and Qmax are the minimum and maximum QPS capabilities of all models, respectively.

Citation Information

Patent Citations

  • Model evaluation method and device, electronic equipment and storage medium

    CN117272011A

  • Multi-dimensional and multi-angle automatic large model testing system and method

    CN117785664A

  • Large language model evaluation method and device, electronic equipment and storage medium

    CN118035807A

  • Multi-index extensible multi-agent code generation evaluation system and method

    CN118838629A

Cited By

  • Test method and device for transverse evaluation of large model performance

    CN120371654A

  • All-in-one machine performance evaluation method and electronic equipment

    CN120950363A

  • Method and system for testing high concurrency performance of arrival time of first token in HTTP (Hyper Text Transport Protocol) streaming transmission

    CN121333977A