Terminal efficiency evaluation method and device for large language model
By constructing a multi-dimensional test data set and setting performance index thresholds, terminal performance evaluation of large language models is solved, and the problem that the existing technology cannot comprehensively evaluate the overall performance of large language models is achieved, and accurate evaluation and iterative tuning support for model end-to-end performance are achieved.
Patent Information
- Application Number
- CN202411824576.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-06
AI Technical Summary
When evaluating the effectiveness of large language models, the prior art focuses on the training and inference efficiency of the model, and cannot fully reflect the overall effectiveness of the model, and the effectiveness evaluation in specific tasks or knowledge areas requires the preparation of test data separately.
Provide a terminal performance evaluation method and device for large language models. By constructing a multi-dimensional test data set and setting performance index thresholds, load requests and performance index records are performed, and performance analysis reports are output.
It realizes a comprehensive, automatic, fast and accurate evaluation of the end-to-end performance of large language models, blocks network transmission losses, provides a performance analysis report that can reflect the user experience of end users, and supports iterative training and tuning of large language models.
Smart Images

Figure CN119938465A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of large model evaluation, and in particular to a terminal performance evaluation method and device for a large language model. Background Art
[0002] At present, large language models (LLMs) are experiencing explosive growth. Various enterprises, universities and research institutions are developing them. There are also many models for large language models in China, including general and subdivided models. In the development and tuning of large language models, performance is also a very important link. At present, the performance evaluation of large language models mostly focuses on the training and reasoning efficiency of the model, which cannot better reflect the overall performance of the model. There is still a certain error between the model reasoning performance and the user's intuitive experience, such as some time-consuming steps such as data forwarding, security recording and filtering, and intelligent analysis within the model.
[0003] Chinese Patent Publication No. CN118035104A discloses a system load stress testing method for a large language model, which aims to solve the problem of low generation efficiency in existing sponge test sample generation methods. Chinese Patent Publication No. CN118070775A discloses a performance evaluation method, device, and computer equipment for a summary generation model. This patent evaluates the ability of a target summary generation model by comparing it with other models. This patent only targets the summary generation ability of the summary model. In general, the prior art generally only conducts an overall performance evaluation of the model in terms of model performance testing to evaluate the overall model output performance. If you want to know the performance of the model in certain specific tasks or knowledge fields, you need to prepare corresponding test data separately for testing. Summary of the invention
[0004] The present invention provides a terminal performance evaluation method and device for a large language model, so as to solve the technical problems existing in the above-mentioned prior art.
[0005] To achieve the above object, the present invention provides a terminal performance evaluation method for a large language model, which includes:
[0006] Data construction step: According to the characteristics of the large language model and the fields involved, a multi-dimensional test data set is constructed from the business direction and general capability direction of the large language model. The multi-dimensional test data set includes multiple types, and each type is further divided into data sets of different difficulty. The data set includes multiple test data;
[0007] Indicator construction step, for multi-dimensional test data sets, performance indicator thresholds are set according to types;
[0008] Model testing step: Use a multi-dimensional test data set to initiate a load request to the large language model, call and execute the groups within the large language model sequentially or randomly, and record various performance indicators and execution results during the execution process;
[0009] The performance evaluation step outputs a performance analysis report based on the performance indicators and execution results recorded in the model testing step.
[0010] In one embodiment of the present invention, the data sets are divided into simple difficulty data sets, medium difficulty data sets and difficult difficulty data sets according to the difficulty.
[0011] In one embodiment of the present invention, in the data construction step, the field involved is the financial and taxation field.
[0012] In one embodiment of the present invention, the types of multidimensional test data sets include table analysis, chart construction, function rewriting, multiple rounds of clarification, information extraction, reading comprehension, text summarization, question generation, accounting calculations, tax law retrieval, operation instructions, interpretation of laws and regulations, and case analysis warning.
[0013] In one embodiment of the present invention, in the indicator construction step, the performance indicator threshold is set based on the initial basic model or the three-way comparison model.
[0014] In one embodiment of the present invention, the performance indicator thresholds include response time, number of replies per second, number of tokens per second, maximum tokens response time, total time consumption, and number of issues processed per unit time.
[0015] In one embodiment of the present invention, in the indicator construction step, for the performance indicator threshold, the search test data is higher than the reading comprehension test data, the accounting calculation test data and the case analysis test data.
[0016] In one embodiment of the present invention, the performance analysis report is divided into an overview version performance analysis report and a precise version performance analysis report. The analysis data of the overview version performance analysis report includes total input, total output, throughput, tps, average number of tokens per second, and data obtained by comparing and analyzing various performance indicators in the execution process of the large language model with the set performance indicator threshold.
[0017] In one embodiment of the present invention, the precise version performance analysis report makes a quality judgment on the responses of the large language model based on the execution results recorded in the model testing step, including marking the responses of the large language model as follows: model rejection, model stuck, irrelevant answers, and redundant and verbose. The precise version performance analysis report analyzes data sets of different types and difficulties separately and analyzes and reports on the response quality of the large language model.
[0018] In one embodiment of the present invention, a method for judging the quality of a response of a large language model includes calculating the repetition rate of an N-gram algorithm, regular matching, and evaluating using a three-party referee model.
[0019] The present invention also provides a terminal performance evaluation device for a large language model, which includes:
[0020] The data construction module is configured to construct a multi-dimensional test data set from the business direction and the general capability direction of the large language model according to the characteristics of the large language model and the fields involved. The multi-dimensional test data set includes multiple types, and each type is further divided into data sets of different difficulties. The data set includes multiple test data;
[0021] The indicator construction module is configured to set performance indicator thresholds for multi-dimensional test data sets according to their types;
[0022] The model testing module is configured to use a multi-dimensional test data set to initiate a load request to the large language model, call and execute the groups within the large language model sequentially or randomly, and record various performance indicators and execution results during the execution process;
[0023] The performance evaluation module is configured to output a performance analysis report based on various performance indicators and execution results recorded in the model testing step.
[0024] The terminal performance evaluation method and device for a large language model provided by the present invention have the following beneficial technical effects:
[0025] (1) The present invention provides a systematic evaluation method that can comprehensively, automatically, quickly and accurately reflect the end-to-end performance of a large language model. The present invention can perform an overall performance evaluation on a large language model by subdividing the test data set and cleaning invalid test data on the basis of shielding network transmission loss (including time-consuming steps such as model combination consumption and security filtering), thereby obtaining performance data from the large language model to the terminal, and finally obtaining a performance analysis report that can reflect the end user's experience;
[0026] (2) The present invention can generate detailed performance data classified in multiple dimensions under one test process, and can analyze performance data separately according to subdivided fields, and the performance data is more accurate;
[0027] (3) The performance data produced by the present invention can provide data support for the iterative training and tuning of the large language model, thereby improving the iterative efficiency of the large language model. During the training and fine-tuning of the large language model, by paying real-time attention to user experience-related data, the tuning effect of the large language model and the quality of the training data can be inferred. In the early stage of the release of the large language model, the use of the method of the present invention can enable the provider of the large language model to have a more comprehensive and detailed understanding of the performance status of the large language model;
[0028] (4) Invalid responses from the large language model are filtered out through a variety of intelligent algorithms, thereby more accurately analyzing the effectiveness of the large language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0030] Figure 1 A schematic diagram of a terminal performance evaluation device for a large language model according to an embodiment of the present invention;
[0031] Figure 2 It is a schematic diagram of obtaining a performance analysis report according to performance indicators according to an embodiment of the present invention. DETAILED DESCRIPTION
[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0033] The terminal performance evaluation method and device for a large language model provided by the present invention are used to evaluate the performance of the large language model felt by the user when using the large language model, that is, the performance from the terminal to the model.
[0034] An embodiment of the present invention provides a terminal performance evaluation method for a large language model, which includes:
[0035] Data construction step: According to the characteristics of the large language model and the fields involved, a multi-dimensional test data set is constructed from the business direction and general capability direction of the large language model. The multi-dimensional test data set includes multiple types, and each type is further divided into data sets of different difficulty. The data set includes multiple test data;
[0036] The field involved in this embodiment is the finance and taxation field, that is, the large language model is a large model in the finance and taxation field. According to the difficulty, the data set is divided into a simple difficulty data set, a medium difficulty data set and a difficult difficulty data set. In other embodiments, it can also be divided into more (4 or more) or fewer (two) difficulty levels according to actual needs.
[0037] In this implementation, the types of multi-dimensional test data sets include table analysis, chart construction, function rewriting, multiple rounds of clarification, information extraction, reading comprehension, text summarization, question generation, accounting calculation, tax law retrieval, operation instructions, legal and regulatory interpretation, and case analysis warning. In other embodiments, when the fields involved are other fields, such as the financial field and the customer service field, the types of multi-dimensional test data sets can be appropriately adjusted to match the actual usage requirements of the field.
[0038] Indicator construction step, for multi-dimensional test data sets, performance indicator thresholds are set according to types;
[0039] In this embodiment, the performance index threshold is set based on the initial basic model or the three-party comparison model (such as ChatGPT, Qianfan large model, etc.). Using the initial basic model as a benchmark can reflect the impact on the model performance in the fine-tuning iteration of the large language model. Using the three-party comparison model as a benchmark can horizontally evaluate the performance level of the tested large language model.
[0040] The performance index thresholds include response time, number of replies per second, number of tokens per second, maximum tokens response time, total time consumption, and number of issues processed per unit time. In other embodiments, other performance index thresholds related to the execution process or result can be selected or several of the above multiple performance index thresholds can be selected for use, depending on actual needs.
[0041] Generally speaking, for the performance indicator threshold, search test data is higher than reading comprehension test data, accounting calculation test data and case analysis test data. For example, in this embodiment, the response time threshold of search test data such as tax law retrieval and interpretation of laws and regulations is set to 300ms, while for reading comprehension test data, accounting calculation test data and case analysis test data, the response time threshold is set to 500ms.
[0042] Model testing step: Use a multi-dimensional test data set to initiate a load request to the large language model, call and execute the groups within the large language model sequentially or randomly, and record various performance indicators and execution results during the execution process;
[0043] Each piece of test data includes a test question, and the execution result generally refers to the response of the large language model to the test question, that is, the answer to the test question.
[0044] In this embodiment, while recording various performance indicators during the execution process, the overall resource usage data of the large language model during the operation process is recorded on the large language model side.
[0045] In the model testing step, when the large language model includes a security model and / or an intelligent analysis model, monitors can be further added before and after the security model and / or the intelligent analysis model to record the overall resource usage data within the security model and / or the intelligent analysis model.
[0046] The performance evaluation step outputs a performance analysis report based on the performance indicators and execution results recorded in the model testing step.
[0047] In this embodiment, the performance analysis report is divided into an overview version performance analysis report and a precise version performance analysis report. The analysis data of the overview version performance analysis report includes total input, total output, throughput, tps, average number of tokens per second, and data obtained by comparing and analyzing various performance indicators during the execution of the large language model with the set performance indicator thresholds.
[0048] The precise version performance analysis report makes a quality judgment on the large language model's responses based on the execution results recorded in the model testing steps, including marking the responses of the large language model as follows: model rejection, model stuck, irrelevant answers, and redundant and wordy. The precise version performance analysis report analyzes data sets of different types and difficulties separately and analyzes and reports on the quality of the large language model's responses.
[0049] Methods for judging the quality of responses from large language models include calculating the repetition rate of the N-gram algorithm, regular matching, and using a three-party referee model for evaluation. Figure 2 This is a schematic diagram of obtaining a performance analysis report based on performance indicators according to an embodiment of the present invention, in which the above three methods are used to judge the quality of the responses of the large language model.
[0050] Calculate the N-gram algorithm repetition rate. Calculate the N-gram repetition rate of the large language model's replies to determine whether the large language model's replies are redundant or the model is stuck.
[0051] According to the execution results recorded in the model testing step, each response of the large language model is evaluated. The specific method is as follows: the response content of the tested large language model is segmented according to punctuation marks, including: comma, period, colon, semicolon (model jamming and redundant verbosity often occur); then the segmented phrases are grouped (usually into 2 groups); finally, the number of n-grams shared between the groups is compared, and the obtained n-gram sequence is counted in the operation monitoring data row.
[0052] Regular matching is to build a rejection answer template for the large language model, match the response of the large language model, and identify whether the model response is a rejection answer.
[0053] According to the execution results recorded in the model testing step, perform regular matching evaluation on each reply of the large language model. The regular expressions include but are not limited to: "(?: I can't answer your question|The service is temporarily unavailable)", "(?: Service|System|Function).*?(?: Unavailable|Error|Failure|Suspended)", "Sorry, .*", and record the number of successful matches.
[0054] Use a three-party referee model to evaluate the responses of the large language model to determine whether the large language model is randomly assembled or answers irrelevant questions, and to determine the relevance of the responses of the large language model. Using a three-party referee model for evaluation is relatively time-consuming, and is generally not required when the performance of the large language model significantly exceeds the performance indicator threshold. The specific method of the referee model evaluation method is: submit each response of the large language model to a three-party referee model (ChatGPT or Wenxinyiyan), and use prompt words to let the referee model determine the relevance of the response of the large language model to the questions in the test question set. In order to improve performance, the accuracy of the model's response is not determined. Record the evaluation results of the referee model.
[0055] Through the above three methods, the quality of the replies of the large language model is judged, and the effective replies are screened out to issue an accurate version performance analysis report. From the above analysis, it can be seen that the overview version performance analysis report and the accurate version performance analysis report perform performance analysis on the large language model from the type dimension, difficulty dimension, and reply quality dimension, so that the performance distribution of each aspect of the model can be quickly, clearly, and in detail. The overall data indicators obtained by the present invention include at least: the number of tokens processed per second (input), the average number of tokens replied per second (output); the detailed data statistical dimensions specifically include: business classification (such as information extraction ability tokens processing ability, summary ability tokens processing ability) dimension statistics, difficulty classification (simple question tokens processing ability, difficult question tokens processing ability) statistical dimension, indicator threshold trigger distribution (statistical analysis test question set that triggers the performance indicator threshold distribution). In other embodiments, the present invention can evaluate the terminal performance of the large language model from multiple dimensions, not limited to the above.
[0056] The present invention also provides a terminal performance evaluation device for a large language model. Figure 1 FIG. 1 is a schematic diagram of a terminal performance evaluation device for a large language model according to an embodiment of the present invention. Figure 1 As shown, it includes:
[0057] The data construction module is configured to construct a multi-dimensional test data set from the business direction and the general capability direction of the large language model according to the characteristics of the large language model and the fields involved. The multi-dimensional test data set includes multiple types, and each type is further divided into data sets of different difficulties. The data set includes multiple test data;
[0058] The indicator construction module is configured to set performance indicator thresholds for multi-dimensional test data sets according to their types;
[0059] The model testing module is configured to use a multi-dimensional test data set to initiate a load request to the large language model, call and execute the groups within the large language model sequentially or randomly, and record various performance indicators and execution results during the execution process;
[0060] The performance evaluation module is configured to output a performance analysis report based on various performance indicators and execution results recorded in the model testing step.
[0061] The terminal performance evaluation method and device for a large language model provided by the present invention have the following beneficial technical effects:
[0062] (1) The present invention provides a systematic evaluation method that can comprehensively, automatically, quickly and accurately reflect the end-to-end performance of a large language model. The present invention can perform an overall performance evaluation on a large language model by subdividing the test data set and cleaning invalid test data on the basis of shielding network transmission loss (including time-consuming steps such as model combination consumption and security filtering), thereby obtaining performance data from the large language model to the terminal, and finally obtaining a performance analysis report that can reflect the end user's experience;
[0063] (2) The present invention can generate detailed performance data classified in multiple dimensions under one test process, and can analyze performance data separately according to subdivided fields, and the performance data is more accurate;
[0064] (3) The performance data produced by the present invention can provide data support for the iterative training and tuning of the large language model, thereby improving the iterative efficiency of the large language model. During the training and fine-tuning of the large language model, by paying real-time attention to user experience-related data, the tuning effect of the large language model and the quality of the training data can be inferred. In the early stage of the release of the large language model, the use of the method of the present invention can enable the provider of the large language model to have a more comprehensive and detailed understanding of the performance status of the large language model;
[0065] (4) Invalid responses from the large language model are filtered out through a variety of intelligent algorithms, thereby more accurately analyzing the effectiveness of the large language model.
[0066] Those skilled in the art can understand that the accompanying drawings are only schematic diagrams of an embodiment, and the modules or processes in the accompanying drawings are not necessarily required to implement the present invention.
[0067] Those skilled in the art can understand that the modules in the device in the embodiment can be distributed in the device in the embodiment according to the description of the embodiment, or can be changed accordingly and located in one or more devices different from the embodiment. The modules in the above embodiment can be combined into one module, or can be further divided into multiple sub-modules.
[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A terminal performance evaluation method for a large language model, characterized in that: include: Data construction step: According to the characteristics of the large language model and the fields involved, a multi-dimensional test data set is constructed from the business direction and general capability direction of the large language model. The multi-dimensional test data set includes multiple types, and each type is further divided into data sets of different difficulty. The data set includes multiple test data; Indicator construction step, for multi-dimensional test data sets, performance indicator thresholds are set according to types; Model testing step: Use a multi-dimensional test data set to initiate a load request to the large language model, call and execute the groups within the large language model sequentially or randomly, and record various performance indicators and execution results during the execution process; The performance evaluation step outputs a performance analysis report based on the performance indicators and execution results recorded in the model testing step.
2. The terminal performance evaluation method for a large language model according to claim 1, characterized in that: According to the difficulty, the datasets are divided into simple difficulty datasets, medium difficulty datasets and difficult difficulty datasets.
3. The terminal performance evaluation method for a large language model according to claim 2, characterized in that: In the data construction step, the areas involved are finance and taxation.
4. The terminal performance evaluation method for a large language model according to claim 3, characterized in that: The types of multi-dimensional test data sets include table analysis, chart construction, function rewriting, multiple rounds of clarification, information extraction, reading comprehension, text summarization, question generation, accounting calculations, tax law retrieval, operation instructions, interpretation of laws and regulations, and case analysis warning.
5. The terminal performance evaluation method for a large language model according to claim 1, characterized in that: In the indicator construction step, the performance indicator threshold is set based on the initial basic model or the three-way comparison model.
6. The terminal performance evaluation method for a large language model according to claim 1, characterized in that: Performance indicator thresholds include response time, number of replies per second, number of tokens per second, maximum tokens response time, total time consumed, and number of issues processed per unit time.
7. The terminal performance evaluation method for a large language model according to claim 1, characterized in that: In the indicator construction step, for the performance indicator threshold, the search test data is higher than the reading comprehension test data, accounting calculation test data, and case analysis test data.
8. The terminal performance evaluation method for a large language model according to claim 1, characterized in that: The performance analysis report is divided into an overview version performance analysis report and a precise version performance analysis report. The analysis data of the overview version performance analysis report includes total input, total output, throughput, tps, average number of tokens per second, and data obtained by comparing and analyzing various performance indicators during the execution of the large language model with the set performance indicator thresholds.
9. The terminal performance evaluation method for a large language model according to claim 8, characterized in that: The precise version performance analysis report makes a quality judgment on the large language model's responses based on the execution results recorded in the model testing steps, including marking the responses of the large language model as follows: model rejection, model stuck, irrelevant answers, and redundant and wordy. The precise version performance analysis report analyzes data sets of different types and difficulties separately and analyzes and reports on the quality of the large language model's responses.
10. The terminal performance evaluation method for a large language model according to claim 9, characterized in that: Methods for judging the quality of responses from large language models include calculating the repetition rate of the N-gram algorithm, regular matching, and evaluating using a three-party referee model.
11. A terminal performance evaluation device for a large language model, characterized in that: include: The data construction module is configured to construct a multi-dimensional test data set from the business direction and the general capability direction of the large language model according to the characteristics of the large language model and the fields involved. The multi-dimensional test data set includes multiple types, and each type is further divided into data sets of different difficulties. The data set includes multiple test data; The indicator construction module is configured to set performance indicator thresholds for multi-dimensional test data sets according to their types; The model testing module is configured to use a multi-dimensional test data set to initiate a load request to the large language model, call and execute the groups within the large language model sequentially or randomly, and record various performance indicators and execution results during the execution process; The performance evaluation module is configured to output a performance analysis report based on various performance indicators and execution results recorded in the model testing step.
Citation Information
Patent Citations
System load pressure testing method for large language model
CN118035104A
Performance evaluation method and device for abstract generation model, and computer equipment
CN118070775A