Large language model reasoning throughput testing method and device and program product

By using precision evaluation data sets and fitting models of request mean and throughput in large language model inference throughput testing, the problem of unreliable measurement results in existing test methods is solved, and more accurate and reliable throughput measurements are achieved.

CN120045896APending Publication Date: 2025-05-27BEIJING INBO DIGITAL TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510058818.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing large language model inference throughput testing methods have the problem that measurement results are unreliable, especially because the output length is not specified and random input may lead to high cache hit rate, affecting the accuracy of measurement results.

Method used

Throughput measurement is performed using the data set of accuracy evaluation, the information of each request is recorded, the request mean value in the pre-filling and decoding stages is calculated, and a function of request mean value and throughput is established, and a throughput prediction model is generated to reduce the deviation between inference practice and measurement throughput.

Benefits of technology

By using a real accuracy evaluation data set for measurement, we ensure that the output length meets the specified requirements, reducing the deviation of the measurement results, and solving the problem that the length does not meet the throughput measurement requirements through the fitting algorithm, improving the credibility of the measurement results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045896A_ABST
    Figure CN120045896A_ABST
Patent Text Reader

Abstract

The invention relates to a large language model reasoning throughput testing method and device and a program product. The method comprises the steps that reasoning evaluation is conducted on a to-be-tested framework based on a precision evaluation data set; recording each piece of request information in the reasoning evaluation; respectively calculating a request mean value in a pre-filling stage and a decoding stage at each moment; and establishing a function of the request mean value and the throughput, and generating a throughput prediction model. Reasoning evaluation is carried out by adopting a real precision evaluation data set, and meanwhile, fitting is carried out through the relationship between the throughput and the input and output length, so that the deviation between reasoning practicability and throughput measurement is reduced. Precision evaluation and throughput measurement are completed at the same time, and the additional calculation amount is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of model testing, and in particular, to a method, apparatus, and program product for testing the inference throughput of large language models. Background Art

[0002] In the past two years, with the introduction and application of large language models such as OpenAI's ChatGPT, a large number of large language model applications have emerged at home and abroad. After large language models enter the application stage, due to the need to process a large number of user service requests, the computing power requirements for inference will exceed those for training. Therefore, a series of frameworks for accelerating large language model inference, such as SGLang and vLLM, have emerged. At the same time, many new frameworks are also constantly being developed. And throughput determines the computing power utilization efficiency of various frameworks. Therefore, it is necessary to have a relatively economical and accurate method to measure the inference throughput of these frameworks.

[0003] Currently, there are already some tool software for measuring the inference throughput of large language models. Existing measurements generally specify the input length and do not limit the output length. If the specified length has not been reached when the output is completed, the signal indicating the end of the output can only be ignored and generation continues. Typically, llmperf developed by ray and the bench_serving module attached to SGLang. Since the output length of large language models is generally not specified and automatically stops when the model infers to the end character. And the length specified by llmperf often cannot be reached, so its setting of the output length is not well adhered to. The random input of the bench_serving module of SGLang is generated by random numbers. If the random number seed is not re-specified, it may result in no change in the tokens in each input request, leading to a particularly high cache hit rate and thus unreliable measurement results. Summary of the Invention

[0004] In view of this, this application proposes a method, apparatus, and program product for testing the inference throughput of large language models. By directly using the dataset for accuracy evaluation to measure throughput, the deviation between inference utility and measured throughput can be reduced.

[0005] According to one aspect of this application, a method for testing the inference throughput of large language models is provided, including: performing inference evaluation on the framework to be tested based on the dataset for accuracy evaluation;

[0006] Recording each request information in the inference evaluation;

[0007] Calculating the average value of requests in the prefill stage and the decoding stage at each moment respectively;

[0008] Establishing a function between the average value of requests and throughput to generate a throughput prediction model.

[0009] In a possible implementation, when performing inference evaluation on the to-be-tested framework based on the dataset for accuracy evaluation, a dataset that covers the target input length and output length with the maximum and minimum input lengths and the maximum and minimum output lengths is selected as the dataset for accuracy evaluation used in this method.

[0010] In a possible implementation, when recording each request message in the inference evaluation, the request message includes: the sending time, start generation time, completion time of each request during the inference evaluation, and the throughput of the server at each moment;

[0011] Among them, the interval from the sending of each request to the start of generation is the pre-filling stage, and the interval from the start of generation to the completion is the decoding stage.

[0012] In a possible implementation, when establishing the function of the request mean and the throughput according to the following steps:

[0013] For each moment, the request mean in the input stage is denoted as L i , the request mean in the output stage is denoted as L o , and the throughput at the current moment is denoted as T;

[0014] Construct a 7-dimensional input The output is 1 / T, and the function is obtained by fitting:

[0015]

[0016] In a possible implementation, when obtaining the dataset for accuracy evaluation, it further includes:

[0017] When the input lengths and output lengths of the selected dataset are concentratedly distributed, sampling is performed;

[0018] The dataset samples obtained by sampling are evenly distributed on the two-dimensional matrix of the input length and the output length;

[0019] Moreover, each sample is collected at most once during sampling.

[0020] In a possible implementation, using the function as a prediction model for throughput, record the maximum and minimum values of the function input and output;

[0021] When given L i , L o is within the range of the maximum and minimum values of the function input and output, the throughput parameter value is obtained according to the prediction model

[0022] According to another aspect of the present application, there is provided an apparatus for testing the inference throughput of a large language model, including an acquisition module, a processing module, and a prediction module;

[0023] The acquisition module is configured to perform inference evaluation on the framework to be tested based on the dataset for accuracy evaluation;

[0024] The processing module is configured to record each request information in the inference evaluation; calculate the average value of requests in the pre-fill stage and the decoding stage at each moment respectively;

[0025] The prediction module is configured to establish a function between the average value of requests and the throughput, and generate a throughput prediction model.

[0026] According to another aspect of the present application, there is provided a device for testing the inference throughput of a large language model, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the above method.

[0027] According to another aspect of the present application, there is provided a non-volatile computer-readable storage medium, on which computer program instructions are stored, wherein, when the computer program instructions are executed by a processor, the above method is implemented.

[0028] According to another aspect of the present application, there is provided a computer program product, including a computer program, wherein, when the computer program is executed by a processor, the steps of the above method are implemented.

[0029] The present application directly runs inference evaluation using the dataset for accuracy evaluation, and at the same time post-processes the throughput data during inference evaluation to obtain the estimated throughput. Using the real dataset for accuracy evaluation can ensure that the test results will not be biased due to incorrect or non-random generated random numbers, or the length not meeting the specified requirements. At the same time, through the algorithm of fitting the relationship between throughput and input / output length, the problem that the length in the real dataset does not meet the requirements of throughput measurement is solved.

[0030] According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present application will become clear. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The drawings included in the specification and constituting a part of the specification show the exemplary embodiments, features and aspects of the present application together with the specification, and are used to explain the principles of the present application.

[0032] Figure 1 A flowchart showing the method for testing the inference throughput of a large language model according to an embodiment of the present application;

[0033] Figure 2A device block diagram of a large language model inference throughput test according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0034] Various exemplary embodiments, features and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0035] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0036] In addition, in order to better illustrate the present application, numerous specific details are provided in the following specific embodiments. It should be understood by those skilled in the art that the present application can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the subject matter of the present application.

[0037] This application is applicable to the inference throughput test of large language models, and uses a real precision evaluation data set for inference evaluation. Since the precision evaluation generally has more data volume or contains more types of data sets, there is more room for choice. At the same time, it is also more practical, which can reduce the deviation between reasoning practicality and measured throughput. It ensures that there will be no deviation in the results during the test due to the fact that the generated random number errors are not completely random, or the length does not meet the specified requirements. At the same time, the algorithm that fits the relationship between throughput and input and output length solves the problem that the length in the real data set does not meet the throughput measurement requirements.

[0038] Example 1

[0039] Figure 1 A flow chart of a method for testing the large language model inference throughput according to an embodiment of the present application is shown. Figure 1 As shown, the method includes:

[0040] Step S100, perform reasoning evaluation on the framework to be tested based on the data set of precision evaluation. Reasoning mainly refers to the process of inputting instructions or requests to a large language model and giving its output. Generally, after the training of a large language model is completed, dozens of data sets will be used to evaluate its accuracy in various languages ​​and various tasks. The throughput is measured directly using the data set of precision evaluation. The data set of precision evaluation is designed according to the actual usage scenario. Therefore, if a certain solution is optimized in a targeted manner, then the throughput that is closest to the actual scenario effect can also be measured.

[0041] Step S200, record each request information in the inference evaluation. The request information includes: the issuance time, start generation time, completion time of each request during the inference evaluation, and the throughput of the server at each moment;

[0042] Step S300, respectively calculates the mean of requests in the pre-filling stage and the decoding stage at each moment. The pre-filling stage is also the input stage of reasoning, which means that each reasoning request of the large language model has user input text or instructions. This part can be calculated as a whole and is computationally intensive. After calculating this part, the calculation results will be cached for use in the output stage. This cache is generally called a KV cache. The decoding stage is the output stage of reasoning, which refers to the process of the large language model outputting tokens. Each output needs to use the information of the previous token, but for the output, the previous token is not given. The amount of calculation of this part is not large, but it is necessary to repeatedly read the cache of the previous calculation results from the memory, and the main consumption is memory reading and writing.

[0043] Step S400: Establish a function of request mean and throughput to generate a throughput prediction model. The throughput prediction model is an algorithm that fits the relationship between throughput and input and output lengths. When the input and output lengths are given later, the parameter value corresponding to the throughput can be directly fitted.

[0044] In one possible implementation, after the conventional large language model is trained, dozens of data sets will be used to evaluate the accuracy of the model in various languages ​​and tasks, and the inference evaluation of the framework to be tested will be performed based on the data set of the accuracy evaluation. The data set of the accuracy evaluation is designed according to the actual usage scenario, and the measured throughput is also the closest to the real scenario effect.

[0045] The input and output lengths of the accuracy test data set cannot be specified, and when measuring throughput, the main influencing factor is the input and output length. Therefore, the method of large language model inference throughput test proposed in this application is to post-process the actual running data to obtain the throughput solution when the input and output length is given.

[0046] Furthermore, when selecting a precision evaluation dataset, select a dataset whose maximum and minimum input lengths and maximum and minimum output lengths cover the target input length and output length. If the target input and output lengths are not specified, select the dataset with the widest coverage so that extrapolation does not occur in the subsequent fitting process, thereby avoiding increasing the estimation error. The input length and output length refer to the number of tokens that can be input and output during the model inference phase.

[0047] If the input lengths and output lengths of the selected dataset are concentrated in a relatively small area, sampling can be performed to ensure that the samples of the dataset are evenly distributed on the two-dimensional matrix of the input length and output. During sampling, ensure that each sample is collected at most once to avoid inaccurate measurements due to caching.

[0048] Based on the dataset for accuracy evaluation, build a large language model service for the framework to be tested (such as vLLM, sgLang, HuggingfaceTGI), and perform inference evaluation according to the logic of accuracy evaluation. Record the data of accuracy evaluation to verify that there are no obvious errors in the implementation of the framework to be tested or that the optimization does not cause accuracy loss, which can avoid inaccurate measurement results due to implementation errors or other links.

[0049] At the same time, record the sending time, start generation time, completion time of each request in the inference evaluation, as well as the throughput of the server at each moment to obtain the inference evaluation dataset.

[0050] For each moment, calculate the mean values of the requests in the input stage and output stage at the current moment, denoted as L i ,L o , and record the throughput statistics at the current moment as T. The method proposed in this application can complete accuracy evaluation and throughput measurement simultaneously. Since accuracy evaluation is a necessary task when training or evaluating the accuracy of a model, no additional computational effort is introduced.

[0051] Establish the functional relationship between 1 / T and L i ,L o . Since according to the calculation complexity inference, the throughput is approximately inversely proportional to the -2 power of L i and L o . Therefore, a 7-dimensional input can be constructed, which are respectively with the output being 1 / T. Use the least squares linear fitting to obtain the final function:

[0052]

[0053] Use this function as the throughput prediction model, and record the function as well as the maximum and minimum values of the input and output. Given the input L i ,L o , as long as L i ,L o is within the maximum and minimum values of the prediction model, the corresponding 1 / T can be fitted, and thus the throughput T can be calculated.

[0054] Furthermore, according to another aspect of this application, a device for testing the inference throughput of a large language model is provided, including an acquisition module, a processing module, and a prediction module, which can correspond to the method described above.

[0055] An acquisition module, configured to perform an inference evaluation on a to-be-tested framework based on a dataset for accuracy evaluation;

[0056] A processing module, configured to record each request message in the inference evaluation; calculate the request mean values in the pre-filling stage and the decoding stage at each moment respectively;

[0057] A prediction module, configured to establish a function between the request mean value and the throughput, and generate a throughput prediction model.

[0058] Furthermore, according to another aspect of the present application, there is also provided a device for testing the inference throughput of a large language model. As Figure 2 shown, it includes a processor 210 and a memory 220 for storing executable instructions that can be executed by the processor 210. Among them, the processor 210 is configured to implement the method for testing the inference throughput of a large language model described in any one of the foregoing when executing the executable instructions.

[0059] Here, it should be noted that the number of processors 210 can be one or more. At the same time, in the device for testing the inference throughput of a large language model in the embodiment of the present application, an input device 230 and an output device 240 may also be included. Among them, the processor 210, the memory 220, the input device 230, and the output device 240 can be connected through a bus or in other ways, which is not specifically limited here.

[0060] The memory 220, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and various modules, such as: programs or modules corresponding to the method for testing the inference throughput of a large language model in the embodiment of the present application. The processor 210 executes various functional applications and data processing of the device for testing the inference throughput of a large language model by running the software programs or modules stored in the memory 220.

[0061] The input device 230 can be used to receive input numbers or signals. Among them, the signal can be a key signal related to user settings and function control of the device / terminal / server. The output device 240 may include a display device such as a display screen.

[0062] According to another aspect of the present application, there is also provided a non-volatile computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by the processor 210, the method for testing the inference throughput of a large language model described in any one of the foregoing is implemented.

[0063] According to another aspect of the present application, there is also provided a computer program product, including a computer program, which when executed by a processor, implements the steps of any of the above-mentioned methods. Its specific implementation manner is consistent with the method implementation manner and can achieve the same beneficial effects, so it will not be elaborated here.

[0064] In summary, the present application directly runs the inference evaluation using the accuracy evaluation data set, and at the same time post-processes the throughput data during the inference evaluation to obtain the estimated throughput. The accuracy evaluation and the throughput measurement are completed simultaneously. Since the accuracy evaluation is also a necessary task when training or evaluating the model accuracy, no additional computational effort is introduced. At the same time, using the real accuracy evaluation data set for evaluation can ensure that there will be no deviation in the results due to the generated random numbers not being completely random or the length not meeting the specified requirements. At the same time, through the algorithm of fitting the relationship between the throughput and the input and output lengths, the problem that the lengths in the real data set do not meet the requirements of the throughput measurement is solved. The present application also ensures that there is no obvious difference in the accuracy on the test data set compared with the standard implementation, which can avoid inaccurate measurement results caused by implementation errors or other links. At the same time, if there is a loss of accuracy in some optimizations of the framework to be tested, it can also be recorded, which is convenient for users to make a trade-off between accuracy and throughput when selecting.

[0065] The various embodiments of the present application have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary technical personnel in the technical field to understand the embodiments disclosed herein.

Claims

1. A method for testing the inference throughput of a large language model, characterized in that: include: Perform reasoning evaluation on the framework to be tested based on the accuracy evaluation dataset; Recording each request information in the reasoning evaluation; Calculate the mean of requests in the pre-filling phase and the decoding phase at each moment respectively; A function of the request mean and throughput is established to generate a throughput prediction model.

2. The method for testing the large language model inference throughput according to claim 1, characterized in that: A data set whose maximum and minimum input lengths and maximum and minimum output lengths cover the target input length and output length is selected as the data set for accuracy evaluation used in this method.

3. The method for testing the large language model inference throughput according to claim 1, characterized in that: When recording each request information in the reasoning evaluation, the request information includes: The issuance time, start generation time, completion time of each request and the throughput of the server at each moment during the inference evaluation; The interval between each request being issued and the start of generation is the pre-filling phase, and the interval from the start of generation to completion is the decoding phase.

4. The method for testing the large language model inference throughput according to claim 1, characterized in that: Follow these steps to build a function of the request mean and throughput: At each moment, the average number of requests in the input phase is recorded as L i , the average value of requests in the output stage is recorded as L o , the throughput at the current moment is recorded as T; Constructing 7-dimensional input The output is 1 / T, and the function is obtained by fitting:

5. The method for testing the large language model inference throughput according to claim 2, characterized in that: When obtaining the data set for the accuracy evaluation, sampling is performed when the input length and output length of the selected data set are concentratedly distributed; The sampled data set samples are evenly distributed on the two-dimensional matrix of input length and output length; Moreover, each sample is collected at most once during sampling.

6. The method for testing the large language model inference throughput according to claim 4, characterized in that: Using the function as a throughput prediction model, recording the maximum and minimum values ​​of the function input and output; When given L i , L o The throughput parameter value is obtained according to the prediction model within the range of the maximum value and the minimum value of the function input and output.

7. A device for testing the throughput of large language model inference, characterized in that: Including acquisition module, processing module and prediction module; The acquisition module is configured to perform reasoning evaluation on the framework to be tested based on the accuracy evaluation data set; The processing module is configured to record each request information in the reasoning evaluation; respectively calculate the mean of the requests in the pre-filling stage and the decoding stage at each moment; The prediction module is configured to establish a function of the request mean and the throughput to generate a throughput prediction model.

8. A device for testing the inference throughput of a large language model, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the method described in any one of claims 1 to 6 when executing the executable instructions.

9. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Large model memory management method and device, electronic equipment and readable storage medium

    CN120353603A