Large language model reasoning performance evaluation and optimization method, electronic equipment and storage medium

By dynamically generating concurrent user requests and real-time monitoring of performance indicators, the problem of inconsistent evaluation results of large language models in different hardware environments is solved, precise evaluation and optimization in high concurrency scenarios are achieved, and user experience and system adaptability are improved.

CN120297409APending Publication Date: 2025-07-11XINZHIHUIXIANG TECHNOLOGY (TIANJIN) CO LTD

Patent Information

Application Number
CN202510359521.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing technology lacks a unified, accurate and real-time method to evaluate the inference performance of large language models (LLMs) in different hardware environments, resulting in large differences in evaluation results in different environments and application scenarios, making it difficult to adapt to a dynamically changing production environment.

Method used

A large language model inference performance evaluation and optimization method is proposed, including initializing the test environment, dynamically generating concurrent user requests, collecting performance data, real-time monitoring and reporting, adjusting the number of users through an adaptive gain algorithm, simulating different loads, and calculating performance indicators such as the number of request processing per second, response delay, first word element time and inter-word element delay, etc.

Benefits of technology

It realizes comprehensive and accurate evaluation of LLM in high concurrency scenarios, can monitor and optimize system resource configuration in real time, improve user experience and system adaptability, reduce manual experience dependence, and provide automated performance evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297409A_ABST
    Figure CN120297409A_ABST
Patent Text Reader

Abstract

The invention discloses a big language model reasoning performance evaluation and optimization method, electronic equipment and a storage medium. The method comprises the following steps: initializing a test environment; dynamically generating concurrent user requests; collecting performance data based on the concurrent user requests; performing real-time monitoring and reporting based on the performance data; and generating a performance evaluation report based on the performance data. Through the above steps, the method can comprehensively evaluate the reasoning performance of the LLM in a high-concurrency scene, and ensures the optimal performance of the system in different hardware environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of software technology, and particularly relates to a method for evaluating and optimizing the inference performance of large language models, an electronic device, and a storage medium. Background Art

[0002] With the continuous update and iteration of large language models (LLMs), and the emergence of more and more new LLMs, it has become particularly important to measure the performance of these models in different hardware environments. Especially on different series of GPU cards such as NVIDIA, Ascend, and Tianyi, the performance of LLMs may vary. By measuring the concurrency of LLMs and the number of input / output tokens per second (Throughput), the actual performance of these models in different hardware environments can be comprehensively evaluated and compared, thereby guiding the selection and optimization of models. This has important reference value and practical significance for enterprises and developers deploying LLMs on specific hardware platforms.

[0003] Although the methods for evaluating the inference performance of LLMs have been relatively mature, there is still a lack of a unified measurement standard in actual applications. Currently, the industry usually measures the performance of LLMs through the following key performance indicators:

[0004] Throughput:

[0005] Throughput is a key indicator for measuring the resource utilization rate and system cost of an LLM service system, indicating the number of requests processed by the system per unit time. To increase throughput, the method of increasing the batch size is usually adopted, changing the serial processing of user requests to parallel processing. However, increasing the batch size may increase the latency of each user to a certain extent, affecting the user experience.

[0006] Latency:

[0007] Latency refers to the time required for a user to receive a complete response from the moment the request is sent. In actual applications, the smaller the latency, the smoother the user experience. Generally, when the latency is no more than 50 ms / token, users can feel a relatively smooth usage experience. Therefore, optimizing latency is crucial for improving user satisfaction.

[0008] Requests per minute (RPS):

[0009] Requests per minute reflects the system's ability to process concurrent requests, especially important when dealing with inputs from multiple users or batch inference workloads. Reasonably adjusting the RPS can ensure the stability of the system in high-concurrency scenarios.

[0010] Time to first token (TTFT):

[0011] In a streaming application, the Time to First Token (TTFT) refers to the time required for the LLM to return the first token. In addition to focusing on the average TTFT, it is also necessary to pay attention to its distribution, such as P50, P90, P95, and P99. Optimizing TTFT helps improve the user's waiting experience.

[0012] Inter-Token Latency (ITL):

[0013] Inter-Token Latency refers to the average time between consecutive output tokens. In practical applications, incorporating TTFT into the calculation of Inter-Token Latency can more comprehensively evaluate the inference performance of the LLM.

[0014] Although the above methods can help evaluate the inference performance of the LLM to a certain extent, due to the lack of a unified measurement standard, there may be significant differences in the evaluation results under different environments and application scenarios. In addition, most of these evaluation methods are offline calculations or results obtained after experiments, with poor real-time performance and difficulty in adapting to dynamic production environments.

[0015] Therefore, when facing different hardware environments, there is an urgent need for a more accurate, standardized, and real-time applicable evaluation method to help developers and enterprises better understand and optimize the inference performance of the LLM. Summary of the Invention

[0016] The object of the present invention is to provide a method for evaluating and optimizing the inference performance of a large language model that can identify interference sources, thereby avoiding strong electromagnetic interference and providing a stable and reliable operating environment for the WAPI wireless environment.

[0017] To achieve the above object, the present invention proposes a method for evaluating and optimizing the inference performance of a large language model, the method comprising: initializing a test environment; dynamically generating concurrent user requests; collecting performance data based on the concurrent user requests; performing real-time monitoring and reporting based on the performance data; generating a performance evaluation report based on the performance data.

[0018] In an alternative embodiment, initializing the test environment specifically includes: configuring test parameters, the test parameters including the URL of the server, the session time, and the upper limit of the number of users;. Preparing test data and parsing the data set to obtain prompt information for testing; initializing a concurrent user generator and a performance metric collector; performing a network latency test.

[0019] In an alternative embodiment, performing a network latency test specifically includes: initiating an asynchronous HTTP GET request using aiohttp.ClientSession; performing multiple ping tests to obtain the average response time; if ping correction is enabled, subtracting a fixed value from the average latency.

[0020] In an alternative embodiment, concurrent user requests are dynamically generated, specifically including: using the initialized concurrent user generator to dynamically generate concurrent user requests; supporting the adjustment of the number of users based on the adaptive gain algorithm to simulate different loads; and independently generating each user request and sending it to the server asynchronously.

[0021] In an alternative embodiment, performance data is collected based on the concurrent user requests, specifically including:

[0022] Set a time window; within the time window, use the initialized performance metric collector to collect performance metrics, where the performance metrics include: maximum number of users, active number of users, number of requests processed per second, response latency, time to first token, inter-token latency, and number of input / output tokens per second.

[0023] In an alternative embodiment, within the time window, using the initialized performance metric collector to collect performance metrics specifically includes:

[0024] Calculate the number of requests processed per second, and the formula is as follows:

[0025]

[0026] Calculate the response latency, specifically including:

[0027] Calculate the average response latency, and the formula is as follows:

[0028]

[0029] where Response_Time i is the response time of the i-th request, and N is the total number of requests within the time window;

[0030] Calculate the response latency distribution, and the formula is as follows:

[0031] Response latency distribution = the sorted value of the response time at the x-th percentile;

[0032] Calculation of time to first token (TTFT) and inter-token latency (ITL):

[0033] Calculate the time to first token, and the formula is as follows:

[0034]

[0035] where First_Token_Time i is the time when the i-th request receives the first token, and N is the total number of requests within a specific time window;

[0036] Calculate the inter-token latency, and the formula is as follows:

[0037]

[0038] Among them, Token_Time i is the output time of the i-th token;

[0039] Calculate the number of tokens output per second, and the formula is as follows:

[0040]

[0041] Among them, the total number of tokens refers to the sum of tokens in all requests processed by the system within a given time window;

[0042] Calculate the number of tokens output per second for a single request, and the formula is as follows:

[0043]

[0044] In an alternative embodiment, real-time monitoring and reporting are performed based on the performance data, specifically including: regularly summarizing and reporting current performance metrics to obtain real-time monitoring data; the reporting time window can be dynamically adjusted; and the number of concurrent users is dynamically adjusted according to the real-time monitoring data.

[0045] In an alternative embodiment, the content of the performance report includes: detailed statistics of the average throughput, response latency, number of requests processed per second, first token time, token inter-arrival delay, and number of input / output tokens per second under different time windows.

[0046] The present invention also provides an electronic device, including: at least one processor; a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any one of the large language model inference performance evaluation and optimization methods.

[0047] The present invention also provides a computer storage medium storing a computer program, and when the computer program is executed by a processor, it implements any one of the large language model inference performance evaluation and optimization methods.

[0048] The beneficial effects of the present invention are as follows:

[0049] 1. It can comprehensively evaluate the inference performance of the LLM in a high-concurrency scenario, is more accurate, standardized, and can be applied in real time, and can ensure the optimal performance of the system in different hardware environments.

[0050] 2. Based on theoretical calculations and dynamic adjustments, it can more accurately evaluate the concurrent processing ability of the system.

[0051] 3. Real-time performance: The dynamic adjustment algorithm can monitor and adjust the utilization of system resources in real time, improving adaptability in the production environment.

[0052] 4. Automation: It provides systematic calculation methods and functions, reduces the dependence on manual experience, and can automatically complete complex performance evaluations. Description of the Drawings

[0053] Figure 1 It is a flowchart of a method for evaluating and optimizing the inference performance of a large language model provided by an embodiment of the present invention;

[0054] Figure 2 It is a block diagram of the composition of an electronic device provided by an embodiment of the present invention.

[0055] Description of the Reference Numerals:

[0056] 110. Processor; 120. Memory. Detailed Embodiments

[0057] The present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0058] The coverage area of the transmission corridor is getting wider and wider, and the transmission lines inevitably pass through various icing areas with complex terrains and cold climates. In recent years, with the increasingly warming global climate, various extreme disasters have occurred frequently, and transmission tower collapse accidents in plateau areas have occurred from time to time. The main reason is that the transmission tower structures designed under the existing specifications can no longer effectively cope with the challenges of the new environment. There is an urgent need to propose a set of transmission tower standards that meet the strength requirements in the new environment, check the structural strength of the existing transmission towers, and thus provide guidance for subsequent upgrades and transformations.

[0059] As Figure 1 shown, according to an embodiment of the present invention, on the one hand, a method for evaluating and optimizing the inference performance of a large language model is provided, including the following steps:

[0060] Step S101: Initialize the test environment;

[0061] Step S103: Dynamically generate concurrent user requests;

[0062] Step S105: Collect performance data based on the concurrent user requests;

[0063] Step S107: Perform real-time monitoring and reporting based on the performance data;

[0064] Step S109: Generate a performance evaluation report based on the performance data.

[0065] In this embodiment, from the initialization of the test environment to the collection of performance data, real-time monitoring and reporting, and then to the generation of the final performance evaluation report, the whole process of performance evaluation is covered, which can effectively improve the inference performance of the large language model, optimize the system resource configuration, and improve the user experience.

[0066] Among them, before evaluating the inference performance of the large language model, it is necessary to initialize the test environment to ensure the accuracy and stability of the test. The user number adjustment strategy based on the Adaptive Increase Multiplicative Decrease (AIMD) algorithm can dynamically adjust the number of concurrent users according to the system performance, ensuring that the test covers various scenarios from low load to high load, comprehensively evaluating the model performance. The dynamically generated concurrent user requests can simulate the access behaviors of real users under different loads, making the performance evaluation closer to the actual application scenario and providing real-time data support for real-time monitoring and dynamic adjustment. The real-time monitoring function enables testers to timely understand the running state of the system, quickly discover performance bottlenecks or abnormal situations, and ensure that the test process can adjust the strategy in a timely manner according to the performance performance. Summarize the performance data in the whole test process and generate a detailed performance evaluation report, which can provide comprehensive and accurate data support for the optimization of the large language model and the selection of hardware.

[0067] Further, based on step S101, initialize the test environment, which specifically includes the following steps:

[0068] Step S1011: Configure the test parameters. The test parameters include the URL of the server, the session time, and the upper limit of the number of users. Among them, set the URL address of the server to ensure that it can correctly connect to the inference service of the large language model. When configuring the session time, define the duration of each test session. Set the upper limit of the number of users to determine the maximum number of concurrent users allowed during the test. Other relevant parameters can also be configured, such as the request timeout time, retry policy, etc.

[0069] Step S1013: Prepare the test data and parse the data set to obtain the prompt information for testing. Among them, select or generate a set of voice data for testing. These data should cover different voice features, speech rates, pitches, and contents to simulate real user scenarios. Convert the voice data into a format suitable for model input, such as a specific audio encoding format, sampling rate, etc. The prompt information can be used for requests in the benchmark test, and the content is as follows:

[0070] open("databricks-dolly-15k.jsonl","wb").write(content)

[0071] dataset=[json.loads(line)for line in content.decode().split("\n")]

[0072] for d in dataset:

[0073] d["question"] = d["context"] + d["instruction"]

[0074] d["input_tokens"] = len(encoder.encode(d["question"]))

[0075] Step S1015: Initialize the concurrent user generator and the performance metric collector. Among them, initialize the concurrent user generator (UserSpawner) for dynamically generating concurrent user requests, and initialize the performance metric collector (MetricsCollector) for collecting and calculating performance data.

[0076] Step S1017: Connect to the target server and conduct a network latency test, which can evaluate and correct the impact of the network on the inference performance evaluation of the large language model.

[0077]

[0078] Among them, based on Step S1017, conduct a network latency test, establish a connection with the target server, and conduct a preliminary ping test to evaluate the basic network latency. The specific steps are as follows:

[0079] Step S10171: Initiate an asynchronous HTTP GET request using aiohttp.ClientSession.

[0080] Step S10173: Conduct multiple ping tests to obtain the average response time, which helps to evaluate the network latency.

[0081] elf.response_head_latency_bucket[math.floor(time.time())] +=

[0082] latency - self.ping_latency]

[0083] Step S10175: If ping correction is enabled, subtract a fixed value from the average latency to correct subsequent latency measurements.

[0084] In this embodiment, the average response time between the test client and the target server is obtained through multiple ping tests, so as to understand the latency of the basic network. This is very important for judging whether the network is stable and whether the latency is within an acceptable range. If it is found that the network latency is high, it may indicate that there are bottlenecks or unstable factors in the network, such as insufficient network bandwidth, network congestion, router configuration problems, etc. This helps to detect problems in advance and take measures, such as optimizing network configuration or changing the network environment, to ensure the accuracy of the test. Through network latency testing, accurate data on the basic network latency can be obtained, so that this part of the network latency can be deducted when correcting the response latency of the model, making the performance evaluation result closer to the true inference performance of the model. By subtracting a small fixed value, such as 0.005 seconds, from the average latency, subsequent latency measurements can be corrected more precisely. This correction can reduce the deviation caused by network fluctuations or measurement errors, and improve the accuracy and reliability of performance evaluation. According to the results of network latency testing, test parameters can be reasonably adjusted, such as setting appropriate request timeout times, retry strategies, etc. If the network latency is high, it may be necessary to increase the timeout time to avoid request failures caused by network problems. In addition, network latency testing can simulate the network conditions that real users may encounter when using large language models. By understanding the impact of network latency on performance, the actual experience of users in different network environments can be better understood, so as to optimize the model and system configuration to improve the user experience. When dynamically generating concurrent user requests, the results of network latency testing can be used as a reference for load adjustment strategies. If the network latency is high, it may be necessary to appropriately reduce the number of concurrent users to avoid performance problems caused by network bottlenecks.

[0085] Further, based on step S103, concurrent user requests are dynamically generated, which specifically includes the following steps:

[0086] Step S1031: Dynamically generate concurrent user requests using the initialized concurrent user generator.

[0087] Step S1033: Support adjusting the number of users based on the adaptive increase and decrease (AIMD) algorithm to simulate different loads.

[0088] Step S1035: Each user request is independently generated and sent to the server asynchronously.

[0089] In this embodiment, the UserSpawner module is used to dynamically generate concurrent user requests. This module supports the user number adjustment strategy based on the adaptive increase and decrease (AIMD) algorithm. By increasing or decreasing the number of concurrent users, the system performance under different loads is simulated.

[0090] while True:

[0091] current_users = len(self.user_list)

[0092] if current_users == self.target_user_count:

[0093] await asyncio.sleep(0.1)

[0094] elif current_users < self.target_user_count:

[0095] self.spawn_user(model_name, encoder)

[0096] sleep_time = max(

[0097] (self.target_time - time.time())

[0098] / (self.target_user_count - current_users),

[0099] 0, )

[0101] await asyncio.sleep(sleep_time)

[0102] elif current_users > self.target_user_count:

[0103] self.user_list.pop().cancel()

[0104] sleep_time = max(

[0105] (time.time() - self.target_time)

[0106] / (current_users - self.target_user_count),

[0107] 0, )

[0109] await asyncio.sleep(sleep_time).

[0110] Among them, each user request is generated independently and sent to the server asynchronously to ensure the test accuracy in high-concurrency scenarios.

[0111] def spawn_user(self, model_name, encoder):

[0112] self.user_list.append(asyncio.create_task(self.user_loop(model_name, encoder)))。

[0113] Furthermore, in step S105, performance data is collected based on concurrent user requests, which specifically includes the following steps:

[0114] Step S1051: Set a time window;

[0115] Step S1053: Within the time window, use the initialized performance metric collector to collect performance metrics, which include: the maximum number of users, the number of active users, the number of requests processed per second, the response latency, the time to first token, the latency between tokens, and the number of input / output tokens per second.

[0116] In this embodiment, within the specified time window, the MetricsCollector module is used to collect and calculate performance metrics, including the maximum number of users, the number of active users, the number of requests processed per second (RPS), the response latency (Latency), the time to first token (TTFT), the latency between tokens (ITL), the number of input / output tokens per second, etc.

[0117] The number of active user requests in the system: By maintaining a counter (self.active_users), increment the count whenever a user starts a new request and decrement the count when the request ends.

[0118] @contextlib.contextmanager

[0119] def collect_user(self):

[0120] self.active_users += 1

[0121] yield

[0122] self.active_users -= 1。

[0123] The number of HTTP requests and the response latency: Record the start time when the request starts, and increment the count of in-progress HTTP requests by 1. After the request processing ends, decrement the count of in-progress requests by 1, update the number of responses, and record the response latency.

[0124] @contextlib.contextmanager

[0125] def collect_http_request(self):

[0126] start_time = time.time()

[0127] self.on_going_requests += 1

[0128] yield

[0129] self.on_going_requests -= 1

[0130] self.response_bucket[math.floor(time.time())] += 1

[0131] self.response_latency_bucket[math.floor(time.time())] +=

[0132] time.time() - start_time - self.ping_latency

[0133] Input and output token count:

[0134] encoder = tiktoken.encoding_for_model("gpt-3.5-turbo")

[0135] request_chunk_len = len(encoder.encode(prompt))

[0136] def collect_request_chunk(self, request_chunk_len: int):

[0137] self.request_word_bucket[math.floor(time.time())] += request_chunk_len

[0138] chunk = result.get("choices", [{}])[0].get("delta", {}).get("content")

[0139] It should be noted that there are some potential issues in the original code, such as incorrect variable naming (`start_time=time.time()` should be `start_time = time.time()`) and incomplete expressions (`self.response_latency_bucket[math.floor(time.time())]+= ` seems incomplete). This translation is based on the provided text while trying to maintain the original code structure as much as possible.def collect_response_chunk(self, chunk: str, encoder: Encoding):

[0140] self.response_word_bucket[math.floor(time.time())] += len(encoder.encode(chunk))。

[0141] Specifically, based on step S1053, within the time window, the initialized performance metric collector is used to collect performance metrics, specifically including:

[0142] Calculate the number of requests processed per second, and the formula is as follows:

[0143]

[0144] Calculate the response latency, specifically including:

[0145] Calculate the average response latency, and the formula is as follows:

[0146]

[0147] Among them, Response_Time i is the response time of the i-th request, and N is the total number of requests within the time window;

[0148] Calculate the response latency distribution, and the formula is as follows:

[0149] Response latency distribution = the sorted value of the response time at the x-th percentile;

[0150] Calculate the time to first token (TTFT) and inter-token latency (ITL):

[0151] Calculate the time to first token, and the formula is as follows:

[0152]

[0153] Among them, First_Token_Time i is the time when the i-th request receives the first token, and N is the total number of requests within a specific time window;

[0154] Calculate the inter-token latency, and the formula is as follows:

[0155]

[0156] Among them, Token_Time i is the output time of the i-th token;

[0157] Calculate the number of tokens output per second, and the formula is as follows:

[0158]

[0159] Among them, the total number of tokens refers to the sum of tokens in all requests processed by the system within a given time window;

[0160] Calculate the number of tokens output per second for a single request, and the formula is as follows:

[0161]

[0162] Furthermore, in step S107, perform real-time monitoring and reporting based on performance data, which specifically includes the following steps:

[0163] Step S1071: Regularly summarize and report the current performance metrics to obtain real-time monitoring data.

[0164] Step S1073: The reporting time window can be dynamically adjusted.

[0165] Step S1075: Dynamically adjust the number of concurrent users according to the real-time monitoring data.

[0166] Through the real-time monitoring function, the system can regularly summarize and report the current performance metrics, such as throughput, response latency, RPS, etc. The reporting time window can be dynamically adjusted according to requirements to adapt to different test scenarios. During the test, dynamically adjust the number of concurrent users according to the real-time monitoring data to ensure that the system runs under the optimal load.

[0167] Specifically, the content of the performance report includes: detailed statistics of the average throughput, response latency, number of requests processed per second, first token time, inter-token latency, and number of input / output tokens per second under different time windows. The performance report can be used to compare the performance of different hardware platforms, such as NVIDIA, Ascend, and TianShu series GPU cards, and provide data support for hardware selection and system optimization.

[0168] On the other hand, as Figure 2 shown, the present invention also provides an electronic device, including: at least one processor 110; a memory 120 communicatively connected to the at least one processor 110; wherein, the memory 120 stores instructions executable by the at least one processor 110, and the instructions are executed by the at least one processor 110 so that the at least one processor 110 can execute any one of the large language model inference performance evaluation and optimization methods.

[0169] On the other hand, the present invention also proposes a computer storage medium storing a computer program, and when the computer program is executed by a processor, it implements any one of the large language model inference performance evaluation and optimization methods.

[0170] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc. Each embodiment in this specification is described in a progressive manner. For the same or similar parts between the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiments.

[0171] The above embodiments are only examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.

Claims

1. A method for evaluating and optimizing the inference performance of a large language model, characterized in that, The method includes: Initializing the test environment; Dynamically generating concurrent user requests; Collecting performance data based on the concurrent user requests; Performing real-time monitoring and reporting based on the performance data; Generating a performance evaluation report based on the performance data.

2. The method for evaluating and optimizing the inference performance of the large language model according to claim 1, wherein Initializing the test environment specifically includes: Configuring test parameters, where the test parameters include the URL of the server, session time, and upper limit of the number of users; Preparing test data and parsing the data set to obtain prompt information for testing; Initializing a concurrent user generator and a performance metric collector; Performing a network latency test.

3. The method for evaluating and optimizing the inference performance of the large language model according to claim 2, wherein Performing a network latency test specifically includes: Initiating an asynchronous HTTP GET request using aiohttp.ClientSession; Performing multiple ping tests to obtain the average response time; If ping correction is enabled, subtract a fixed value from the average latency.

4. The method for evaluating and optimizing the inference performance of a large language model according to claim 2, wherein Dynamically generating concurrent user requests specifically includes: Dynamically generating concurrent user requests using the initialized concurrent user generator; Supporting the adjustment of the number of users based on an adaptive gain algorithm to simulate different loads; Each user request is independently generated and sent to the server asynchronously.

5. The method for evaluating and optimizing the inference performance of a large language model according to claim 2, wherein Collecting performance data based on the concurrent user requests specifically includes: Setting a time window; Within the time window, using the initialized performance metric collector to collect performance metrics, where the performance metrics include: maximum number of users, active users, requests processed per second, response latency, time to first token, inter-token latency, and tokens input / output per second.

6. The method for evaluating and optimizing the inference performance of the large language model according to claim 5, wherein, Within the time window, using the initialized performance metric collector to collect performance metrics specifically includes: Calculating the requests processed per second, with the formula as follows: Calculating the response latency specifically includes: Calculating the average response latency, with the formula as follows: Among them, Response_Time i is the response time of the i-th request, and N is the total number of requests within the time window; Calculating the response latency distribution, with the formula as follows: Response latency distribution = sorted value of the response time at the x% percentile; Calculation of time to first token (TTFT) and inter-token latency (ITL): Calculating the time to first token, with the formula as follows: Among them, First_Token_Time i is the time when the i-th request receives the first token, and N is the total number of requests within a specific time window; Calculating the inter-token latency, with the formula as follows: Among them, Token_Time i is the output time of the i-th token; Calculating the tokens output per second, with the formula as follows: Where the total number of tokens refers to the sum of tokens in all requests processed by the system within a given time window; Calculating the tokens output per second per request, with the formula as follows:

7. The method for evaluating and optimizing the inference performance of a large language model according to any one of claims 2 to 6, characterized in that Performing real-time monitoring and reporting based on the performance data specifically includes: Regularly summarizing and reporting the current performance metrics to obtain real-time monitoring data; The reporting time window can be dynamically adjusted; Dynamically adjusting the number of concurrent users according to the real-time monitoring data.

8. The method for evaluating and optimizing the inference performance of a large language model according to any one of claims 1 to 6, characterized in that, The content of the performance report includes: detailed statistics of average throughput, response latency, requests processed per second, time to first token, inter-token latency, and tokens input / output per second under different time windows.

9. An electronic device, characterized in that, Includes: At least one processor; A memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the large language model inference performance evaluation and optimization method as described in any one of claims 1 to 8.

10. A computer storage medium, characterized in that, A computer program is stored, and when the computer program is executed by a processor, it implements a method for evaluating and optimizing the inference performance of a large language model as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Performance pressure testing system and method applied to large model streaming interface

    CN119645840A

Cited By

  • Service performance evaluation method, computer program product and electronic device

    CN120723342A

  • All-in-one machine performance evaluation method and electronic equipment

    CN120950363A

  • All-in-one machine performance evaluation method and electronic device

    CN120950363B

  • Model performance test method and device, electronic equipment and storage medium

    CN121117539A