Test method and device for transverse evaluation of large model performance
Through the testing method for horizontal evaluation of large-scale model performance, the problem that existing tools cannot capture the core performance characteristics of streaming responses is solved, and refined monitoring and comparative analysis of large-scale model performance is realized, scientific basis is provided for model selection and resource scheduling, and the operability and insight of test results are improved.
Patent Information
- Application Number
- CN202510874374.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing performance testing tools cannot effectively capture the core performance characteristics of large-scale model streaming response, such as first token latency, inter-token latency, token generation rate and long text context processing performance decay, making it difficult to meet the needs of in-depth horizontal evaluation of large-scale model performance.
It provides a test method for horizontal evaluation of large models, including intelligent test data generation and management, streaming performance monitoring based on Token timestamps, dynamic concurrency and traffic simulation, unified access and execution of multiple models, in-depth result analysis and visualization, and resource monitoring. Through these steps, key indicators are obtained and detailed performance comparison data are generated.
It realizes refined monitoring and comparative analysis of the performance of large models, provides scientific basis for model selection and resource scheduling, improves the operability and insight of test results, and supports multi-model testing and practical application scenario simulation.
Smart Images

Figure CN120371654A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of large artificial intelligence models, and specifically provides a test method and device for horizontal evaluation of large model performance. Background Art
[0002] With the rapid development of large model technology, there are a wide variety of models with significant performance differences. Traditional performance testing tools (such as API stress testing tools) mainly focus on request-level metrics (such as total response time, throughput, error rate), and cannot effectively capture the core performance characteristics of large model streaming responses (Token-by-Token generation), for example: Time to First Token (TTFT): A key metric for users to perceive the response speed.
[0003] Inter-Token Latency: Affects the response fluency.
[0004] Tokens per Second: A direct manifestation of the model inference efficiency.
[0005] Performance degradation in processing long text contexts: The stability of the model when processing long inputs.
[0006] The impact of different input complexities (semantics, length) on performance.
[0007] Existing tools lack the ability to finely monitor, compare and analyze these key metrics of the streaming process, as well as test schemes for specific scenarios of large models (such as long contexts, multi-turn conversations), and it is difficult to meet the needs of in-depth horizontal evaluation of large model performance. Summary of the Invention
[0008] The present invention aims at the deficiencies of the above-mentioned existing technologies and provides a practical test method for horizontal evaluation of large model performance.
[0009] A further technical task of the present invention is to provide a test device for horizontal evaluation of large model performance with reasonable design, safety and applicability.
[0010] The technical solution adopted by the present invention to solve its technical problems is: A test method for horizontal evaluation of large model performance has the following steps: S1. Intelligent test data generation and management; S2. Streaming performance monitoring based on Token timestamps; S3. Dynamic concurrency and traffic simulation; S4. Unified access and execution of multiple models; S5. Depth result analysis and visualization; S6. Resource monitoring.
[0011] Furthermore, in step S1, it includes: S1-1. Gradient input length, provide standard length character options to cover long-tail scenarios; S1-2. Generate or select input data by length, or by semantics and task complexity, evaluate the impact of different task types on performance, and use NLP technology to assist in defining complexity; S1-3. The execution mode is sequential execution or random execution; S1-4. Optimize the data loading speed; S1-5. Construct context input for simulating multi-round dialogue scenarios and test the performance of the model in continuous interactions.
[0012] Furthermore, in step S2, intercept and record each Token returned by the large model streaming API and its corresponding server-side generated timestamp; It includes: S2-1. Parse the streaming response at the client or proxy layer and assign an accurate timestamp to each valid Token; S2-2. Real-time calculation of key metrics: Time to First Token TTFT: `t0 - request_send_time`; Time to Nth Token; Inter-Token Latency: `t_i - t_{i - 1}`; Tokens per Second, TPS: instantaneous rate, average rate, Time to Last Token; S2-3. Higher-order metric calculation; Latency distribution statistics: Calculate the percentile values of TTFT and Inter-Token Latency to evaluate response stability and tail latency; Throughput fluctuation analysis: Calculate the variance of TPS or draw a time series diagram to identify fluctuations or stutters during the generation process; Long text performance decay analysis: Compare the change rates of TTFT, average Inter-Token Latency, and TPS when the same model processes short texts and long texts, and quantify the performance loss brought by long contexts; S2-4. Data storage: Ensure the complete storage of the original Token timestamp sequence and the calculation results of higher-order metrics.
[0013] Further, in step S3, it includes: S3-1. Achieving high concurrency through multithreading or coroutines or asynchronous I / O; S3-2. Dynamically setting the number of concurrent users; S3-3. Setting the request interval; S3-4. Designing a warm-up stage to make the system enter a stable state and collecting performance data in the steady state stage; S3-5. Dynamically adjusting the concurrency number according to the time script or function to simulate tidal traffic and test the elastic ability of the model.
[0014] Further, in step S4, it includes: S4-1. Presetting a model access configuration template; S4-2. Defining the access to any large model API that conforms to the specification through the API URL, authentication information, request, and response structure; S4-3. Making all the models under test run in the same time period, using the same test data set and concurrent pressure configuration.
[0015] Further, in step S5, deeply analyzing, comparing, and visualizing the collected raw data at the Token level and high-order metrics to generate a horizontal comparison report, including: S5-1. Conducting a side-by-side comparison of the key metrics of different models in the same test scenario and generating box plots, line charts, and bar charts for visual display; S5-2. Combining the Token-level latency data with model and infrastructure knowledge to analyze the source of bottlenecks and provide analysis clues; S5-3. Automatically analyzing the change curves of the key metrics of each model under different concurrencies and identifying performance inflection points; S5-4. Summarizing the detailed performance data and comparative analysis results of all the models under test in the same test batch into a single report file, supporting Excel, SQL, and interactive HTML reports, and allowing users to select the metrics, models, and scenarios they are interested in to generate reports.
[0016] Further, in step S6, when deploying the test model in a controllable environment, synchronously monitoring the system resource consumption during the model operation; Integrating Prometheus, NVIDIA DCGM, or operating system-level monitoring tools to collect data and performing correlation analysis with performance metrics to help determine whether the performance bottleneck is due to insufficient computing resources.
[0017] A test device for horizontal evaluation of the performance of large models, including: at least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable programs and execute a test method for horizontal evaluation of large model performance.
[0018] Compared with the prior art, a test method and device for horizontal evaluation of large model performance according to the present invention have the following prominent beneficial effects: (1) The fine monitoring method based on Token timestamps can obtain key indicators such as TTFT, inter-Token delay, TPS and its distribution, and truly reveal the internal efficiency of the large model response generation process.
[0019] (2) Systematically test the influence of different input lengths (especially long texts) and optional complexities on performance, quantify the performance attenuation of long contexts, and provide key basis for model selection (especially for long document applications).
[0020] (3) Ensure that multiple models run under completely consistent test conditions (data set, concurrency mode, time window), and make comparisons based on refined metrics calculated uniformly, and the results are more comparable and persuasive.
[0021] (4) Beyond simple data collection, provide automated in-depth comparative analysis, visualization reports and bottleneck attribution clues (such as performance inflection point identification, delay distribution analysis, long text attenuation rate calculation), greatly improve the operability and insight depth of test results, and help R & D personnel quickly locate the optimization direction.
[0022] (5) Support basic concurrent stress testing, long text special testing, multi-round dialogue simulation, traffic fluctuation simulation, etc., and are closer to actual application scenarios.
[0023] (6) By concurrently testing multiple models, automated processes and intelligent report generation, significantly shorten the cycle from test execution to obtaining in-depth insights.
[0024] (7) Modular design, supporting flexible addition of new models, new monitoring metrics, new test scenarios and analysis dimensions. Description of the Drawings
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 It is a schematic flowchart of a test method for horizontal evaluation of large model performance. Detailed Embodiments
[0027] To enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0028] The following gives a best embodiment: Such as Figure 1 As shown, a test method for horizontal evaluation of large model performance in this embodiment has the following steps: S1. Intelligent test data generation and management; Including: S1-1. Provide standard length options (100, 500, 1000, 5000, 10000, 20000 characters), focusing on covering long-tail scenarios (>5K characters) to test the long context processing ability and performance degradation of the model.
[0029] S1-2. Generate or select input data not only by length, but also by semantic / task complexity classification (such as: simple Q&A, complex reasoning, code generation, summarization tasks, etc.) to evaluate the impact of different task types on performance. NLP techniques (such as perplexity, syntactic complexity) can be used to assist in defining complexity.
[0030] S1-3. Execution mode: sequential execution or random execution.
[0031] S1-4. Data caching: Optimize data loading speed.
[0032] S1-5. Support the construction of context input for simulating multi-round dialogue scenarios to test the performance of the model in continuous interaction.
[0033] S2. Streaming performance monitoring based on Token timestamps; Precisely intercept and record each Token returned by the large model streaming API and its corresponding server-side generation timestamp (or client-side reception timestamp).
[0034] Including: S2-1. Parse the streaming response (such as SSE, Websocket) at the client or proxy layer, and mark each valid Token with an accurate timestamp (t0, t1, t2,..., tn).
[0035] S2-2. Real-time calculation of key metrics: First Token Time TTFT: `t0 - request_send_time`; Time to Nth Token Inter-Token Latency: `t_i - t_{i - 1}` Tokens per Second, TPS: Instantaneous rate, average rate Time to Last Token S2-3, High-order metric calculation: Latency distribution statistics: Calculate percentile values such as P50, P90, P95, P99, Max of TTFT and Inter-Token Latency to evaluate response stability and tail latency.
[0036] Throughput fluctuation analysis: Calculate the variance of TPS or draw a time series diagram to identify fluctuations or lags in the generation process.
[0037] Long text performance degradation analysis: Compare the change rates of TTFT, average Inter-Token Latency, and TPS when the same model processes short texts and long texts to quantify the performance loss caused by long contexts.
[0038] S2-4, Data storage: Ensure the complete storage of the original Token timestamp sequence and the calculation results of high-order metrics.
[0039] S3, Dynamic concurrency and traffic simulation; Simulate the real user access pattern and apply controllable concurrency pressure.
[0040] Including: S3-1, Achieve high concurrency through multi-threading or coroutines or asynchronous IO.
[0041] S3-2, Dynamically set the number of concurrent users.
[0042] S3-3, Set the request interval (fixed, random, following a distribution such as Poisson distribution).
[0043] S3-4, Design a warm-up phase (gradually increase concurrency) to make the system enter a stable state, and collect the main performance data in the steady state phase to avoid the impact of cold start.
[0044] S3-5, Support dynamically adjusting the concurrency number according to a time script or function to simulate tidal traffic and test the elastic ability of the model.
[0045] S4, Unified access and execution of multiple models; Configure, manage, and concurrently execute test tasks of multiple different large models.
[0046] Including: S4-1. Access configuration templates for pre-set mainstream models (such as Qwen-plus, deepseek-r1, grok-3, chatgpt-4, etc.).
[0047] S4-2. Flexible custom model access: Support accessing any large model API that complies with the specification through API URL, authentication information, and request / response structure definitions.
[0048] S4-3. Concurrent execution engine: Ensure that all tested models run in the same time period with the same (or equivalent) test data sets and concurrent pressure configurations to ensure the fair comparability of test results.
[0049] S5. In-depth result analysis and visualization; Conduct in-depth analysis, comparison, and visualization of the collected raw Token-level data and high-level metrics to generate an easy-to-understand horizontal comparison report.
[0050] Including: S5-1. Comparative analysis of multi-model metrics, juxtaposing key metrics (TTFT, average / tail Inter-Token Latency, TPS, long text decay rate) of different models under the same test scenarios (same input length / complexity, same concurrency).
[0051] S5-2. Visualization display: Generate box plots (showing latency distribution), line charts (showing performance changes under different lengths / concurrencies), bar charts (directly comparing metric values), etc.
[0052] S5-3. Combine Token-level latency data with model and infrastructure knowledge to attempt to analyze possible sources of performance bottlenecks (e.g., high initial generation latency may involve Prompt processing and pre-filling; high inter-Token latency may involve autoregressive decoding efficiency; significant long text decay may involve KV Cache efficiency or attention calculation overhead). Provide analysis clues.
[0053] S5-4. Automatically analyze the change curves of key metrics (such as TPS, error rate, latency percentile values) of each model under different concurrencies to identify performance inflection points (such as throughput saturation points, latency steep increase points).
[0054] S5-5, Report Output: Aggregate the detailed performance data (raw + aggregated) and comparative analysis results of all tested models in the same test batch into a single report file. Support for Excel (for easy viewing of details), SQL (for easy persistence and complex queries), and interactive HTML reports (an innovation point, including rich visual charts and dynamic filtering). Allow users to select the metrics, models, and scenarios they are interested in for report generation.
[0055] S6, Resource Monitoring; When the test model is deployed in a controllable environment (such as a self-owned GPU server), synchronously monitor the system resource consumption during model operation (such as GPU utilization, video memory occupancy, CPU utilization, memory occupancy).
[0056] Collect data by integrating Prometheus, NVIDIA DCGM, or operating system-level monitoring tools (such as psutil), and perform correlation analysis with performance metrics (TPS, latency) to help determine whether the performance bottleneck stems from insufficient computing resources (such as slow token generation caused by GPU bottlenecks).
[0057] Based on the above method, a test device for horizontal performance evaluation of large models in this embodiment includes: at least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program and execute a test method for horizontal performance evaluation of large models.
[0058] The above specific implementation manners are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific implementation manners. Any technical solution that conforms to the technical solutions described in the above specific implementation manners of the present invention and any appropriate changes or substitutions made by those of ordinary skill in the art shall fall within the patent protection scope of the present invention.
[0059] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A test method for horizontal evaluation of large model performance, characterized in that, The steps are as follows: S1. Intelligent test data generation and management; S2, streaming performance monitoring based on Token timestamp; S3, dynamic concurrency and traffic simulation; S4, unified access and execution of multiple models; S5, in-depth result analysis and visualization; S6. Resource monitoring.
2. The test method for horizontal evaluation of large model performance according to claim 1, wherein In step S1, it includes: S1-1, input length gradient, providing standard length character options to cover long-tail scenarios; S1-2, generate or select input data by length, semantics, or task complexity, evaluate the impact of different task types on performance, and use NLP technology to assist in defining complexity; S1-3, the execution mode is sequential execution or random execution; S1-4, optimize data loading speed; S1-5. Construct context input to simulate multi-round dialogue scenarios and test the performance of the model in continuous interaction.
3. The test method for horizontal evaluation of large model performance according to claim 2, wherein In step S2, each Token returned by the large model streaming API and its corresponding server-side generated timestamp are intercepted and recorded; include: S2-1. Parse the streaming response at the client or proxy layer and add an accurate timestamp to each valid token. S2-2. Real-time calculation of key indicators: First Token time TTFT: `t0-request_send_time`; Time to Nth Token; Inter-Token Latency: `t_i-t_{i-1}`; Token generation rate Tokens per Second, TPS: instantaneous rate, average rate, Response completion time Time to Last Token; S2-3, calculation of high-level indicators; Latency distribution statistics: calculate the percentile values of TTFT and Inter-Token Latency to evaluate response stability and tail latency; Throughput fluctuation analysis: Calculate the variance of TPS or draw a timing diagram to identify fluctuations or freezes in the generation process; Long text performance degradation analysis: Compare the change rates of TTFT, average Inter-Token Latency, and TPS when the same model processes short text and long text, and quantify the performance loss caused by long context; S2-4. Data storage: Ensure the complete storage of the original Token timestamp sequence and high-level indicator calculation results.
4. The test method for horizontal evaluation of large model performance according to claim 3, wherein, In step S3, it includes: S3-1, multithreading or coroutine or asynchronous IO to achieve high concurrency; S3-2, dynamically set the number of concurrent users; S3-3, set the request interval; S3-4, design the preheating phase to make the system enter a stable state, and collect performance data during the steady-state phase; S3-5. Dynamically adjust the number of concurrency according to the time script or function, simulate tidal flow, and test the elasticity of the model.
5. The test method for horizontal performance evaluation of large models according to claim 4, wherein In step S4, it includes: S4-1, preset model access configuration template; S4-2, access any large model API that complies with the specification through API URL, authentication information, request, and response structure definitions; S4-3. Run all the tested models during the same time period, using the same test data set and concurrent stress configuration.
6. The test method for horizontal evaluation of large model performance according to claim 5, wherein, In step S5, conduct in-depth analysis, comparison, and visualization of the collected raw Token-level data and high-level metrics to generate a cross-evaluation report, including: S5-1. Compare the key metrics of different models under the same test scenario side by side, and generate box plots, line charts, and bar charts for visual display; S5-2. Combine the Token-level latency data with model and infrastructure knowledge to analyze the sources of bottlenecks and provide analysis clues; S5-3. Automatically analyze the change curves of the key metrics of each model under different concurrencies to identify performance inflection points; S5-4. Summarize the detailed performance data and comparative analysis results of all the tested models in the same test batch into a single report file, supporting Excel, SQL, and interactive HTML reports, and allowing users to select the metrics, models, and scenarios of interest for report generation.
7. A test method for horizontal evaluation of large model performance according to claim 6, characterized in that In step S6, when deploying the test model in a controllable environment, synchronously monitor the system resource consumption during the model operation; Integrate Prometheus, NVIDIA DCGM, or operating system-level monitoring tools to collect data and conduct correlation analysis with performance metrics to help determine whether the performance bottleneck is due to insufficient computing resources.
8. A test device for horizontal performance evaluation of large models, characterized in that, Including: At least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program and execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method, device and equipment for evaluating model performance, medium and program product
CN117648140A
NLP model performance evaluation method and device, storage medium and electronic equipment
CN117724965A
Travel comfort prediction method based on artificial intelligence and sensing technology
CN118113995A
Model quality monitoring system based on machine learning operation and maintenance
CN118839791A
Data analysis system based on artificial intelligence
CN119066423A
Cited By
Test method of retrieval enhancement generation system and electronic equipment
CN121349828A
End-to-end AI agent evaluation method, system, equipment and medium
CN121900800A