Pressure measurement method and device for large model and electronic equipment

CN120011189APending Publication Date: 2025-05-16SHANGHAI IQIYI NEW MEDIA TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411983311.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Large models may be in an overload state when processing a large number of user requests, which affects their operating stability and response capabilities, resulting in the inability to process user requests in a timely manner.

Method used

A pressure measurement method for a large model is provided. By obtaining the target number of test requests and sending them to the target large model in parallel, receiving response results, and determining the reception time information according to the data output method, and analyzing the pressure measurement results of the large model in high concurrency situations.

Benefits of technology

Evaluate the load capacity of the large model through pressure measurement, help control data flow, reduce the possibility of the large model being in an overloaded operating state, and ensure its stability and responsiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011189A_ABST
    Figure CN120011189A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a pressure measurement method and device for a large model and electronic equipment, and relates to the technical field of computers, and the method comprises the steps: obtaining a target number of test requests for carrying out the pressure measurement of a target large model; the target number of test requests are sent to the target large model in parallel, so that the target large model responds to the received target number of test requests and feeds back response results; receiving a response result corresponding to each test request fed back by the target large model; determining target receiving time information corresponding to each test request according to the target data output mode; and based on the target receiving time information corresponding to each test request, analyzing a pressure test result of the target large model when the target number is used as the concurrent request number. Therefore, through the scheme, the load capacity of the large model can be evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a large-model stress testing method, device, and electronic equipment. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, big model technology, as an important driving force and implementation means of its technological development, has also developed rapidly.

[0003] Moreover, as the number of Internet users increases, the number of requests that the large model needs to process increases dramatically. For example, during certain periods of time, the number of user requests increases dramatically, and the large model needs to process a large number of user requests at the same time. The number of user requests that it needs to process may exceed the maximum number of concurrent processing supported, that is, processing these user requests at the same time may cause the large model to be in an overloaded operating state, affecting its operating stability, and even causing the large model to be unable to respond to user requests in a timely manner.

[0004] Based on this, there is a need to perform stress testing on large models to evaluate their load capacity, so that data flow control and other operations can be performed based on the load capacity of the large models, thereby reducing the possibility of the large model being in an overloaded operating state.

[0005] Therefore, a large model stress testing method is urgently needed to evaluate the load capacity of large models. Summary of the invention

[0006] The purpose of the embodiments of the present application is to provide a large model stress testing method, device and electronic device, so as to evaluate the load capacity of the large model through stress testing of the large model, so as to perform data flow control and other operations based on the load capacity of the large model, thereby reducing the possibility of the large model being in an overloaded operating state. The specific technical solution is as follows:

[0007] In a first aspect provided by an embodiment of the present application, a stress testing method for a large model is first provided, the method comprising:

[0008] Get the target number of test requests for stress testing the target large model;

[0009] Sending the target number of test requests to the target large model in parallel, so that the target large model responds to the received target number of test requests and feeds back the response results;

[0010] Receive a response result corresponding to each test request fed back by the target large model; wherein the response result is a result containing one or more object contents;

[0011] According to the target data output mode, the target receiving time information corresponding to each test request is determined; wherein the target data output mode is the content presentation mode of the target macro model when feeding back the object content in the response result; the target receiving time information corresponding to each test request is the receiving time information of the object content in the response result corresponding to the test request;

[0012] Based on the target receiving time information corresponding to each test request, the stress test results of the target large model when the target number is used as the number of concurrent requests are analyzed.

[0013] In a second aspect provided by an embodiment of the present application, a large model stress testing device is provided, the device comprising:

[0014] An acquisition module is used to obtain a target number of test requests for stress testing a target large model;

[0015] A sending module, used for sending the target number of test requests to the target large model in parallel, so that the target large model responds to the received target number of test requests and feeds back the response result;

[0016] A receiving module, used to receive a response result corresponding to each test request fed back by the target large model; wherein the response result is a result containing one or more object contents;

[0017] A determination module, used to determine the target reception time information corresponding to each test request according to the target data output mode; wherein the target data output mode is the content presentation mode of the target macro model when feeding back the object content in the response result; the target reception time information corresponding to each test request is the reception time information of the object content in the response result corresponding to the test request;

[0018] The analysis module is used to analyze the stress test results of the target large model when the target number is used as the number of concurrent requests based on the target reception time information corresponding to each test request.

[0019] In the third aspect provided by the embodiment of the present application, there is also provided an electronic device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory is used to store computer programs; the processor is used to implement the stress testing method of any large model provided in the first aspect when executing the program stored in the memory.

[0020] In another aspect provided by an embodiment of the present application, a computer-readable storage medium is provided, wherein a computer program is stored in the computer storage medium, and when the computer program is executed by a processor, the stress testing method of any large model provided in the first aspect above is implemented.

[0021] In another aspect of the embodiments of the present application, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute any large model stress testing method provided in the first aspect above.

[0022] An embodiment of the present application provides a stress testing method for a large model. First, a target number of test requests for stress testing a target large model are obtained, and a target number of test requests are sent to the target large model in parallel, so that the target large model responds to the received target number of test requests and feeds back the response results. Thus, after receiving the response result corresponding to each test request fed back by the target large model, the target receiving time information corresponding to each test request is determined according to the target data output mode of the target large model when feeding back the response result. Then, based on the target receiving time information corresponding to each test request, the stress testing result of the target large model when the target number is used as the number of concurrent requests is analyzed.

[0023] Based on this, the solution provided by the embodiment of the present application is applied. Under different data output modes, the target large model presents the object content in the feedback response result in different ways. Therefore, the receiving time information of the object content in the response result corresponding to each test request can be determined according to the target data output mode of the target large model, as the target receiving time information corresponding to the test request. Thus, the obtained target receiving time information is used to analyze and obtain the stress test results of the target large model when the target number is used as the number of concurrent requests. Thus, based on the obtained stress test results, stress testing of the large model is implemented, and then, the load capacity of the large model is evaluated, so that data flow control and other operations can be performed based on the load capacity of the large model in the future, thereby further reducing the possibility of the large model being in an overloaded operating state. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art are briefly introduced below.

[0025] Figure 1 A schematic diagram of a flow chart of a first large-model stress testing method provided in an embodiment of the present application;

[0026] Figure 2 A schematic diagram of a flow chart of a second large model stress testing method provided in an embodiment of the present application;

[0027] Figure 3 A schematic diagram of a flow chart of a third large model stress testing method provided in an embodiment of the present application;

[0028] Figure 4A schematic diagram of a structure of a specific embodiment provided for the present application;

[0029] Figure 5 A schematic diagram of the first stress test result provided in an embodiment of the present application;

[0030] Figure 6 A schematic diagram of a second stress test result provided in an embodiment of the present application;

[0031] Figure 7 A schematic diagram of a third stress test result provided in an embodiment of the present application;

[0032] Figure 8 A schematic diagram of a fourth stress test result provided in an embodiment of the present application;

[0033] Fig. 9 A schematic diagram of the structure of a large-scale stress testing device provided in an embodiment of the present application;

[0034] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0035] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0036] In order to better understand the present application, the professional terms involved in the embodiments of the present application are first introduced below.

[0037] In general, large models can be divided into two types: streaming output and non-streaming output according to the data output method when feeding back the response results. Among them, the so-called streaming output means that when the model feeds back the results or generates the answer, it gradually presents the results to the user one word or one phrase at a time, simulating the typewriter-style output effect, thereby giving the user an animation feeling that the answer gradually appears. This method is common in natural language processing models, especially large language models (LLMs), which can generate coherent and detailed text responses. The so-called non-streaming output means that when the model feeds back the results or generates the answer, it can generate an entire sentence or paragraph, or even an entire document at one time. This output method allows the model to generate more natural and coherent text in a short time, and is suitable for application scenarios that require rapid generation of a large amount of content, such as generative pre-training models.

[0038] Stress testing (also known as stress testing) is mainly aimed at verifying the stability and performance of large models under high load.

[0039] Large models usually refer to machine learning models with large-scale parameters and computing power. These models are usually built with deep neural networks, have billions or even hundreds of billions of parameters, can process massive amounts of data, and exhibit strong learning and generalization capabilities. However, fine-tuning large models on specific tasks often requires a lot of computing resources and time, which limits their flexibility in practical applications.

[0040] LoRA (Low-Rank Adaptation) is a fine-tuning technology for large models. It aims to adapt to downstream tasks by training a small number of parameters, thereby reducing the trainable parameters of the model while minimizing the loss of model performance, so as to make the fine-tuning of large models on specific tasks more efficient and flexible. For example, when adjusting the parameters of a large model with large-scale parameters, modifying any of the parameters may affect the more complex parameters in the entire large model, and even cause the large model to fail to operate normally. This is extremely time-consuming and difficult. However, using LoRA, some small switches can be added to the entire large model. In this way, by adjusting these small switches, the machine behavior of the large model can be changed without touching the more complex parameters, so that the large model can quickly adapt to different business scenarios and improve the flexibility of the large model.

[0041] The above-mentioned small switch can be understood as a LoRA model built in a large model based on LoRA technology. This type of model has the advantages of fast training speed, low computing requirements, and small training weights. By freezing the original model, the required LoRA model is built or adjusted without directly modifying the weights of the original model, so as to add some trainable parameters for weight adjustment, so that the large model can quickly adapt to different business scenarios and support processing requests of different request types.

[0042] A container instance usually refers to a single application or service instance running in containerization technology. Containerization technology allows developers to package applications and their dependencies into an independent, portable container, thereby achieving consistent operating results in different environments. In cloud computing and microservice architectures, container instances are widely used to deploy and manage application services. In other words, large models or LoRA models can be deployed to multiple container instances to achieve distributed processing and scalability of the model.

[0043] For example, currently, teams in various companies use large models such as LLM for business output, that is, to process requests sent by users, and in order to improve the utilization of GPU (Graphics Processing Unit) resources, multi-LoRA integration is usually used to meet different business needs (processing different types of requests), that is, by integrating multiple LoRA models for processing different types of requests, the large model can support processing requests of different request types, expanding the flexibility and adaptability of the large model. When the business demand increases, the performance of the container instance where the model is deployed may become a bottleneck for the entire large model to respond to user needs in a timely manner.

[0044] During the application of the large model, the number of user requests increases dramatically in certain periods of time. The large model needs to process a large number of user requests at the same time. The number of user requests that it needs to process exceeds the maximum number of concurrent processing supported. That is, processing these user requests at the same time may cause the large model to be in an overloaded operating state, affecting its operating stability and even causing the large model to be unable to respond to user requests in a timely manner.

[0045] Therefore, there is a need to perform stress testing on large models to evaluate the load capacity of the large models, so as to perform data flow control and other operations based on the load capacity of the large models, thereby reducing the possibility of the large models being in an overloaded operating state. Therefore, a stress testing method for large models is urgently needed to evaluate the load capacity of large models.

[0046] In order to solve the above technical problems, the embodiments of the present application provide a large-model stress testing method, device and electronic device.

[0047] Among them, the method can be applied to various application scenarios of stress testing large models. For example, stress testing natural language processing models, stress testing generative pre-trained models, etc. In addition, the method can be applied to various electronic devices such as laptops, tablet computers, and desktop computers that store applications or modules for executing a stress testing method for a large model provided in an embodiment of the present application, hereinafter referred to as electronic devices. Based on this, the embodiment of the present application does not limit the application scenarios and execution subjects of the method.

[0048] A large model stress testing method provided in an embodiment of the present application may include the following steps:

[0049] Get the target number of test requests for stress testing the target large model;

[0050] Sending the target number of test requests to the target large model in parallel, so that the target large model responds to the received target number of test requests and feeds back the response results;

[0051] Receive a response result corresponding to each test request fed back by the target large model; wherein the response result is a result containing one or more object contents;

[0052] According to the target data output mode, the target receiving time information corresponding to each test request is determined; wherein the target data output mode is the content presentation mode of the target macro model when feeding back the object content in the response result; the target receiving time information corresponding to each test request is the receiving time information of the object content in the response result corresponding to the test request;

[0053] Based on the target receiving time information corresponding to each test request, the stress test results of the target large model when the target number is used as the number of concurrent requests are analyzed.

[0054] As can be seen from the above, an embodiment of the present application provides a stress testing method for a large model. First, a target number of test requests for stress testing a target large model are obtained, and a target number of test requests are sent to the target large model in parallel, so that the target large model responds to the received target number of test requests and feeds back the response results. Thus, after receiving the response result corresponding to each test request fed back by the target large model, the target receiving time information corresponding to each test request is determined according to the target data output mode of the target large model when feeding back the response result. Then, based on the target receiving time information corresponding to each test request, the stress testing result of the target large model when the target number is used as the number of concurrent requests is analyzed.

[0055] Based on this, the solution provided by the embodiment of the present application is applied. Under different data output modes, the target large model presents the object content in the feedback response result in different ways. Therefore, the receiving time information of the object content in the response result corresponding to each test request can be determined according to the target data output mode of the target large model, as the target receiving time information corresponding to the test request. Thus, the obtained target receiving time information is used to analyze and obtain the stress test results of the target large model when the target number is used as the number of concurrent requests. Thus, based on the obtained stress test results, stress testing of the large model is implemented, and then, the load capacity of the large model is evaluated, so that data flow control and other operations can be performed based on the load capacity of the large model in the future, thereby further reducing the possibility of the large model being in an overloaded operating state.

[0056] The following is a detailed description of a large model stress testing method provided in an embodiment of the present application in conjunction with the accompanying drawings.

[0057] Figure 1 A schematic diagram of a large model stress testing method provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the method may include the following steps:

[0058] S101: Obtain a target number of test requests for stress testing a target large model.

[0059] In this application, for the request types supported by the target large model and the target number of concurrent requests to be tested by the target large model at the current moment, a target number of test requests for stress testing the target large model is obtained. It should be noted that the above target number of test requests are data requests processed by the target large model.

[0060] Among them, the target large model can be an LLM model, or a GPT-3 (Generative Pre-trained Transformer 3) model, etc., and this application does not make any specific limitations on this.

[0061] In addition, the target large model may support multiple or one request types, which is reasonable. Moreover, the above request type may be a request type for text or a request type for image, which is not specifically limited in the embodiment of the present application.

[0062] Optionally, when the target large model supports multiple request types, each execution of the stress testing method based on the target number of test requests can be a request type. For example, the target number is 10, and the target large model supports text and image request types. When the stress test is performed on the target large model for the first time, the target number is 10. At this time, all 10 test requests are text type test requests. When the stress test is performed on the target large model for the second time, all 10 test requests are image type test requests. The above method for determining the request type for the target number of test requests can also be called a polling method.

[0063] Optionally, in one implementation, the above step S101, obtaining a target number of test requests for stress testing the target large model, may include the following steps:

[0064] Step A: In a predetermined request database corresponding to the target large model, a target number of test requests for stress testing the target large model are obtained.

[0065] Among them, when the target large model is an online large model, the predetermined request database is the historical request database corresponding to the target large model; when the target large model is an offline large model, the predetermined request database is a request database manually constructed for the target large model.

[0066] In this implementation, a target number of test requests for stress testing the target large model are obtained in the predetermined request database corresponding to the target large model. In addition, it is considered that the target large model can be a large model that has been online and has historical request data, or a large model that has not been online and has no historical request data. When the target large model is a large model that has been online, the above-mentioned predetermined request database is the historical request database corresponding to the target large model, and accordingly, the obtained test requests are the historical requests of the target large model. When the target large model is a large model that has not been online, since the target large model has not actually run and processed user requests, but the request types supported by the target large model are known, a request database for the target large model can be manually constructed in advance to simulate the requests sent by users, and a target number of test requests can be obtained from the request database.

[0067] This implementation method directly obtains the target number of test requests from the predetermined request database corresponding to the target large model, while ensuring that the request type obtained is the request type supported and processed by the target large model, and provides a sufficient number of test requests for subsequent stress testing on the target large model to support stress testing on the target large model.

[0068] S102: Send a target number of test requests in parallel to the target large model, so that the target large model responds to the received target number of test requests and feeds back the response results.

[0069] S103: Receive a response result corresponding to each test request fed back by the target large model;

[0070] The response result is a result containing an or object content.

[0071] In the present application, a target number of test requests are sent in parallel to the target large model, so that the target large model responds after receiving the target number of test requests, and then feeds back the response results to the electronic device, which receives the response results corresponding to each test request fed back by the target large model; wherein the response results include one or more object contents.

[0072] Optionally, the electronic device calls the thread pool executor to send a target number of test requests to the target large model in a concurrent manner, that is, the electronic device calls the thread pool executor to create a target thread for the target large model for stress testing the target large model. Specifically, the electronic device calls the thread pool executor to create a first function and a second function in the above-mentioned target thread, and then uses the first function to obtain a target number of test requests, and submits the target number of test requests to be sent to the target large model to the second function, so that the second function sends the target number of test requests to the target large model, and receives the response results corresponding to each test request fed back by the target large model.

[0073] Among them, the first function can be functions such as submit_request (a function for submitting a request), and the second function can be llm_request (a function for processing requests for large language models), etc., and this embodiment of the present application does not specifically limit this. It should be noted that the above-mentioned second function is a request function for sending and receiving a target large model, which is a function type that can communicate with the target large model. For example, when the target large model is an LLM model, the above-mentioned second function is the llm_request function, and when the target large model is a GPT-3 model, the above-mentioned second function is the gpt-3_request function.

[0074] S104: Determine target receiving time information corresponding to each test request according to the target data output mode.

[0075] Among them, the target data output method is the content presentation method of the target large model when feeding back the object content in the response result; the target reception time information corresponding to each test request is the reception time information of the object content in the response result corresponding to the test request.

[0076] In this application, it is considered that under different data output modes, the target large model presents the object content in the feedback response result in different ways. Accordingly, the data information that needs to be paid attention to and used for stress testing analysis when determining the received response result is different.

[0077] For example, when a large model with streaming output mode feeds back the object content in the response result, the content presentation method is: the result is gradually presented to the user one word or one phrase at a time, simulating the typewriter-style output effect. Therefore, when stress testing this type of large model, the focus is on the reception time information of a single object content in the response result corresponding to the test request.

[0078] Exemplarily, if the object content of the response result is text content, it is output in characters, and when stress testing this type of large model, the focus is on the receiving time information of each character; if the object content of the response result is image content, it is output in sub-images corresponding to each part of the image, and when stress testing this type of large model, the focus is on the receiving time information of each sub-image; if the object content of the response result is audio content, it is output in audio frames, and when stress testing this type of large model, the focus is on the receiving time information of each audio frame.

[0079] For example, when a large model with non-streaming output mode presents the object content in the feedback response result in the following way: an entire sentence or complete text is generated at one time and presented to the user. Therefore, when stress testing such a large model, the focus is on the receiving time information of the response result itself corresponding to the test request.

[0080] Exemplarily, if the object content of the response result is text content, the output is performed with the entire text content as the output unit, and when stress testing this type of large model, the focus is on the receiving time information of the last character; if the object content of the response result is image content, the output is performed with the entire image as the output unit, and when stress testing this type of large model, the focus is on the receiving time information of the last sub-image; if the object content of the response result is audio content, the output is performed with the entire audio segment as the output unit, and when stress testing this type of large model, the focus is on the receiving time information of the last audio frame.

[0081] Therefore, after receiving the response result corresponding to each test request fed back by the target large model, the target receiving time information corresponding to each test request can be determined according to the target data output mode of the target large model.

[0082] The target receiving time information of each test request may include: the object content of the response result, the delay time between adjacent object contents in the response result, the request sending time of the test request, the result receiving time, and the total request time of the test request, etc. This embodiment of the present application does not make specific limitations on this.

[0083] Optionally, the electronic device calls the thread pool executor to create a third function in the target thread, and uses the third function to record the data information generated in the response result process of each test request received, and determines the target reception time information corresponding to each test request in the received data information according to the target data output method.

[0084] It should be noted that the third function may be process_results (a function for processing data information), etc. The data information may include: the number of requests successfully processed by the target large model, the error rate of the target large model processing requests, the delay time of each test request when processing to different percentiles, and the model throughput of the target large model, etc. This embodiment of the application does not specifically limit this.

[0085] Among them, the above-mentioned different percentiles are used to characterize the progress of each test request being processed. For example, when the processing progress of a test request is 25%, its corresponding percentile is 25%. The interval between the receiving time when the processing progress of the test request is 25% and the request sending time of the test request is the delay time when the test request is processed to 25%.

[0086] S105: Based on the target receiving time information corresponding to each test request, analyze the stress test results of the target large model when the target number is used as the number of concurrent requests.

[0087] In this application, after obtaining the target receiving time information of each test request, the stress test result of the target large model when the target number is used as the number of concurrent requests can be analyzed based on each target receiving time information. In other words, by stress testing the target large model, the limit of the number of parallel processing requests of the target large model is determined, and thus, the load capacity of the target large model is evaluated through the determined number limit.

[0088] Optionally, the electronic device calls the thread pool executor to create a fourth function in the target thread, uses the fourth function to count the target reception time information of each test request, and sends the target reception time information concerned by the target data output method to the statistical module according to the target data output method of the target large model, so as to use the statistical module to analyze the target reception time information and obtain the stress test result of the target large model when the target number is used as the number of concurrent requests. Among them, the above-mentioned fourth function can be metrics_summary (a function used to perform stress test indicator statistics, that is, a function used to count the target reception time information), etc., and the above-mentioned statistical module can be Pandas Data Frame (panda data frame, which is a drawing component of the open source library in python), etc. In this regard, the embodiments of the present application do not make specific limitations.

[0089] Optionally, in order to improve the richness of the stress testing results, the above stress testing results also include the processing time of each test request when processing to different percentiles, the average, minimum, maximum and standard deviation of the delay time between each object content, and the content quantity of the object content input and output of each test request.

[0090] As can be seen from the above, an embodiment of the present application provides a stress testing method for a large model. First, a target number of test requests for stress testing a target large model are obtained, and a target number of test requests are sent to the target large model in parallel, so that the target large model responds to the received target number of test requests and feeds back the response results. Thus, after receiving the response result corresponding to each test request fed back by the target large model, the target receiving time information corresponding to each test request is determined according to the target data output mode of the target large model when feeding back the response result. Then, based on the target receiving time information of each test request, the stress testing result of the target large model when the target number is used as the number of concurrent requests is analyzed.

[0091] Based on this, the solution provided by the embodiment of the present application is applied. Under different data output modes, the target large model presents the object content in the feedback response result in different ways. Therefore, the receiving time information of the object content in the response result corresponding to each test request can be determined according to the target data output mode of the target large model, as the target receiving time information corresponding to the test request. Thus, the obtained target receiving time information is used to analyze and obtain the stress test results of the target large model when the target number is used as the number of concurrent requests. Thus, based on the obtained stress test results, stress testing of the large model is implemented, and then, the load capacity of the large model is evaluated, so that data flow control and other operations can be performed based on the load capacity of the large model in the future, thereby further reducing the possibility of the large model being in an overloaded operating state.

[0092] Optionally, in one implementation, Figure 2 A flow chart of another large model stress testing method provided in an embodiment of the present application is shown as follows: Figure 2 As shown, the above step S104, determining the target receiving time information corresponding to each test request according to the target data output mode, may include the following steps:

[0093] S1041: If the target data output mode is a mode of outputting data in units of single object content, for each test request, determine the first receiving time of each object content of the response result of the test request, and obtain target receiving time information corresponding to each test request;

[0094] Accordingly, the above step S105, based on the target receiving time information corresponding to each test request, analyzes the stress test result of the target large model when the target number is used as the number of concurrent requests, and may include the following steps:

[0095] S1051: for each test request, based on the sending time of the test request and the first receiving time of each object content of the response result of the test request, analyzing the first delay time and each second delay time corresponding to the test request;

[0096] The first delay duration corresponding to the test request is used to characterize the interval duration between the first object content of the response result of the test request and the sending time of the test request, and each second delay duration corresponding to the test request is used to characterize the interval duration between adjacent object contents in the response result of the test request;

[0097] S1052: Based on the first delay duration and each second delay duration corresponding to each test request, determine the stress test result of the target large model when the target number is used as the number of concurrent requests.

[0098] In this implementation, if the target data output method is a method of outputting with a single object content as the output unit, this method focuses on the receiving time information of each object content. At this time, for each test request, the first receiving time of each object content that receives the response result of the test request is determined, and the target receiving time information corresponding to each test request is obtained. Thus, based on the sending time of the test request and the first receiving time of each object content that receives the response result of the test request, the first delay duration and each second delay duration corresponding to the test request are analyzed, that is, the interval duration of the first object content that receives the response result of the test request relative to the sending time of the test request is analyzed, and the interval durations used to characterize adjacent object contents in the response result of the test request are analyzed.

[0099] For example, if the object content in the response result corresponding to the test request is text content, the first delay duration corresponding to each test request is used to represent the interval duration between the first character of the response result received to the test request and the sending time of the test request, and each second delay duration is used to represent the interval duration between two adjacent characters in the response result received to the test request; if the object content in the response result corresponding to the test request is image content, the first delay duration corresponding to each test request is used to represent the interval duration between the first sub-image of the response result received to the test request and the sending time of the test request, and each second delay duration is used to represent the interval duration between two adjacent sub-images in the response result received to the test request; if the object content in the response result corresponding to the test request is audio content, the first delay duration corresponding to each test request is used to represent the interval duration between the first audio frame of the response result received to the test request and the sending time of the test request, and each second delay duration is used to represent the interval duration between two adjacent audio frames in the response result received to the test request.

[0100] Thus, based on the first delay duration and each second delay duration corresponding to each test request, the stress test result of the target large model when the target number is used as the number of concurrent requests is determined.

[0101] In this implementation, when the data output method of the target large model when feeding back the response result is to output in a single object content unit, what it focuses on is the receiving time information of each object content, so that the receiving time of each object content of the response result of each test request is determined as the target receiving time information of the test request. In this way, for different data output methods, the receiving time information it focuses on is determined as the target time receiving information for stress testing analysis, so as to improve the accuracy of the stress testing results for this type of large model.

[0102] Optionally, in one implementation, the above step S1052, based on the first delay duration and each second delay duration corresponding to each test request, determines the stress test result of the target large model when the target number is used as the number of concurrent requests, and may include the following steps:

[0103] Step B1: Determine the first delay duration and each second delay duration corresponding to each test request in the target number as the stress test result of the target large model when the target number is used as the number of concurrent requests.

[0104] In this implementation, the first delay duration and each second delay duration corresponding to each test request in the target number are determined as the stress test result of the target large model when the target number is used as the concurrent request number, that is, the stress test result of each test request of the target number is used as the stress test result of the target large model when the target number is used as the concurrent request number.

[0105] In this implementation, the stress test results of each test request of the target number are counted in detail, and the accuracy of the determined stress test results is improved by improving the detail level of the stress test results.

[0106] Optionally, in one implementation, the above step S1052, based on the first delay duration and each second delay duration corresponding to each test request, determines the stress test result of the target large model when the target number is used as the number of concurrent requests, and may include the following steps:

[0107] Step B2: Calculate first result information based on the first delay duration corresponding to each test request, and calculate second result information based on each second delay duration corresponding to each test request; determine the first result information and the second result information as stress test results of the target large model when the target number is used as the number of concurrent requests;

[0108] Among them, the first result information is used to characterize the delay duration of the sending time relative to the response result receiving time when the target large model uses the target number as the concurrent request number; the second result information is used to characterize the interval duration of adjacent object contents when the response result is received when the target large model uses the target number as the concurrent request number.

[0109] In this implementation, considering that the number of object contents contained in the response result of a test request may be large, when there are multiple test requests, the amount of data required for statistics is large. Therefore, based on the first delay duration corresponding to each test request, the delay duration of the target large model's request sending time relative to the response result receiving time when the target number is used as the concurrent request number can be calculated as the first result information; and based on the second delay durations corresponding to each test request, the interval duration of adjacent object contents when the response result is received when the target large model uses the target number as the concurrent request number is calculated as the second result information. Thus, the first result information and the second result information are determined as the stress test results of the target large model when the target number is used as the concurrent request number.

[0110] Among them, the delay duration represented by the first result information can be the average value of the first delay duration of the target number of test requests calculated based on the mean algorithm, or the standard deviation of the first delay duration of the target number of test requests calculated based on the standard deviation algorithm; the interval duration represented by the second result information can be the average value of the second delay duration of the target number of test requests calculated based on the mean algorithm, or the standard deviation of the second delay duration of the target number of test requests calculated based on the standard deviation algorithm; in this regard, the embodiments of the present application do not make specific limitations. It should be understood that in order to ensure the accuracy of the stress test results, the calculation methods of the two result information must remain consistent.

[0111] In this implementation, by processing the first delay duration and the second delay duration of each test request of the target number, the obtained first result information and the second result information are determined as the stress testing results of the target large model when the target number is used as the number of concurrent requests, thereby reducing the impact of possible extreme stress testing results on the overall stress testing results, thereby improving the objectivity and universality of the obtained stress testing results.

[0112] Optionally, in one implementation, Figure 3 A flow chart of another large model stress testing method provided in an embodiment of the present application is shown as follows: Figure 3 As shown, the above step S104, determining the target receiving time information corresponding to each test request according to the target data output mode, may include the following steps:

[0113] S1042: If the target data output mode is an output mode that uses a single response result itself as an output unit, for each test request, determine the second receiving time of the last object content of the response result of the test request, and obtain the target receiving time information corresponding to each test request;

[0114] Accordingly, the above step S105, based on the target receiving time information corresponding to each test request, analyzes the stress test result of the target large model when the target number is used as the number of concurrent requests, and may include the following steps:

[0115] S1053: for each test request, based on the sending time of the test request and the target receiving time information corresponding to the test request, analyzing the processing time of the target large model to process the test request;

[0116] S1054: Based on the processing time of each test request of the target large model, determine the stress test result of the target large model when the target number is used as the number of concurrent requests.

[0117] In this implementation, if the target data output method is an output method that uses a single response result itself as the output unit, this method focuses on the receiving time information of all object contents contained in the response result. At this time, for each test request, the second receiving time of the last object content of the response result of the test request is determined. Thus, based on the sending time and the second receiving time of the test request, the processing time of the target large model to process the test request is analyzed. Furthermore, based on the processing time of the target large model to process each test request, the stress test result of the target large model when the target number is used as the number of concurrent requests is determined.

[0118] For example, if the object content in the response result corresponding to the test request is text content, then the target receiving time information corresponding to each test request is: the receiving time of the last character of the response result of the test request; if the object content in the response result corresponding to the test request is image content, then the target receiving time information corresponding to each test request is: the receiving time of the last sub-image of the response result of the test request; if the object content in the response result corresponding to the test request is audio content, then the target receiving time information corresponding to each test request is: the receiving time of the last audio frame of the response result of the test request.

[0119] Optionally, in one implementation, the above step S1054, based on the processing time of the target large model to process each test request, determines the stress test result of the target large model when the target number is used as the number of concurrent requests, and may include the following steps:

[0120] Step C: Based on the processing time of each test request processed by the target large model, analyze the actual number of requests processed concurrently within the predetermined period and the average processing time for processing each test request within the predetermined period when the target large model uses the target number as the concurrent request number, and obtain the stress test result of the target large model when the target number is used as the concurrent request number.

[0121] In this implementation, the target data output method is an output method that uses a single response result itself as the output unit. This method focuses on the receiving time information of the entire response result itself. That is to say, when performing stress testing analysis on this type of large model, what is concerned is the processing of the entire request by the large model.

[0122] Therefore, based on the determined processing time of the target large model to process each test request, we can analyze and obtain the actual number of requests concurrently processed by the target large model within the predetermined period when the target number is used as the number of concurrent requests, as well as the average processing time for processing each test request within the predetermined period, and obtain the stress testing results of the target large model when the target number is used as the number of concurrent requests.

[0123] For example, the target number is 10. After determining the processing time of each test request for the target large model, the actual number of requests processed concurrently by the target large model within the predetermined period is 8 when 10 is used as the concurrent request number. The stress test result of the target large model is: when the target number is used as the concurrent request number, the actual number of requests processed concurrently within the predetermined period is 8. At this time, 8<10, and the actual number of requests did not reach the target number, indicating that when 10 is used as the concurrent request number, the target large model is in an overloaded operating state.

[0124] For another example, the target number is 10. After determining the processing time of each test request processed by the target large model, it is analyzed that when the target large model uses 10 as the number of concurrent requests, the average processing time for processing each test request within the predetermined period is 5. The stress test result of the target large model is: when the target number is used as the number of concurrent requests, the average processing time for processing each test request within the predetermined period is 5. At this time, the user determines that the average processing time for processing each test request within the predetermined period is 3 based on his own experience, and 3<5, which means that when 10 is used as the number of concurrent requests, the average processing time for the test request is significantly greater than 3, and the target large model is in an overloaded operating state.

[0125] Among them, the above-mentioned average processing time for processing each test request within the predetermined time period can be understood as the average processing time between each test request being sent from the electronic device to the target large model and receiving the response result of each test request fed back by the target large model. Therefore, the above-mentioned average processing time for processing each test request within the predetermined time period can also be called end-to-end delay (that is, the test request and its response result, the interval time for transmission between the electronic device end and the electronic device equipped with the target large model).

[0126] In this implementation, when the data output mode of the target large model when feeding back the response result is an output mode in which a single response result itself is used as the output unit, what it focuses on is the receiving time information of the entire response result itself, thereby, the second receiving time of the last object content of the response result of the test request is received, and is determined as the target receiving time information of the test request. In this way, for different data output modes, the receiving time information it focuses on is determined as the target time receiving information for stress testing analysis, so as to improve the accuracy of the stress testing results for this type of large model.

[0127] Optionally, in one embodiment, a large model stress testing method provided in an embodiment of the present application may further include the following steps:

[0128] Step D1: Determine whether the stress test end condition is met; if so, execute step D2; if not, execute step D3;

[0129] Step D2: End the stress test on the target large model and output the stress test results;

[0130] Step D3: adjust the target number according to the predetermined adjustment range of the request number, and return to the step of obtaining the target number of test requests for stress testing the target large model.

[0131] In this embodiment, the stress testing end condition is preset, so that by determining whether the target large model meets the above stress testing condition, it is determined whether to continue stress testing the target large model.

[0132] When the conditions are met, the stress test on the target large model is ended, and the analyzed stress test results are output so that the user can understand the load capacity of the target large model based on the stress test results and evaluate its load capacity.

[0133] When it is not satisfied, it indicates that the target number of test requests sent to the target large model during the stress testing process does not cause the target large model to be in an overloaded operating state. At this time, the target number can be adjusted according to the predetermined request number range, and then, return to the above step S101, obtain the target number of test requests for stress testing the target large model, send the adjusted target number of test requests to the target large model in parallel, and stress test the target large model again.

[0134] Optionally, in one implementation, the stress test end condition includes: the target quantity has reached a predetermined test quantity threshold; or, specified content in the analyzed stress test result meets a predetermined content condition.

[0135] In this implementation, different large models can support different numbers of concurrently processed requests, but because the computing resources are fixed, the number of concurrently processed requests that each large model can support has a certain upper limit. Therefore, the above upper limit can be used as a predetermined test quantity threshold, and when the target number has reached the predetermined test quantity threshold, it indicates that the target large model has reached the stress test end condition and the stress test can be ended.

[0136] Furthermore, in order to improve the stress testing efficiency of the large model, it is possible to determine whether the target large model meets the stress testing end condition based on whether the specified content in the analyzed stress testing result meets the predetermined content condition.

[0137] For example, before stress testing a large model, a predetermined time period is preset for the large model, that is, within the predetermined time period after the stress test start time, and the large model is stress tested. In this way, after the time when the test request is first sent to the large model and the time when the test request is sent to the large model again reaches the predetermined time period, it indicates that the large model has reached the stress test end condition and the stress test can be ended.

[0138] For another example, before stress testing a large model, a predetermined feedback duration is pre-set for the large model. The above feedback duration is the interval between the time the request is sent and the time the last object content of the response result corresponding to the request is received. In this way, when the feedback duration of multiple consecutive test requests exceeds the above predetermined feedback duration, it indicates that the target large model may be in an overloaded operating state, and the target large model has reached the stress testing end condition, and the stress testing can be ended.

[0139] In this embodiment, the stress testing process on the target large model is stopped in time through the set stress testing end condition, so as to reduce the time consumed by the stress testing of the large model, thereby improving the stress testing efficiency of the large model.

[0140] Optionally, in an embodiment, in a stress testing method for a large model provided in an embodiment of the present application, before the above step S101, obtaining a target number of test requests for stress testing the target large model, may further include the following steps:

[0141] Step E1: Obtain target command line parameters for the target large model to be stress tested;

[0142] Step E2: Parse the target command line parameters to obtain the specified model parameters of the target large model;

[0143] The specified model parameters include the number of targets and the target data output method.

[0144] In this embodiment, the command line parameters allow the user to customize the behavior of the program. Therefore, the user can pass the data information about the large model data to be stress-tested to the electronic device in the form of command line parameters. In this way, when stress-testing a target large model, by parsing the above command line parameters, the electronic device can obtain various data information about the model, so as to stress-test the target large model based on the data information.

[0145] Therefore, after obtaining the target command line parameters of the target large model to be stress tested, the specified model parameters of the target large model are obtained by parsing the obtained target command line parameters. Thus, through the specified model parameters, the target number of concurrent requests for the target large model and the target data output mode of the target large model can be determined, and then, according to the target data output mode of the target large model, the target receiving time information corresponding to each test request can be determined.

[0146] The target command line parameters may be command line parameters received and input by the user based on the model parameters of the target large model to be stress tested, or command line parameters corresponding to previously adjusted model parameters stored in the electronic device itself. This embodiment of the application does not specifically limit this.

[0147] Optionally, before parsing the target command line parameter, the target command line parameter may be verified to determine the legitimacy and correctness of the data indicated by the target command line parameter, thereby reducing the possibility of errors in the model due to parameter errors.

[0148] Among them, the specified model parameters of the target large model may also include the model name, the average and standard deviation of the input and output tokens, and the timeout duration of the stress test, etc., which are not specifically limited in the embodiments of the present application.

[0149] The so-called input and output tokens refer to the data units processed by the model during the model reasoning process, that is, the object content in the embodiment of the present application. Among them, the input token refers to the data unit received by the model, usually text, image or other forms of input data. In natural language processing, the input token is usually a character in a word, phrase or sentence. For example, when processing a sentence, each word or character can be regarded as an input token. For another example, in image processing, the model needs to divide the image into small blocks (such as pixels or areas), and each small block can be regarded as an input token. For another example, in speech recognition, the speech signal is converted into a series of audio features, each feature can be regarded as an input token. The output token refers to the data unit generated by the model, usually a prediction result or generated text. In the task of natural language processing, the output token can be a character in a word, phrase or sentence. For example, in a machine translation task, each word or phrase in the translated text generated by the model can be regarded as an output token.

[0150] The mean and standard deviation of input and output tokens are limits on the number of contents of the object contents of the input data and output data.

[0151] Optionally, the specified model parameters of the above-mentioned target large model may also include the request types supported for processing. In this way, the electronic device generates a request configuration for a test request for stress testing the target large model based on the model name, input and output tokens, and the request types supported for processing in the obtained specified model parameters of the target large model. Thereby, based on the request configuration, a target number of test requests that meet the request configuration are generated or obtained from a predetermined request database corresponding to the target large model.

[0152] In this embodiment, by parsing the command line parameters to obtain the specified model parameters about the target large model, the electronic device can allocate test requests to the target large model based on the target number in the specified model parameters when performing stress testing on the target large model, and determine the target reception time information of the target large model for each test request according to the target data output method in the specified model parameters, so as to improve the stress testing efficiency and stress testing accuracy of the stress testing results.

[0153] Optionally, in one embodiment, the obtained stress testing results are saved in a storage area pre-set for the target large model, and the stored stress testing results are output after the stress testing of the target large model is completed.

[0154] For example, the stress testing results obtained each time the stress testing method is executed are summarized, and the summary results and the response results of each test request are stored in the result directory specified in the storage area respectively, and the json (Java ScriptObject Notation, JS object notation, is a lightweight data exchange format) module is called to save the above result directory. The above json module is a tool for processing json format data.

[0155] In order to facilitate understanding of a large model stress testing method provided by the present application, a specific embodiment is described below. In this specific embodiment, the target large model is an LLM model, and the object content is text content.

[0156] In the existing technology, the LLM large model is used for business output, and in order to improve the utilization rate of GPU resources, a multi-LoRA integration method is usually adopted to meet different business needs, that is, different LoRAs are responsible for processing different types of requests. When the business demand increases, the performance of the container instance deployed with the LoRA model may become a bottleneck. Therefore, it is urgent to evaluate the instance load capacity through stress testing to realize the evaluation of the load capacity of the large model.

[0157] After investigation, the output of the LLM large model is mainly divided into two categories, namely streaming output and non-streaming output. Among them, the non-streaming output focuses on the number of concurrency per unit time (that is, the actual number of requests processed concurrently within the predetermined period in the embodiment of the present application), the average time consumption of a single task (that is, the average processing time for processing each test request within the predetermined period in the embodiment of the present application); the streaming output focuses on indicators such as the first character delay time (that is, the first delay time in the embodiment of the present application), the delay between characters (that is, the second delay time in the embodiment of the present application), etc.

[0158] Taking into account the different data output methods of the LLM large model when feeding back response results, the large model stress testing method provided in this application needs to have the ability of hybrid stress testing, that is, the stress testing method can be used to stress test both large models with streaming output and large models with non-streaming output.

[0159] By constructing a request distribution stress test LLM for different request types, the requests sent to the LLM model are more in line with the actual business scenario. For the online large model, the historical request data is retrieved from the online database corresponding to the large model, and the number of concurrent requests sent to the large model is set in parallel. The historical requests corresponding to the historical request data of the concurrent request number are sent to the large model to obtain the response results fed back by the large model; for the large model that is not online, a batch of requests is manually constructed according to the request types supported by the large model, and the constructed requests are imported into an EXCEL table (spreadsheet). According to the set number of concurrent requests sent to the large model in parallel, the constructed requests are read from the code in the table and sent to the large model to obtain the response results fed back by the large model. Finally, the target receiving time information of each request is determined, and based on the target receiving time information of each request, the stress test results of the large model at the set number of concurrent requests are analyzed. Finally, based on the analyzed stress test results, an intuitive stress test indicator result interface is provided for users to analyze.

[0160] Combine the following Figure 4 The structural diagram of a specific embodiment shown in the figure specifically describes an electronic device for executing a large model stress testing method provided in an embodiment of the present application, and a stress testing process performed by the electronic device. Figure 4 The electronic device shown includes a command line parameter parsing and setting module 410, an initialization request configuration module 420, a concurrent request sending and result receiving module 430, a result processing and performance measurement module 440 and a report generation and result storage module 450.

[0161] The command line parameter parsing and setting module 410 is used to accept various configuration options about the large model to be stress tested through command line parameters, that is, the specified model parameters in the embodiment of the present application, such as the model name, the average and standard deviation of the input and output tokens, the number of concurrent requests, the timeout duration of the stress test, etc. Then, the specified model parameters are sent to the initialization request configuration module 420. It should be noted that in order to ensure the legitimacy and correctness of the data indicated by the command line parameters, the command line parameters need to be verified before parsing them.

[0162] The initialization request configuration module 420 is used to generate a request configuration for the test request of the large model based on the model name, the number of input and output tokens, etc., included in the specified model parameters after receiving the specified model parameters. Then, the generated request configuration is sent to the concurrent request sending and result receiving module 430.

[0163] The concurrent request sending and result receiving module 430 is used to create a thread pool executor, and after receiving the request configuration, use the submit_request function to generate and submit a test request, wherein each test request is actually sent to the LLM large model through the llm_request function, and receives the response result fed back by the large model. Then, the result of each test request is processed using the process_results function, and then the result of each test request processed is sent to the result processing and performance measurement module 440.

[0164] The result of each test request processed by the process_results function is the target reception time information corresponding to each test request in the embodiment of the present application. Exemplarily, the result of each test request may also include generated text and various performance metrics (such as delay between tokens, total request duration, etc.).

[0165] The result processing and performance measurement module 440 is used to receive the response results of all test requests, calculate various statistical data to obtain results, and then summarize and analyze the results through Pandas Data Frame to calculate the processing time of each test request when processing to different percentiles, the number of input characters and output characters of each test request, and the average, minimum, maximum and standard deviation of the delay time between each character. Then, the calculated data is sent to the report generation and result saving module 450.

[0166] Among them, the statistical data include the number of successful requests, error rate, delays and throughputs at different percentiles, etc. The above statistical data are the data information generated in the response result process of each test request received in the embodiment of the present application; the number of successful requests is the number of requests successfully processed by the large model in the embodiment of the present application; the error rate is the error rate of the large model processing requests in the embodiment of the present application; the delays at different percentiles are the delay durations of each test request in the embodiment of the present application when it is processed to different percentiles; the throughput is the model throughput of the large model in the embodiment of the present application. The above data sent to the report generation and result saving module 450 are the stress test results in the embodiment of the present application.

[0167] The report generation and result saving module 450 is used to save the summary results and the response of each individual request from the received data to the specified result directory, and use the json module to save the results, and append if the file already exists. Then, based on the saved data information, a detailed performance report is generated, including various statistical indicators and charts.

[0168] Exemplarily, taking the case where the object content of the response result is text content, in this case, a single object content is a character. Figure 5-Figure 8 A schematic diagram of the stress test results provided in the embodiment of the present application. Figure 5 and Figure 6 The output mode of the large model is non-streaming output, and the corresponding charts in the performance report are obtained. Figure 7 and Figure 8 The output mode for large models is streaming output, and the corresponding charts in the performance report are obtained. Figure 5-Figure 8 The horizontal axis represents the number of concurrent requests sent in parallel to the large model for processing each time. Figure 5 The vertical axis represents the number of requests completed per minute; Figure 6 The vertical axis represents the average processing time of each request processed by the large model under the number of concurrent requests; Figure 7 The vertical axis represents the delay time of the first character under the number of concurrent requests, that is, the interval time from the sending time of each request to the receiving time of the first character; Figure 8 The vertical axis represents the interval between adjacent characters in the response result of each request under the number of concurrent requests, which is used to measure the speed at which the large model spits out characters.

[0169] It should be noted that during the entire stress testing process, the submit_request function will continuously generate and submit test requests, and end the test after a timeout (that is, the specified content in the stress testing results analyzed in the embodiment of the present application meets the predetermined content conditions) or the maximum number of completed requests is reached (that is, the target number in the embodiment of the present application has reached the predetermined test number threshold) to ensure the operating stability and efficiency of the large model through an appropriate concurrency control mechanism.

[0170] Compared with conventional stress testing methods that focus more on the number of concurrent requests per second and the average request duration, the stress testing method for a large model provided in this embodiment focuses more on the usage scenarios of the large model. According to the streaming output and non-streaming output data output methods, the large model is stress tested using the corresponding stress testing method to evaluate the load capacity of the large model. This helps in the subsequent actual operation of the large model, by controlling the number of concurrent requests, so that the large model is not in an overloaded operating state, thereby achieving the purpose of improving the utilization rate of GPU resources.

[0171] Corresponding to the above method embodiment, the present application embodiment also provides a large model stress testing device, such as Fig. 9 FIG. 1 is a schematic diagram of a structure of a large-scale stress testing device provided in an embodiment of the present application, wherein the device comprises:

[0172] An acquisition module 910 is used to acquire a target number of test requests for stress testing a target large model;

[0173] The sending module 920 is used to send the target number of test requests to the target large model in parallel, so that the target large model responds to the received target number of test requests and feeds back the response result;

[0174] The receiving module 930 is used to receive a response result corresponding to each test request fed back by the target large model; wherein the response result is a result containing one or more object contents;

[0175] The determination module 940 is used to determine the target receiving time information corresponding to each test request according to the target data output mode; wherein the target data output mode is the content presentation mode of the target macro model when feeding back the object content in the response result; the target receiving time information corresponding to each test request is the receiving time information of the object content in the response result corresponding to the test request;

[0176] The analysis module 950 is used to analyze the stress test result of the target large model when the target number is used as the concurrent request number based on the target reception time information corresponding to each test request.

[0177] As can be seen from the above, an embodiment of the present application provides a stress testing method for a large model. First, a target number of test requests for stress testing a target large model are obtained, and a target number of test requests are sent to the target large model in parallel, so that the target large model responds to the received target number of test requests and feeds back the response results. Thus, after receiving the response result corresponding to each test request fed back by the target large model, the target receiving time information corresponding to each test request is determined according to the target data output mode of the target large model when feeding back the response result. Then, based on the target receiving time information corresponding to each test request, the stress testing result of the target large model when the target number is used as the number of concurrent requests is analyzed.

[0178] Based on this, the solution provided by the embodiment of the present application is applied. Under different data output modes, the target large model presents the object content in the feedback response result in different ways. Therefore, the receiving time information of the object content in the response result corresponding to each test request can be determined according to the target data output mode of the target large model, as the target receiving time information corresponding to the test request. Thus, the obtained target receiving time information is used to analyze and obtain the stress test results of the target large model when the target number is used as the number of concurrent requests. Thus, based on the obtained stress test results, stress testing of the large model is implemented, and then, the load capacity of the large model is evaluated, so that data flow control and other operations can be performed based on the load capacity of the large model in the future, thereby further reducing the possibility of the large model being in an overloaded operating state.

[0179] Optionally, in an implementation manner, the determining module 940 is specifically configured to:

[0180] If the target data output mode is a mode of outputting data in units of single object content, for each test request, determining the first receiving time of each object content of the response result of the test request, and obtaining the target receiving time information corresponding to each test request;

[0181] The analysis module 950 includes:

[0182] A first analysis unit is used to analyze, for each test request, a first delay duration and each second delay duration corresponding to the test request based on the sending time of the test request and the first receiving time of each object content of the response result received to the test request; wherein the first delay duration corresponding to the test request is the interval duration between the first object content of the response result received to the test request and the sending time of the test request, and each second delay duration corresponding to the test request is used to characterize the interval duration between adjacent object contents in the response result received to the test request;

[0183] The first determination unit is used to determine the stress test result of the target large model when the target number is used as the number of concurrent requests based on the first delay duration and each second delay duration corresponding to each test request.

[0184] Optionally, in an implementation manner, the first determining unit is specifically configured to:

[0185] Determine the first delay duration and each second delay duration corresponding to each test request in the target number as the stress test result of the target large model when the target number is used as the number of concurrent requests; or, calculate the first result information based on the first delay duration corresponding to each test request, and calculate the second result information based on the each second delay duration corresponding to each test request; determine the first result information and the second result information as the stress test result of the target large model when the target number is used as the number of concurrent requests;

[0186] Among them, the first result information is used to characterize the delay duration of the sending time relative to the response result receiving time when the target large model uses the target number as the concurrent request number, and the second result information is used to characterize the interval duration of adjacent object contents when the response result is received when the target large model uses the target number as the concurrent request number.

[0187] Optionally, in an implementation manner, the determining module 940 is specifically configured to:

[0188] If the target data output mode is an output mode that uses a single response result itself as an output unit, for each test request, determine the second receiving time of the last object content of the response result of the test request, and obtain the target receiving time information corresponding to each test request;

[0189] The analysis module 950 includes:

[0190] A second analysis unit is used to analyze, for each test request, a processing time of the target large model processing the test request based on the sending time of the test request and the target receiving time information corresponding to the test request;

[0191] The second determining unit is used to determine the stress testing result of the target large model when the target number is used as the number of concurrent requests based on the processing time of each test request processed by the target large model.

[0192] Optionally, in an implementation manner, the second determining unit is specifically configured to:

[0193] Based on the processing time of each test request processed by the target large model, the actual number of requests concurrently processed within the predetermined period and the average processing time for processing each test request within the predetermined period are analyzed when the target large model uses the target number as the concurrent request number, and the stress testing result of the target large model when the target number is used as the concurrent request number is obtained.

[0194] Optionally, in an implementation manner, the device further includes:

[0195] The end determination module is used to determine whether the stress test end conditions are met; if they are met, the end module is triggered; if not, the update module is triggered;

[0196] The end module is used to end the stress test on the target large model and output the stress test result;

[0197] The update module is used to adjust the target number according to the predetermined request number adjustment range, and return to the step of obtaining the target number of test requests for stress testing the target large model. Optionally, in one implementation, the device further includes:

[0198] A parameter acquisition module is used to obtain target command line parameters about the target large model to be stress tested before obtaining a target number of test requests for stress testing the target large model; parse the target command line parameters to obtain specified model parameters of the target large model; wherein the specified model parameters include the target number and the target data output method.

[0199] Optionally, in an implementation, the acquisition module 910 is specifically configured to:

[0200] In the predetermined request database corresponding to the target large model, a target number of test requests for stress testing the target large model are obtained; wherein, when the target large model is an online large model, the predetermined request database is a historical request database corresponding to the target large model; when the target large model is an offline large model, the predetermined request database is a request database manually constructed for the target large model.

[0201] The present application also provides an electronic device, such as Fig.10 As shown, it includes a processor 1001, a communication interface 1002, a memory 1003 and a communication bus 1004, wherein the processor 1001, the communication interface 1002, and the memory 1003 communicate with each other through the communication bus 1004, and the memory 1003 is used to store computer programs; the processor 1001 is used to implement the stress testing method of any large model provided in the above-mentioned embodiments of the present application when executing the program stored in the memory 1003.

[0202] The communication bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0203] The communication interface is used for communication between the above terminal and other devices.

[0204] The memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0205] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0206] In another embodiment provided in the present application, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, the stress testing method for a large model described in any of the above embodiments is implemented.

[0207] In another embodiment provided in the present application, a computer program product including instructions is also provided, which, when executed on a computer, enables the computer to execute the large model stress testing method described in any one of the above embodiments.

[0208] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive Solid State Disk (SSD)), etc.

[0209] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0210] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, electronic device embodiment, computer-readable storage medium embodiment, and computer program product embodiment, since they are basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0211] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are included in the protection scope of the present application.

Claims

1. A stress testing method for a large model, characterized in that: The method comprises: Get the target number of test requests for stress testing the target large model; Sending the target number of test requests to the target large model in parallel, so that the target large model responds to the received target number of test requests and feeds back the response results; Receive a response result corresponding to each test request fed back by the target large model; wherein the response result is a result containing one or more object contents; According to the target data output mode, the target receiving time information corresponding to each test request is determined; wherein the target data output mode is the content presentation mode of the target macro model when feeding back the object content in the response result; the target receiving time information corresponding to each test request is the receiving time information of the object content in the response result corresponding to the test request; Based on the target receiving time information corresponding to each test request, the stress test results of the target large model when the target number is used as the number of concurrent requests are analyzed.

2. The method according to claim 1, characterized in that The step of determining target receiving time information corresponding to each test request according to the target data output mode includes: If the target data output mode is a mode of outputting data in units of single object content, for each test request, determining the first receiving time of each object content of the response result of the test request, and obtaining the target receiving time information corresponding to each test request; The target receiving time information corresponding to each test request is analyzed based on the stress test result of the target large model when the target number is used as the number of concurrent requests, including: For each test request, based on the sending time of the test request and the first receiving time of each object content of the response result received to the test request, analyze the first delay duration and each second delay duration corresponding to the test request; wherein the first delay duration corresponding to the test request is used to characterize the interval duration between the first object content of the response result received to the test request and the sending time of the test request, and each second delay duration corresponding to the test request is used to characterize the interval duration between adjacent object contents in the response result received to the test request; Based on the first delay duration and each second delay duration corresponding to each test request, the stress test result of the target large model when the target number is used as the number of concurrent requests is determined.

3. The method according to claim 2, characterized in that The step of determining the stress test result of the target large model when the target number is used as the number of concurrent requests based on the first delay duration and each second delay duration corresponding to each test request includes: Determine the first delay duration and each second delay duration corresponding to each test request in the target number as the stress test result of the target large model when the target number is used as the number of concurrent requests; or, Calculate first result information based on a first delay duration corresponding to each test request, and calculate second result information based on each second delay duration corresponding to each test request; determine the first result information and the second result information as stress test results of the target large model when the target number is used as the number of concurrent requests; Among them, the first result information is used to characterize the delay duration of the sending time relative to the response result receiving time when the target large model uses the target number as the concurrent request number, and the second result information is used to characterize the interval duration of adjacent object contents when the response result is received when the target large model uses the target number as the concurrent request number.

4. The method according to claim 1, characterized in that: The step of determining target receiving time information corresponding to each test request according to the target data output mode includes: If the target data output mode is an output mode that uses a single response result itself as an output unit, for each test request, determine the second receiving time of the last object content of the response result of the test request, and obtain the target receiving time information corresponding to each test request; The target receiving time information corresponding to each test request is analyzed based on the stress test result of the target large model when the target number is used as the number of concurrent requests, including: For each test request, based on the sending time of the test request and the target receiving time information corresponding to the test request, analyzing the processing time of the target large model processing the test request; Based on the processing time of the target large model to process each test request, the stress test result of the target large model when the target number is used as the number of concurrent requests is determined.

5. The method according to claim 4, characterized in that The step of determining the stress test result of the target large model when the target number is used as the number of concurrent requests based on the processing time of each test request processed by the target large model includes: Based on the processing time of each test request processed by the target large model, the actual number of requests concurrently processed within the predetermined period and the average processing time for processing each test request within the predetermined period are analyzed when the target large model uses the target number as the concurrent request number, and the stress testing result of the target large model when the target number is used as the concurrent request number is obtained.

6. The method according to claim 1, characterized in that The method further comprises: Determine whether the stress test end conditions are met; If satisfied, the stress test on the target large model is terminated and the stress test result is output; If not satisfied, the target number is adjusted according to a predetermined adjustment range of the request number, and the process returns to the step of obtaining a target number of test requests for stress testing the target large model.

7. The method according to claim 1, characterized in that Before obtaining the target number of test requests for stress testing the target large model, the method further includes: Get the target command line parameters of the target large model to be stress tested; Parse the target command line parameters to obtain designated model parameters of the target large model; wherein the designated model parameters include the target quantity and the target data output mode.

8. The method according to claim 1, characterized in that The step of obtaining a target number of test requests for stress testing a target large model includes: In a predetermined request database corresponding to the target large model, a target number of test requests for stress testing the target large model are obtained; Among them, when the target big model is an online big model, the predetermined request database is a historical request database corresponding to the target big model; when the target big model is an offline big model, the predetermined request database is a request database manually constructed for the target big model.

9. A large-scale model pressure measurement device, characterized in that: The device comprises: An acquisition module is used to obtain a target number of test requests for stress testing a target large model; A sending module, used for sending the target number of test requests to the target large model in parallel, so that the target large model responds to the received target number of test requests and feeds back the response result; A receiving module, used to receive a response result corresponding to each test request fed back by the target large model; wherein the response result is a result containing one or more object contents; A determination module, used to determine the target reception time information corresponding to each test request according to the target data output mode; wherein the target data output mode is the content presentation mode of the target macro model when feeding back the object content in the response result; the target reception time information corresponding to each test request is the reception time information of the object content in the response result corresponding to the test request; The analysis module is used to analyze the stress test results of the target large model when the target number is used as the number of concurrent requests based on the target reception time information corresponding to each test request.

10. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1 to 8 when executing a program stored in a memory.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • All-in-one machine performance evaluation method and electronic equipment

    CN120950363A

  • All-in-one machine performance evaluation method and electronic device

    CN120950363B