Performance test method and device, electronic equipment and computer readable storage medium
By dynamically generating load and target testing tasks, using load generation models or PID controllers to simulate the real environment, the problem of inaccurate stability evaluation of AI accelerator servers in the prior art is solved, and more accurate performance evaluation and fault detection are achieved.
Patent Information
- Application Number
- CN202510491228.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
AI Technical Summary
Testing with the highest power consumption load in the prior art cannot effectively and accurately evaluate the server stability of the AI accelerator, and cannot truly reflect the performance and stability of the server in actual operation, which may lead to the neglect of potential problems.
By building a load generation model or a PID controller to generate target load dynamically, simulate load changes in the real environment, and accelerate the calculation of target test tasks through the target AI accelerator, obtain performance indicators in real time, and generate target test tasks.
It achieves a more accurate and realistic stability evaluation of AI accelerator server, can dynamically simulate load changes, and obtain performance indicators that are more in line with the real operation, providing an effective data foundation for discovering and solving potential failures.
Smart Images

Figure CN120353678A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of performance testing, and particularly relates to a performance testing method, device, electronic device, and computer-readable storage medium. Background Art
[0002] An artificial intelligence accelerator (also known as an AI accelerator) is hardware used to accelerate the computing tasks of AI models. After deploying an AI accelerator in a server, the server may experience performance degradation or failures due to the huge computations of the AI accelerator. Therefore, it is very important to perform stability testing on the server deployed with the AI accelerator.
[0003] When performing stability testing in related technologies, a load with the highest power consumption is usually applied to the AI accelerator. However, in real application scenarios, the AI accelerator does not always operate at the load with the highest power consumption. Testing with the load of the highest power consumption cannot effectively and accurately evaluate the stability of the server deployed with the AI accelerator. Summary of the Invention
[0004] The present application provides a performance testing method, device, electronic device, and computer-readable storage medium, which at least solves the problem that testing with a load of the highest power consumption in related technologies cannot effectively and accurately evaluate the stability of the server deployed with the AI accelerator.
[0005] The present application provides a performance testing method, including: Generating a target load for the current period according to the period sequence of the current period through a constructed load generation model; or obtaining the target load for the current period through a PID controller according to the target loads and actual loads of each historical period; the target loads of two adjacent periods are different; Generating a target test task to be executed in the current period based on the target load of the current period; Sending a target request to a target server to be tested, where a target AI accelerator is deployed in the target server, and the target request is used to indicate allocating the target AI accelerator for a target AI model to accelerate the process of calculating the execution of the target test task by the target AI model through the target AI accelerator; During the process of accelerating the calculation by the target AI accelerator, obtaining a plurality of performance metrics reported in real time by the target server in the current period.
[0006] The present application also provides a performance testing device, including: A load generation module, configured to generate a target load for the current period according to the period sequence of the current period through a constructed load generation model; or obtain the target load for the current period through a PID controller according to the target loads and actual loads of each historical period; the target loads of two adjacent periods are different; A task generation module, configured to generate a target test task to be executed in the current period based on the target load in the current period; A transmission module, configured to send a target request to a target server to be tested, where a target AI accelerator is deployed in the target server, and the target request is used to instruct to allocate the target AI accelerator to a target AI model, so as to accelerate the process of calculating the execution of the target test task by the target AI accelerator through the target AI accelerator; A data analysis module, configured to obtain multiple performance metrics reported in real time by the target server in the current period during the process of accelerating the calculation by the target AI accelerator.
[0007] This application further provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above performance testing methods when executing the computer program.
[0008] This application further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above performance testing methods are implemented.
[0009] This application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above performance testing methods are implemented.
[0010] Through this application, since the target load in the current period is generated by the constructed load generation model according to the cycle sequence of the current period; or, the target load in the current period is obtained by the PID controller according to the target load and the actual load in each historical period, the target load in each period can be dynamically generated, and the dynamically changing target load is more in line with the load in the real operating environment. Therefore, the problem that the related technology cannot effectively and accurately evaluate the stability of the server deploying the AI accelerator by testing with the load of the highest power consumption can be solved, and the technical effect of being able to truly, effectively and accurately evaluate the stability of each performance metric can be achieved.
[0011] In addition, generating the target test task in the current period through the load in the current period, and the process of accelerating the execution of the target test task is also relatively close to the real operating situation, and the obtained multiple performance metrics are also relatively in line with the real operating results, which can more truly reflect the performance of the system and provide an effective data basis for subsequent discovery and solution of potential faults. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0013] Figure 1 One of the flow schematic diagrams of a performance testing method provided by an embodiment of the present application; Figure 2 The system architecture schematic diagram of a data center provided by an embodiment of the present application; Figure 3 The second flow schematic diagram of a performance testing provided by an embodiment of the present application.
[0014] Figure 4 The structure schematic diagram of a performance testing device provided by an embodiment of the present application. Detailed implementation manners
[0015] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0016] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0017] In the related art, the software pressure state for pressurizing and stability testing of an artificial intelligence accelerator is constant. The artificial intelligence accelerator and the entire artificial intelligence server are always operating at the highest power consumption or a relatively stable power consumption, and the load does not change. It is impossible to simulate the dynamic change situation in the real environment, and it is impossible to truly reflect the stability of the server, so it cannot effectively reflect the performance and stability of the server during actual operation, which may lead to inaccurate evaluation results and may also lead to the neglect of some potential problems, such as performance degradation or instability, attenuation, and thermal failure after load migration.
[0018] To solve the above at least one technical problem, the embodiments of the present application provide a performance testing method, device, electronic device, and computer-readable storage medium.
[0019] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0020] An embodiment of the present application provides a performance testing method, which will be described in detail in combination with the execution process of the performance testing method.
[0021] A performance testing method provided by an embodiment of the present application can be executed by a server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0022] See Figure 1 , which is one of the flow diagrams of a performance testing method provided by an embodiment of the present application, including the following steps 110, 120, 130, and 140.
[0023] Step 111: Generate the target load of this cycle according to the cycle sequence of this cycle through the constructed load generation model; or, obtain the target load of this cycle through a PID controller according to the target load and actual load of each historical cycle; the target loads of two adjacent cycles are different.
[0024] An AI accelerator refers to a hardware device or chip designed specifically to accelerate the computing of artificial intelligence (AI) tasks. Compared with traditional central processing units (CPUs) or graphics processing units (GPUs), AI accelerators usually have an optimized architecture, can process complex calculations in artificial intelligence, machine learning, deep learning and other tasks more efficiently, can significantly improve the processing speed, reduce power consumption, and improve efficiency.
[0025] AI accelerators include, but are not limited to, graphics processing units (GPUs), tensor processing units (TPUs), field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), neural processing units (NPUs), and vision processing units (VPUs).
[0026] Among them, GPUs were originally used for graphics rendering, but due to their powerful parallel computing capabilities, they are widely used in deep learning and artificial intelligence tasks.
[0027] TPU optimizes tensor operations to support the training and inference of deep learning models.
[0028] FPGA chips allow users to reconfigure the circuit according to their needs to support different algorithms and tasks, providing a good balance between flexibility and acceleration capabilities.
[0029] ASICs are designed for specific tasks, with high performance and low power consumption, suitable for large-scale deep learning workloads.
[0030] NPUs are designed specifically for performing neural network operations and are usually integrated in mobile devices and edge computing devices, with the advantages of low power consumption and fast response.
[0031] VPUs focus on the processing of visual data and are commonly used in computer vision applications such as image recognition and video processing.
[0032] The applications of AI accelerators in the training and inference processes are very extensive, especially in fields such as data centers, cloud computing platforms, and intelligent devices.
[0033] Since in the real environment, the load of the AI accelerator is constantly changing, therefore, to imitate the real environment load, the embodiments of the present application generate a dynamically changing target load that changes periodically, that is, the target loads of two adjacent cycles are different.
[0034] The target load refers to the pre-set load level, which is used to estimate the load or pressure borne by the server when the task to be executed is executed, with the purpose of evaluating the performance of the target server under different load conditions.
[0035] The real load refers to the actual load borne by the server when the task is executed, which is affected by various factors such as user behavior and environmental changes.
[0036] The target load is mainly used for testing and evaluation, while the real load reflects the performance in real operation.
[0037] The embodiments of the present application can generate the target load of the current cycle according to the constructed load generation model or through a Proportional-Integral-Derivative (PID) controller.
[0038] Among them, the load generation model is a model created according to the load amplitude, load angular frequency, initial phase, and load linear growth rate, and is used to generate the target load of each cycle based on the cycle order of each cycle.
[0039] The PID controller is a feedback controller commonly used in automatic control systems to regulate the output of the system. In the embodiments of this application, the output of the PID controller is the target load for each period.
[0040] The PID controller generates the target load for the current period based on the load deviation between the target load and the actual load in the historical periods.
[0041] The cycle order of the current period can be input into the constructed load generation model to obtain the target load for the current period output by the load generation model; or the target loads and actual loads of each historical period can be input into the PID controller to obtain the target load for the current period output by the PID controller.
[0042] Step 120: Generate the target test task to be executed in the current period based on the target load for the current period.
[0043] The target test tasks to be executed in the embodiments of this application belong to artificial intelligence tasks, including but not limited to image processing tasks (such as image recognition tasks, image classification tasks, etc.), text processing tasks (such as text classification, question answering tasks, etc.), scientific computing tasks (such as simulating matrix operations, Monte Carlo simulations, etc.), benchmark tasks, etc.
[0044] For image processing tasks, image processing tasks can be generated through ImageNet or COCO datasets, etc.
[0045] For text processing tasks, text classification and question answering tasks can be generated based on the SQuAD or GLUE datasets.
[0046] For scientific computing tasks, such as simulating matrix operations (such as generating large-scale matrix multiplications using NumPy), Monte Carlo simulations, etc.
[0047] Benchmark tasks usually refer to standard tasks for performance evaluation in a certain field or specific environment.
[0048] It can be understood that the loads of tasks of the same task type are relatively similar, and there are significant differences in the loads of tasks of different task types. Usually, each task type has a corresponding load range, that is, there is a correspondence between the task type and the load range, and the load ranges corresponding to different task types are different.
[0049] The target load range to which the target load belongs can be determined from multiple preset load ranges, the task type corresponding to the target load range can be determined, and the target test task to be executed in the current period can be generated according to the task type and the target load.
[0050] The task types of the target test tasks for two adjacent cycles can be the same or different.
[0051] In addition, the memory usage of the target test task can be restricted by cgroups. For example, set the memory usage threshold memory = 32G. When the memory usage is greater than the memory usage threshold, the target test task can be discarded.
[0052] Furthermore, custom scripts can be written based on Locust or Apache JMeter to dynamically adjust the number of concurrent requests and task types.
[0053] Step 130: Send a target request to the target server to be tested. The target AI accelerator is deployed in the target server. The target request is used to indicate allocating the target AI accelerator for the target AI model to accelerate the calculation process of the target AI model executing the target test task.
[0054] The target test task is executed by the target AI model. The model type of the target AI model needs to be the same as the task type of the target test task. The model types include but are not limited to image processing types, text processing models, etc.
[0055] For example, the task types that an image classification model can handle are image classification types, such as classic CNN models like ResNet50, YOLOv5, etc.
[0056] The task types that a text processing model can handle are text processing types, such as models like BERT-Large, GPT-3, etc.
[0057] The target AI model can be located on the target server to be tested or on other servers, such as the main server.
[0058] AI model services can be deployed through TensorFlow Serving or Triton Inference Server.
[0059] The target AI model is the main body for executing the target test task. However, during the process of executing the target test task, a large amount of computing resources are required, and these computing resources can be provided by the target AI accelerator. The target AI accelerator can accelerate the calculation process of the target AI model executing the target test task.
[0060] Accelerated computing refers to the simultaneous use of specialized hardware such as accelerators and optimized algorithms to improve the execution efficiency of computing tasks. It is particularly suitable for applications that need to process large amounts of data and perform complex calculations. It has a wide range of applications in the fields of artificial intelligence, scientific computing, image processing, etc., enabling tasks in these fields to be completed in a shorter time and capable of processing larger-scale data.
[0061] The target AI accelerator can be at least one of the above-mentioned GPUs, TPUs, FPGAs, ASICs, NPUs, and VPUs.
[0062] In the embodiments of the present application, the server where the target AI accelerator is located is the target server to be tested. The target server is usually a computing server. In an artificial intelligence computing center (also known as a data center) or a distributed system, the target server usually belongs to a slave server.
[0063] In a data center or a distributed system, the number of computing servers can be one or more. The AI accelerators in each computing server can be one or more. The AI accelerators in different computing servers can be the same or different.
[0064] As Figure 2 shown, the embodiments of the present application provide a schematic diagram of the system architecture of a data center, including: Switch 1, Switch 2, Switch 3, a main server, and an artificial intelligence computing cluster. The artificial intelligence computing cluster includes multiple computing servers, namely Slave Server 1, Slave Server 2, Slave Server 3, and Slave Server 4. Switch 1 is respectively connected to Switch 2 and Switch 3. Slave Server 1 and Slave Server 2 are respectively connected to Switch 2. Slave Server 3 and Slave Server 4 are respectively connected to Switch 3. The main server is connected to Switch 3. The connections between the servers and the switches, as well as the connections between the switches, can be Ethernet connections. The target server can be any of the slave servers.
[0065] In the embodiments of the present application, the main server tests the target server. Therefore, before the target AI model executes the target test task, the main server will send a target request to the target server through Ethernet. The target request is used to indicate classifying the target AI accelerator for the target AI model, and the target request can be sent through the gRPC / REST interface.
[0066] The target server will allocate a target AI accelerator for the target AI model according to the model type of the target AI model. For example, when the model type is an image processing type, the target AI model is a GPU, VPN, etc.
[0067] During the process of the target AI model executing the target test task, the target AI accelerator accelerates the calculation process of the target AI model executing the target test task, thereby improving the processing efficiency.
[0068] In addition, for some target accelerators, such as GPUs, the NVIDIA MIG technology can be used to divide the GPU into multiple GPU instances to run different tasks respectively.
[0069] Step 140: During the process of the target AI accelerator performing accelerated calculation, obtain multiple performance metrics reported by the target server in real time during this cycle.
[0070] In the process of the target AI accelerator in the embodiment of the present application performing accelerated calculation, it is also a process of performing performance testing on the target server. Therefore, during this process, it is necessary to monitor multiple performance metrics on the target server in real time, and the target server needs to report multiple performance metrics to the main server in real time. Each performance metric is used to evaluate the performance of the target server subsequently. The multiple performance metrics include at least two of CPU temperature, AI accelerator temperature, actual load, CPU memory utilization, AI accelerator memory utilization, request latency, throughput, task success rate, and fault recovery duration.
[0071] The performance testing method provided by the embodiment of the present application can generate the target load of this cycle according to the cycle sequence of this cycle through the constructed load generation model; or, obtain the target load of this cycle through the PID controller according to the target load and actual load of each historical cycle. It can dynamically generate the target load of each cycle, and the target loads of adjacent two cycles are different. The dynamically changing target load is more in line with the load of the real operating environment, which can solve the problem that the related technology cannot effectively and accurately evaluate the stability of the server deploying the AI accelerator by testing with the load of the highest power consumption, and achieve the technical effect of being able to truly, effectively, and accurately evaluate the stability of each performance metric.
[0072] In addition, according to the target load of this cycle, the target test task to be executed in this cycle is generated. The process of the target AI accelerator accelerating the calculation of the target AI model executing the target test task, generating the target test task, and accelerating the execution process of the target test task are also relatively close to the real operating environment. The multiple performance metrics obtained are also more in line with the real operating conditions, which can more truly reflect the performance of the system and provide an effective data basis for discovering and solving potential faults subsequently.
[0073] In some embodiments, generating the target load of this cycle according to the cycle sequence of this cycle through the constructed load generation model includes: The load generation model obtains the first load increment of this cycle according to the load amplitude, load angular frequency, initial phase, and cycle order of this cycle; the first load increment characterizes the magnitude change of the load amplitude; the first load increment changes according to the sine function law with the cycle order. The load generation model obtains the second load increment of this cycle according to the load linear growth rate and the cycle order of this cycle, and the second load increment characterizes the load growth rate. The load generation model performs a summation process on the first load increment, the second load increment, a preset basic load, and load noise to obtain the target load of this cycle.
[0074] The load amplitude A characterizes the peak value of the load fluctuation, with the unit of TFLOPS (floating-point operation volume) or Tasks / s (number of tasks per second). The recommended value range is A ∈ [0.5P_max, 2P_max], where P_max is the maximum load capacity of the server. P_max (maximum load capacity) can be set based on the server hardware specifications (such as GPU computing power, memory bandwidth). For example, the P_max of NVIDIA A100 = 312 TFLOPS. The load amplitude A is usually determined according to the historical peak load. For example, take 1.5 times the actual business peak (A = 1.5 × Ppeak), or it can also be a preset value.
[0075] The load angular frequency ω characterizes the rate of controlling the load change. The load angular frequency ω can be calculated according to formula (1): ω = 2πf (1), where f is the load fluctuation frequency (unit: Hz). In typical scenarios, f ∈ [0.01, 1] Hz, simulating load fluctuations from minutes to seconds. The load angular frequency ω is usually set according to the business cycle. For example, the daily fluctuation corresponds to ω = 2π86400 rad / s (i.e., the cycle is 24 hours). The initial phase φ refers to the initial phase angle, which is used to control the starting position of the waveform and is usually set as φ ∈ [0, 2π].
[0076] The load linear growth rate B characterizes the linear growth rate of the constant load, with the unit of TFLOPS / s. It is used to simulate the cumulative effect of long-term task loads. It is recommended that B ≤ 0.1P_max / s. The load linear growth rate B is usually set according to the business annual growth rate. For example, B = 0.1 × Pmax / s.
[0077] The basic load C, also known as the load offset, is the basic load level, representing the reference load when the system is idle, and is usually set as C = 0.1P_max.
[0078] The load noise N(t) is usually Gaussian white noise, simulating the random fluctuations of the actual load, and is defined as N(t) ~ Normal(0, ), the standard deviation of Gaussian noise σ = 0.2A, simulating the random fluctuation of the actual load.
[0079] The first load L1 can be characterized by the following formula (2): L1 = A·sin(ωt + φ) (2), where A represents the load amplitude; ω represents the load angular frequency, t represents the cycle order of this cycle, and φ represents the initial phase.
[0080] The second load L2 can be characterized by the following formula (3): L2 = B·t (3), where B represents the load linear growth rate.
[0081] Then, the target load L(t) of this cycle can be characterized by the following formula (4): L(t) = L1 + L2 + C + N(t) = A·sin(ωt + φ) + B·t + C + N(t) (4).
[0082] That is, the target load of this cycle is the sum of the first load increment, the second load increment, the preset basic load, and the load noise of this cycle.
[0083] The embodiment of the present application considers the frequency and amplitude of the applied load through the load generation model. The target load of each cycle generated by the load generation model can change sinusoidally and linearly with the cycle order, and can more realistically test the performance of the AI server.
[0084] In some embodiments, the target load of this cycle is obtained by the PID controller according to the target load and the actual load of each historical cycle, including: Determining the load deviation amount between the target load and the actual load of the previous historical cycle by the PID controller, the load integral amount obtained by integrating the load deviation amounts of each historical cycle, and the load differential amount obtained by differentiating the load deviation amounts of each historical cycle; Obtaining the proportional gain, integral gain, and differential gain of the PID controller; Weighting the load deviation amount and the proportional gain, weighting the load integral amount and the integral gain, and weighting the load differential amount and the differential gain by the PID controller to obtain the first weighted load, the second weighted load, and the third weighted load respectively; Performing a summation process on the first weighted load, the second weighted load, and the third weighted load by the PID controller to obtain the target load of this cycle.
[0085] The embodiment of the present application can also adjust the target load of each cycle through the PID controller.
[0086] The PID parameters include three gain parameters, namely the K_p (proportional) gain, the K_i (integral) gain, and the K_d (derivative) gain. The K_p gain affects the rapid adjustment of the load; the K_i gain eliminates long-term errors and ensures the accuracy of the load in the steady state; the K_d gain helps to reduce oscillations and overshoots and increases the stability of the load.
[0087] The values of the K_p gain, the K_i gain, and the K_d gain can be calibrated through experiments. For example, K_p = 0.8, K_i = 0.2, and K_d = 0.05.
[0088] If the load deviation between the target load and the actual load in the previous historical cycle is e(t - 1), then the first weighted load is K_p * e(t - 1); If the load integral obtained by integrating the load deviation of each historical cycle is ∫e(t - 1)dt, then the second weighted load is K_i * ∫e(t - 1)dt; If the load derivative obtained by differentiating the load deviation of each historical cycle is de(t - 1) / d(t - 1), then the third weighted load is K_d * de(t - 1) / d(t - 1).
[0089] Based on the above, the target load of this cycle can be obtained through formula (5): L(t)=K_p * e(t)+K_i * ∫e(t)dt+K_d * de(t) / dt (5).
[0090] Adjusting the target load through the PID controller can ensure that the system adjusts the target load stably, efficiently, and accurately. The PID controller adjusts the control target load through real-time feedback, which can reduce errors, eliminate unstable factors, improve the response speed, and increase the degree of automation.
[0091] In some embodiments, based on the target load of this cycle, a target test task to be executed in this cycle is generated, including: Determine the target load interval to which the target load belongs from a plurality of preset load intervals; Determine the target task type corresponding to the target load interval; each load interval has a task type corresponding to it, and the task types include at least one of an image processing task, a text processing task, and a scientific computing task; Based on the target task type and the target load, generate a target test task to be executed in this cycle.
[0092] As described in the foregoing embodiments, the target loads of the same task type are relatively close, and the target loads of the same task type belong to the same load range, that is, each load range has a corresponding task type. For example: when the target load L(t) > 0.8P_max, the task type is an image processing task, and the image processing task can specifically be high-resolution image processing (such as 512x512). When the target load L(t) < 0.3P_maxL, the task type is a text processing task. When 0.3P_max < target load L(t) < 0.8P_max, the task type is a scientific computing task.
[0093] In the embodiments of the present application, the target load range of the target load can be mapped to a specific target task type, and then, based on the target task type and the target load, a target test task to be executed in the current period is generated, so as to generate a target test task that matches the changing target load, and the adaptability of the target server to different load conditions can be evaluated more effectively.
[0094] In some embodiments, generating a target test task to be executed in the current period based on the target task type and the target load includes: Determining a target container type corresponding to the target task type; each container encapsulates a program for generating a task of the corresponding task type; Based on the target load and the target task type, determining task parameters to be adjusted; Obtaining the container type of the container currently running on the target server; When the container type of the currently running container is the same as the target container type, inputting the task parameters into the currently running container to generate a target test task based on the currently running container; When the container type of the currently running container is different from the target container type, switching the currently running container to a container corresponding to the target container type, and inputting the target test task parameters into the container corresponding to the target container type to generate a target test task based on the container corresponding to the target container type.
[0095] In the embodiments of the present application, containers can be orchestrated for each task type, specifically, Kubernetes can be used to orchestrate containers, and each container encapsulates a program for generating a task of the corresponding task type.
[0096] When generating a target test task, it is necessary to determine whether the currently running container is a container corresponding to the target task type, that is, to determine whether the container type of the currently running container is the same as the target container type corresponding to the target task type.
[0097] When the container type of the currently running container is the same as the target container type, the task parameters can be directly input into the currently running container, so as to directly generate a target test task based on the currently running container.
[0098] The task parameters can be the input resolution, the convolution kernel size, etc.
[0099] When the container type of the currently running container is different from the target container type, it is necessary to switch the currently running container to the container corresponding to the target container type, and then input the target test task parameters into the container corresponding to the target container type, so as to generate a target test task based on the container corresponding to the target container type.
[0100] By setting containers for generating tasks of different task types in the embodiments of the present application, it is possible to isolate the programs for generating each task, ensure the mutual isolation between tasks, and also achieve flexible task management and simplified program deployment.
[0101] In addition, by switching the currently running program, it is possible to quickly switch the task type, improving the elasticity, scalability, and reliability of the system.
[0102] In some embodiments, after sending a target request to the target server to be tested, the method further includes: When a fault injection instruction is obtained, obtaining the fault occurrence frequency of this period; the fault occurrence frequency is the ratio between the total number of faults and the unit time length; Based on the fault occurrence frequency of this period, triggering a target fault to the target server; The target fault is at least one of a hardware fault, a network interruption, and a software crash.
[0103] In the embodiments of the present application, when a fault injection instruction is obtained, a fault can be injected into the target server, and the fault injection instruction is used to indicate injecting a target fault.
[0104] The target fault can be at least one of a hardware fault, a network interruption, and a software crash.
[0105] Among them, the hardware fault includes at least one of an accelerator downclocking and a memory error. The accelerator downclocking is, for example, a GPU downclocking (limiting the GPU power consumption to 200W); the memory error is, for example, injecting a memory read / write error using Memtester.
[0106] The network interruption includes at least one of random packet loss (such as a packet loss rate of 30%) and latency fluctuation (such as setting the latency to 100ms).
[0107] The software crash includes randomly killing a process.
[0108] The above faults can be generated based on the Poisson distribution (such as once every 30 minutes on average), and load-related injection can also be performed. For example, when the load exceeds 0.8P_max, the failure rate is increased to λ(t)=0.01 times / minute.
[0109] First, the fault occurrence frequency of the current period can be obtained. The fault occurrence frequency can be preset or dynamically changed according to the target load. Based on the fault occurrence frequency of the current period, a target fault can be triggered to the target server.
[0110] Introducing a fault injection mechanism in the test to simulate faults such as hardware failures, network interruptions, and software crashes can test the recovery ability and performance of the system, and can effectively evaluate the stability and reliability of the system under real fault conditions.
[0111] In some embodiments, obtaining the fault occurrence frequency of the current period includes: Obtaining the fault frequency increment of the current period based on the target load of the current period and the preset load influence coefficient; Taking the sum of the fault frequency increment of the current period and the preset basic fault occurrence frequency as the fault occurrence frequency of the current period.
[0112] It can be understood that in the real operating environment, the greater the target load, the higher the fault occurrence frequency, and the smaller the target load, the relatively smaller the fault occurrence frequency. Therefore, the fault occurrence frequency of each period can be dynamically adjusted based on the target load of each period.
[0113] Specifically, the fault occurrence frequency λ(t) of the current period can be obtained through the following formula (6): λ(t)=λ_0+α*L(t) (6), where λ_0 represents the preset basic fault occurrence frequency (such as 0.001 times / minute), α represents the load influence coefficient, L(t) represents the target load of the current period, and the product α*L(t) of the load influence coefficient α and the target load L(t) of the current period is the fault frequency increment of the current period.
[0114] The embodiment of the present application adjusts the fault occurrence frequency of the current period according to the target load of the current period, can comprehensively simulate various stress test situations that the system may encounter in actual operation, and can effectively verify the fault tolerance ability of the target server.
[0115] In some embodiments, the multiple performance indicators include the actual recovery duration of the fault; After obtaining the multiple performance indicators reported by the target server in real time in the current period, it further includes: Comparing the target load with the preset load threshold, and obtaining the predicted recovery duration of the target fault based on the comparison result; When the actual recovery duration of the fault is greater than the predicted recovery duration, outputting a first warning, where the first warning is used to indicate that the actual recovery duration is greater than the predicted recovery duration.
[0116] An important indicator for evaluating the fault recovery ability in the embodiments of this application is the actual recovery duration of the fault. Due to different target loads, the fault occurrence frequencies are different, and accordingly, the fault durations are naturally different. Generally, the larger the target load, the longer the actual recovery duration of the fault; on the contrary, the smaller the target load, the smaller the actual recovery duration of the fault.
[0117] In the embodiments of this application, the recovery duration of the fault can be predicted based on the comparison result between the target load and the preset load threshold, and the predicted recovery duration of the fault is called the predicted recovery duration.
[0118] When the actual recovery duration of the fault is not greater than the predicted recovery duration, it indicates that the fault recovery ability of the target server is qualified; when the actual recovery duration of the fault is greater than the predicted recovery duration, it indicates that the fault recovery ability is unqualified. In this case, a first alarm can be output. This first alarm is used to indicate that the actual recovery duration is greater than the predicted recovery duration.
[0119] In the embodiments of this application, the recovery duration of the fault is predicted through the target load to obtain the predicted recovery duration of the fault. Based on the magnitude relationship between the predicted recovery duration of the fault and the actual recovery duration of the fault, the fault recovery ability of the target server can be effectively evaluated. When the actual recovery duration of the fault is greater than the predicted recovery duration, a first alarm is output, and the user can perform potential bottleneck and abnormal behavior analysis based on this first alarm to reduce the probability of the target server having an abnormality after the test is completed.
[0120] In some embodiments, comparing the target load with the preset load threshold and obtaining the predicted recovery duration of the target fault based on the comparison result includes: When the comparison result indicates that the target load is not greater than the preset load threshold, determining the preset minimum recovery duration as the predicted recovery duration; When the comparison result indicates that the target load is greater than the preset load threshold, obtaining a recovery duration increment based on the preset load influence coefficient and the first load difference between the target load and the preset load threshold, and taking the sum value of the recovery duration increment and the preset minimum recovery duration as the predicted recovery duration.
[0121] Referring to formulas (7) and (8), the predicted recovery duration R(f) of the fault can be described as: D_min; L(t) ≤ L_threshold (7), R(f) = { D_min + k·(L(t) - L_threshold); L(t) > L_threshold (8), Among them, D_min represents a preset minimum recovery duration (such as 5 seconds), L_threshold represents a preset load threshold (such as 0.8P_max), and k represents a load impact coefficient (such as k = 0.1 second / TFLOPS). L(t) - L_threshold represents the first load difference between the target load and the preset load threshold, and k·(L(t) - L_threshold) is the increment of the recovery duration affected by the target load.
[0122] The above formula (7) represents that when the target load is not greater than the preset load threshold, the preset minimum recovery duration is determined as the predicted recovery duration.
[0123] The above formula (8) represents that when the target load is greater than the preset load threshold, the sum of the recovery duration increment and the preset minimum recovery duration is used as the predicted recovery duration.
[0124] Based on the comparison result between the target load and the preset load threshold, the embodiment of the present application can relatively accurately predict the recovery duration of the fault and obtain a relatively accurate predicted recovery duration.
[0125] In some embodiments, the multiple performance metrics include the actual load; After obtaining the multiple performance metrics reported in real time by the target server in this period, the method further includes: Determine the second load difference between the actual load and the target load in this period; Obtain the error percentage between the target load and the actual load according to the second load difference and the actual load; When the error percentage is greater than the error percentage threshold, output a second warning; the second warning is used to indicate that the error percentage between the target load and the actual load is greater than the error percentage threshold.
[0126] When the target AI accelerator performs acceleration calculation, the error percentage between the target load and the actual load of the target server is also an important indicator for evaluating whether there is a potential anomaly in the target server.
[0127] Specifically, the error percentage between the target load and the actual load can be verified through a small-scale pre-test, and the ratio of the second load difference between the target load and the actual load in this period to the actual load can be used as the error percentage between the target load and the actual load.
[0128] When the error percentage is greater than the error percentage threshold (a preset value, such as ±5%), output a second warning; the second warning is used to indicate that the error percentage between the target load and the actual load is greater than the error percentage threshold.
[0129] Embodiments of the present application provide strong evidence for determining whether a target server has a fault by determining the error percentage between the target load and the actual load. When the error percentage is greater than the percentage threshold, a second alarm is output, and the user can analyze potential bottlenecks and abnormal behaviors based on this second alarm, reducing the probability of the target server having an abnormality after the test is completed.
[0130] In some embodiments, after obtaining multiple performance metrics reported in real time by the target server in this cycle, the method further includes: Based on the maximum metric threshold and the minimum metric threshold of each performance metric, normalize each performance metric respectively to obtain the normalized performance metric; Obtain the metric weight of each performance metric; Perform weighted processing on each normalized performance metric and the metric weight to obtain the evaluation value of each performance metric; Perform summation processing on the evaluation values of each performance metric to obtain the total evaluation value of the performance; When the total evaluation value of the performance is less than the evaluation value threshold, output a third alarm, and the third alarm is used to indicate that the total evaluation value of the performance is less than the evaluation value threshold; The performance metrics include at least two of CPU temperature, AI accelerator temperature, actual load, CPU memory utilization, AI accelerator memory utilization, request latency, throughput, task success rate, and fault recovery duration.
[0131] Embodiments of the present application can calculate the total evaluation value of each performance metric. The total evaluation value S can be represented by the following formula (9): S = Σ(Wi·Di) (9), where, wi represents the metric weight of each performance metric (the metric weights of each metric can be determined by the Analytic Hierarchy Process AHP), and Di is the normalized performance metric.
[0132] Di can be obtained based on the following formula (10): Di = (Xi - Xmin) / (Xmax - Xmin) (10), where, Xi represents the actual performance metric, Xmin represents the minimum metric threshold, and Xmax represents the maximum metric threshold.
[0133] Refer to Table 1 below. Table 1 shows the values corresponding to the metric weights and the normalization ranges (minimum metric threshold and maximum metric threshold) of multiple performance metrics including temperature (temp), power consumption (power), task success rate (success), and recovery time (recovery) in a certain scenario.
[0134]
[0135] Based on Table 1 above, the overall evaluation value S can be determined as S = 0.3D_temp + 0.25D_power + 0.25D_success + 0.2D_recovery, where D_temp represents the normalized temperature, D_power represents the normalized power consumption, D_success represents the normalized task success rate, and D_recovery represents the normalized recovery time.
[0136] It can be understood that the higher the overall evaluation value of the performance, the better the performance of the target server, and the lower the overall score of the performance, the worse the performance of the target server.
[0137] In addition, after obtaining the overall evaluation value of the performance, it is also necessary to compare the overall score of the performance with the evaluation value threshold (preset value). When the overall evaluation value of the performance is not less than the evaluation value threshold, it indicates that the overall performance of the target server is relatively good. When the overall evaluation value of the performance is less than the evaluation value threshold, it indicates that the overall performance of the target server is relatively poor. In this case, a third warning can be output, and the third warning is used to indicate that the overall evaluation value of the performance is less than the evaluation value threshold.
[0138] In the embodiments of the present application, the overall performance of the target server is evaluated through the overall evaluation value of each performance index. When the overall performance is relatively poor, the user can analyze potential bottlenecks and abnormal behaviors based on this second warning, reducing the probability of the target server having an abnormality after the test is completed.
[0139] In some embodiments, after obtaining each performance index of each cycle in the embodiments of the present application, each performance index and log can be stored and analyzed. The data analysis tool can be used to evaluate each performance index and log, generate a detailed stability report, and output data visualization charts and conclusions.
[0140] For data storage, a Telegraf agent is deployed in each computing server. The Telegraf agent is used to collect performance metrics and push them to the database InfluxDB, which is sharded by timestamp (such as one shard per hour).
[0141] During the push process, Apache Kafka can be used to transmit performance metrics in real time. One Kafka is used to transmit one type of performance metric, and the transmission through Kafka can ensure low latency (<1 second).
[0142] For logs, the ELKStack (Elasticsearch + Logstash + Kibana) can be used to store fault logs and task logs.
[0143] In addition, the analysis results include curves corresponding to each indicator. Grafana can be used to configure a dynamic dashboard to display real-time actual load curves, target load curves, temperature heat maps, and fault event markers.
[0144] Threshold alarms can also be set for each performance indicator. For example, when the GPU temperature > 85°C, an email notification is triggered.
[0145] In some embodiments, data analysis can also use the Isolation Forest algorithm to identify abnormal points in temperature and power consumption, predict load and temperature changes in the next 10 minutes based on the LSTM model, call Pandas through Python scripts to calculate key indicators (such as MTBF, mean time to recovery), generate a PDF report, and obtain analysis results such as scatter plots of load-temperature relationships and fault impact timelines drawn by Matplotlib for visualization charts.
[0146] In some embodiments, based on the analysis results, system optimization suggestions can be output, such as suggesting to replace with an efficient cooling solution, adding redundant hardware to improve fault recovery capabilities, etc.
[0147] Specifically, when there is a heat dissipation problem in the target server, it can output a suggestion to deploy a liquid cooling system, reducing the thermal resistance R_th from 0.05°C / W to 0.03°C / W; when the computing resources are insufficient, it can output a suggestion to add a spare GPU card to achieve failover through NvLink; when there is a load balancing problem, tasks can be dynamically allocated based on the real-time load (such as using Kubernetes HPA for automatic scaling). The checkpoint function can also be implemented to resume from the most recent state after the task is interrupted; the GPU driver parameters can also be adjusted, such as adjusting the Max Clock and Power Limit to balance performance and power consumption; the kernel parameters can also be adjusted to optimize the TCP buffer size to improve network throughput.
[0148] See Figure 3 , the embodiments of the present application provide the following steps: Step 301: Generate the target load for this period according to the period sequence of this period through the constructed load generation model; or, obtain the target load for this period through a PID controller based on the target load and actual load of each historical period; Step 302: Determine the target task type corresponding to the target load interval to which the target load belongs; Step 303: Determine the target container type corresponding to the target task type; each container encapsulates a program for generating tasks of the corresponding task type; Step 304: Determine the task parameters to be adjusted based on the target load and the target task type; Step 305: Obtain the container type of the currently running container on the target server; Step 306: Determine whether the container type of the currently running container is the same as the target container type; if so, execute Step 307; if not, execute Step 308; Step 307: Input the task parameters into the currently running container to generate a target test task based on the currently running container; then execute Step 309; Step 308: Switch the currently running container to the container corresponding to the target container type, input the target test task parameters into the container corresponding to the target container type, and generate a target test task based on the container corresponding to the target container type; then execute Step 309; Step 309: Send a target request to the target server to be tested. There is a target AI accelerator deployed in the target server. The target request is used to indicate allocating the target AI accelerator for the target AI model, so as to accelerate the process of calculating the execution of the target test task by the target AI accelerator; Step 310: Obtain multiple performance metrics reported in real time by the target server in this cycle; Step 311: Analyze the multiple performance metrics reported in each cycle to obtain an analysis result; Step 312: Output optimization suggestions based on the analysis result.
[0149] For the detailed execution processes of Steps 301 to 312, refer to the foregoing embodiments and will not be elaborated here.
[0150] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0151] See Figure 4 , the embodiments of the present application also provide a structural schematic diagram of a performance testing device. The performance testing device includes: A load generation module 410, configured to generate a target load for this cycle according to the cycle sequence of this cycle through the constructed load generation model; or obtain the target load for this cycle through a PID controller according to the target load and the actual load of each historical cycle; the target loads of two adjacent cycles are different; A task generation module 420, configured to generate a target test task to be executed in this cycle based on the target load of this cycle; A transmission module 430 is configured to send a target request to a target server to be tested. A target AI accelerator is deployed in the target server. The target request is used to instruct to allocate the target AI accelerator for a target AI model, so as to accelerate the process of calculating the execution of a target test task by the target AI accelerator using the target AI accelerator; A data analysis module 440 is configured to obtain a plurality of performance metrics reported in real time by the target server in the current period during the process of accelerating the calculation by the target AI accelerator.
[0152] In some embodiments, the load generation module 410 specifically includes: A first generation sub-module is configured to: Obtain a first load increment of the current period through a load generation model according to a load amplitude, a load angular frequency, an initial phase, and a period sequence of the current period; the first load increment characterizes the magnitude change of the load amplitude; the first load increment changes according to a sine function law with the period sequence; Obtain a second load increment of the current period through the load generation model according to a load linear growth rate and the period sequence of the current period, where the second load increment characterizes the load growth rate; Perform a summation process on the first load increment, the second load increment, a preset base load, and a load noise through the load generation model to obtain a target load of the current period.
[0153] In some embodiments, the load generation module 410 specifically includes: A second generation sub-module is configured to: Determine a load deviation amount between a target load and an actual load in the previous historical period through a PID controller, a load integral amount obtained by performing an integral process on the load deviation amounts of each historical period, and a load differential amount obtained by performing a differential process on the load deviation amounts of each historical period; Obtain a proportional gain, an integral gain, and a differential gain of the PID controller; Perform weighting on the load deviation amount and the proportional gain, on the load integral amount and the integral gain, and on the load differential amount and the differential gain through the PID controller, respectively, to obtain a first weighted load, a second weighted load, and a third weighted load; Perform a summation process on the first weighted load, the second weighted load, and the third weighted load through the PID controller to obtain a target load of the current period.
[0154] In some embodiments, the task generation module 420 is specifically configured to: Determine a target load interval to which the target load belongs from a plurality of preset load intervals; Determine a target task type corresponding to the target load interval; each load interval has a task type corresponding thereto, and the task type includes at least one of an image processing task, a text processing task, and a scientific computing task; Generate a target test task to be executed in the current cycle based on the target task type and target load.
[0155] In some embodiments, the performance testing device further includes: a fault injection module configured to: when a fault injection instruction is obtained, obtain the fault occurrence frequency in the current cycle; the fault occurrence frequency is the ratio between the total number of faults and the unit time duration; based on the fault occurrence frequency in the current cycle, trigger a target fault to the target server; the target fault is at least one of a hardware fault, a network interruption, and a software crash.
[0156] In some embodiments, the fault injection module is specifically further configured to: obtain a fault frequency increment in the current cycle based on the target load in the current cycle and a preset load impact coefficient; use the sum of the fault frequency increment in the current cycle and a preset basic fault occurrence frequency as the fault occurrence frequency in the current cycle.
[0157] In some embodiments, the multiple performance metrics include the actual recovery duration of the fault; the data analysis module 440 is configured to: compare the target load with a preset load threshold, and obtain a predicted recovery duration of the target fault based on the comparison result; when the actual recovery duration of the fault is greater than the predicted recovery duration, output a first alert, and the first alert is used to indicate that the actual recovery duration is greater than the predicted recovery duration.
[0158] In some embodiments, the data analysis module 440 is configured to: when the comparison result indicates that the target load is not greater than the preset load threshold, determine the preset minimum recovery duration as the predicted recovery duration; When the comparison result indicates that the target load is greater than the preset load threshold, obtain a recovery duration increment based on the preset load impact coefficient and a first load difference between the target load and the preset load threshold, and use the sum of the recovery duration increment and the preset minimum recovery duration as the predicted recovery duration.
[0159] In some embodiments, the multiple performance metrics include the actual load; the data analysis module 440 is further configured to: determine a second load difference between the actual load in the current cycle and the target load; obtain an error percentage between the target load and the actual load based on the second load difference and the actual load; when the error percentage is greater than an error percentage threshold, output a second alert; the second alert is used to indicate that the error percentage between the target load and the actual load is greater than the error percentage threshold.
[0160] In some embodiments, the data analysis module 440 is further configured to: respectively perform normalization processing on each performance metric based on the maximum metric threshold and the minimum metric threshold of each performance metric to obtain a normalized performance metric; obtain the metric weight of each performance metric; Perform weighted processing on each normalized performance metric and metric weight to obtain the evaluation value of each performance metric; perform sum processing on the evaluation values of each performance metric to obtain the total evaluation value of the performance; In the case where the total evaluation value of the performance is less than the evaluation value threshold, output a third warning, where the third warning is used to indicate that the total evaluation value of the performance is less than the evaluation value threshold; the performance metrics include at least two of CPU temperature, AI accelerator temperature, actual load, CPU memory utilization, AI accelerator memory utilization, request latency, throughput, task success rate, and fault recovery duration.
[0161] For the description of the features in the embodiments corresponding to the performance testing device, reference can be made to the relevant descriptions in the embodiments corresponding to the performance testing method, which will not be elaborated here one by one.
[0162] An embodiment of the present application further provides an electronic device, including a memory and a value processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned performance testing method embodiments.
[0163] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-mentioned performance testing method embodiments when running.
[0164] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk, or optical disc, etc., various media that can store computer programs.
[0165] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-mentioned performance testing method embodiments.
[0166] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-mentioned performance testing method embodiments.
[0167] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered as exceeding the scope of this application.
[0168] The above has introduced in detail a performance testing method, device, electronic device, and computer-readable storage medium provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A performance testing method, characterized in that, Including: Generating the target load of the current period according to the period sequence of the current period through the constructed load generation model; or, obtaining the target load of the current period through a PID controller according to the target loads and actual loads of each historical period; the target loads of two adjacent periods are different; Generating the target test task to be executed in the current period based on the target load of the current period; Sending a target request to the target server to be tested, where a target AI accelerator is deployed in the target server, and the target request is used to instruct to allocate the target AI accelerator for the target AI model to accelerate the process of calculating the execution of the target test task by the target AI model through the target AI accelerator; During the process of accelerating the calculation by the target AI accelerator, obtaining multiple performance metrics reported by the target server in real time in the current period.
2. The performance testing method according to claim 1, characterized in that, The generating the target load of the current period according to the period sequence of the current period through the constructed load generation model includes: Obtaining the first load increment of the current period through the load generation model according to the load amplitude, load angular frequency, initial phase, and the period sequence of the current period; the first load increment characterizes the magnitude change of the load amplitude; the first load increment changes according to the sine function law with the period sequence; Obtaining the second load increment of the current period through the load generation model according to the load linear growth rate and the period sequence of the current period, and the second load increment characterizes the load growth rate; Performing a summation process on the first load increment, the second load increment, a preset basic load, and load noise through the load generation model to obtain the target load of the current period.
3. The performance testing method according to claim 1, wherein The obtaining the target load of the current period through a PID controller according to the target loads and actual loads of each historical period includes: Determining, through the PID controller, the load deviation amount between the target load and the actual load of the previous historical period, the load integral amount obtained by integrating the load deviation amounts of each historical period, and the load differential amount obtained by differentiating the load deviation amounts of each historical period; Obtaining the proportional gain, integral gain, and differential gain of the PID controller; Respectively obtaining the first weighted load, the second weighted load, and the third weighted load by weighting the load deviation amount and the proportional gain, weighting the load integral amount and the integral gain, and weighting the load differential amount and the differential gain through the PID controller; Performing a summation process on the first weighted load, the second weighted load, and the third weighted load through the PID controller to obtain the target load of the current period.
4. The performance testing method according to claim 1, characterized in that The generating the target test task to be executed in the current period based on the target load of the current period includes: Determining the target load interval to which the target load belongs from a preset plurality of load intervals; Determining the target task type corresponding to the target load interval; each load interval has a task type corresponding to it; the task type includes at least one of an image processing task, a text processing task, and a scientific computing task; Generate a target test task to be executed in the current cycle based on the target task type and the target load.
5. The performance testing method according to claim 1, wherein After sending the target request to the target server to be tested, the method further includes: When a fault injection instruction is obtained, obtain the fault occurrence frequency in the current cycle; the fault occurrence frequency is the ratio between the total number of faults and the unit time duration; Based on the fault occurrence frequency in the current cycle, trigger a target fault to the target server; The target fault is at least one of a hardware fault, a network interruption, and a software crash.
6. The performance testing method according to claim 5, characterized in that The obtaining the fault occurrence frequency in the current cycle includes: Obtain the fault frequency increment in the current cycle based on the target load in the current cycle and a preset load impact coefficient; Use the sum of the fault frequency increment in the current cycle and a preset basic fault occurrence frequency as the fault occurrence frequency in the current cycle.
7. The performance testing method according to claim 5 or 6, characterized in that, The multiple performance indicators include the actual recovery duration of the fault; After obtaining the multiple performance indicators reported in real time by the target server in the current cycle, it further includes: Compare the target load with a preset load threshold, and obtain the predicted recovery duration of the target fault based on the comparison result; When the actual recovery duration of the fault is greater than the predicted recovery duration, output a first warning, where the first warning is used to indicate that the actual recovery duration is greater than the predicted recovery duration.
8. The performance testing method according to claim 7, wherein The comparing the target load with a preset load threshold and obtaining the predicted recovery duration of the target fault based on the comparison result includes: When the comparison result indicates that the target load is not greater than the preset load threshold, determine the preset minimum recovery duration as the predicted recovery duration; When the comparison result indicates that the target load is greater than the preset load threshold, obtain a recovery duration increment based on the preset load impact coefficient and the first load difference between the target load and the preset load threshold, and use the sum of the recovery duration increment and the preset minimum recovery duration as the predicted recovery duration.
9. The performance testing method according to claim 1, characterized in that The multiple performance indicators include the actual load; Determine the second load difference between the actual load and the target load in the current cycle; Based on the second load difference and the actual load, obtain the error percentage between the target load and the actual load; When the error percentage is greater than the error percentage threshold, output a second warning; The second warning is used to indicate that the error percentage between the target load and the actual load is greater than the error percentage threshold.
10. The performance testing method according to claim 1, wherein After obtaining the multiple performance indicators reported in real time by the target server in the current cycle, the method further includes: Based on the maximum index threshold and the minimum index threshold of each performance indicator, perform normalization processing on each performance indicator respectively to obtain the normalized performance indicators; Obtain the index weight of each performance indicator; Perform weighted processing on each normalized performance indicator and the index weight to obtain the evaluation value of each performance indicator; Perform a sum processing on the evaluation values of each performance indicator to obtain the total evaluation value of the performance; When the total evaluation value of the performance is less than the evaluation value threshold, output a third warning, where the third warning is used to indicate that the total evaluation value of the performance is less than the evaluation value threshold; The performance metrics include at least two of CPU temperature, AI accelerator temperature, actual load, CPU memory utilization, AI accelerator memory utilization, request latency, throughput, task success rate, and fault recovery duration.
11. A performance testing device, characterized in that, Comprising: A load generation module, configured to generate a target load for the current period according to the period sequence of the current period through the constructed load generation model; or, obtain the target load for the current period through a PID controller according to the target loads and actual loads of each historical period; the target loads of two adjacent periods are different; A task generation module, configured to generate a target test task to be executed in the current period based on the target load of the current period; A transmission module, configured to send a target request to a target server to be tested, where a target AI accelerator is deployed in the target server, and the target request is used to instruct to allocate the target AI accelerator to a target AI model to accelerate the process of calculating the execution of the target test task by the target AI model through the target AI accelerator; A data analysis module, configured to obtain a plurality of performance metrics reported in real time by the target server in the current period during the process of accelerating the calculation by the target AI accelerator.
12. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to implement the steps of the performance test method according to any one of claims 1 to 10 when executing the computer program.
13. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the performance test method according to any one of claims 1 to 10 when executed by a processor.
Citation Information
Cited By
Performance test method and device of storage equipment, electronic equipment and storage medium
CN120564815A
Load prediction-based server management method, program product and device
CN120950342A
A server management method, program product, and device based on load prediction
CN120950342B
Accelerator card testing method and system and electronic equipment
CN122470459A