Performance evaluation method, computer program product, electronic device, and storage medium
By obtaining the computing power, storage and network design parameters of the server cluster and calculating the performance evaluation results of the simulated training tasks, the problems of time-consuming and high-cost cluster performance evaluation are solved, and fast and low-cost cluster performance evaluation is achieved.
Patent Information
- Application Number
- CN202510838342.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-06-20
AI Technical Summary
Cluster performance evaluation in actual server clusters is time-consuming and expensive, making it difficult to achieve lightweight and fast evaluation, which increases the cost of design and hardware deployment.
By obtaining the computing power, storage and network design parameters of the server cluster, determining the simulation training task, calculating the computing performance of the server cluster in executing the simulation training task, and obtaining the performance evaluation results.
It enables lightweight and rapid evaluation of server cluster performance, significantly reducing time and overhead, lowering designers' research and hardware deployment costs, and improving design efficiency.
Smart Images

Figure CN120371674B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of performance evaluation, and in particular to a performance evaluation method, a computer program product, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the development of artificial intelligence (AI) technology, the advantages of neural network models have gradually become apparent, and various fields have begun to invest heavily in the research of larger neural network models. However, due to the high computing and storage requirements of training, the difficulty of training neural network models has increased exponentially. To solve the training problem of neural network models, the construction of large-scale intelligent computing server clusters has become the main solution. However, as the scale of server clusters increases, the time and cost of cluster performance evaluation also increase rapidly. How to reduce the time and cost of cluster performance evaluation is currently a pressing issue. Summary of the Invention
[0003] The present application provides a performance evaluation method, a computer program product, an electronic device, and a computer-readable storage medium to at least solve the problem in the related art that cluster performance evaluation in an actual server cluster is time-consuming and expensive.
[0004] This application provides a performance evaluation method for a server cluster. The performance evaluation method includes:
[0005] Obtain the computing power design parameters of the server cluster;
[0006] Obtain storage design parameters for the server cluster;
[0007] Obtain network design parameters of the server cluster;
[0008] Determine the simulation training task, and calculate the computing performance of the server cluster in executing the simulation training task based on the computing power design parameters, storage design parameters, and network design parameters to obtain the performance evaluation results.
[0009] The present application also provides a computer program product, applied to a server cluster, comprising:
[0010] The computing power parameter collection module is used to obtain the computing power design parameters of the server cluster;
[0011] Storage parameter collection module, used to obtain storage design parameters of the server cluster;
[0012] Network parameter collection module, used to obtain network design parameters of the server cluster;
[0013] The performance evaluation module is used to determine the simulation training task and calculate the computing performance of the server cluster in executing the simulation training task based on the computing power design parameters, storage design parameters and network design parameters to obtain the performance evaluation results.
[0014] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned performance evaluation methods when executing the computer program.
[0015] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned performance evaluation methods are implemented.
[0016] Through this application, by determining the simulation training task and calculating the computing performance of the server cluster in executing the simulation training task based on the computing power design parameters, storage design parameters and network design parameters of the server cluster, the performance evaluation result is obtained. Therefore, the technical problem of time-consuming and high-cost cluster performance evaluation in actual server clusters can be solved, and the performance of the server cluster can be evaluated lightly and quickly, which significantly reduces time and cost, and can reduce the research cost of designers and the actual hardware deployment cost, thereby improving the design efficiency of the server cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 A flow chart of a performance evaluation method provided in an embodiment of the present application;
[0019] Figure 2 A schematic diagram of the interconnection topology of a server cluster provided in an embodiment of the present application;
[0020] Figure 3 A flow chart of a performance evaluation method provided in an embodiment of the present application;
[0021] Figure 4 A flow chart of a performance evaluation method provided in an embodiment of the present application;
[0022] Figure 5 A flow chart of a performance evaluation method provided in an embodiment of the present application;
[0023] Figure 6A flow chart of a performance evaluation method provided in an embodiment of the present application;
[0024] Figure 7 A flow chart of a performance evaluation method provided in an embodiment of the present application;
[0025] Figure 8 A schematic diagram of a computer program product provided in accordance with an embodiment of the present invention;
[0026] Figure 9 A schematic diagram of a module of an electronic device provided in an embodiment of the present application;
[0027] Figure 10 A schematic diagram of a computer-readable storage medium provided in accordance with an embodiment of the present application. DETAILED DESCRIPTION
[0028] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0029] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0030] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0031] See also Figure 1 and Figure 2 , an embodiment of the present application provides a performance evaluation method, and the method is described in detail in conjunction with the execution process of the performance evaluation method.
[0032] The performance evaluation method is applied to the server cluster 100, and the performance evaluation method includes:
[0033] 010: Obtain computing power design parameters of server cluster 100;
[0034] 020: Obtain storage design parameters of the server cluster 100;
[0035] 030: Obtain network design parameters of the server cluster 100;
[0036] 040: Determine the simulation training task, and calculate the computing performance of the server cluster 100 in executing the simulation training task based on the computing power design parameters, storage design parameters, and network design parameters to obtain a performance evaluation result.
[0037] In the performance evaluation method of the embodiment of the present application, by determining the simulation training task and calculating the computing performance of the server cluster 100 in executing the simulation training task based on the computing power design parameters, storage design parameters and network design parameters of the server cluster 100, a performance evaluation result is obtained. Therefore, the technical problem of time-consuming and high-cost cluster performance evaluation in the actual server cluster 100 can be solved, and the performance of the server cluster 100 can be evaluated lightly and quickly, which significantly reduces time and cost, and can reduce the research cost of designers and the actual hardware deployment cost, thereby improving the design efficiency of the server cluster 100.
[0038] Specifically, the server cluster 100 may be an intelligent computing server cluster 100. Designers may design the server cluster 100 according to actual application requirements and obtain computing power design parameters, storage design parameters, and network design parameters of the server cluster 100 before constructing the actual server cluster 100.
[0039] The designer designs parameters for the computing power of server cluster 100, also known as computing power design parameters. These computing power design parameters include parameters related to the computation of server cluster 100. The designer also designs parameters related to the storage of server cluster 100, also known as storage design parameters. The designer also designs parameters for the network of server cluster 100, also known as network design parameters. These network design parameters include parameters related to the communication of server cluster 100.
[0040] The computing power design parameters, storage design parameters and network design parameters can all use the corresponding parameters in the server 10 product performance manual, or the corresponding parameters obtained by performance testing of the actual server 10 product, or the parameters directly input by the designer.
[0041] The corresponding simulated training task is determined based on the application requirements of the server cluster 100, for example, based on the training task actually required to be executed by the server cluster 100. After obtaining the computing power design parameters, storage design parameters, and network design parameters, the computing performance of the server cluster 100 in executing the simulated training task can be calculated based on the computing power design parameters, storage design parameters, and network design parameters to obtain a performance evaluation result.
[0042] It should be noted that the performance evaluation results are calculated using the parameters of the server cluster 100 and do not require the actual server cluster 100 to actually execute the simulation training task.
[0043] After obtaining the performance evaluation results, if the performance evaluation results are unqualified, the designer can adjust the server cluster 100 according to the performance evaluation results to improve the performance evaluation results; if the performance evaluation results are qualified, the actual server cluster 100 deployment can be carried out according to the current server cluster 100 design.
[0044] In this way, by comprehensively considering the computing, storage, and communication factors through the computing power design parameters, storage design parameters, and network design parameters of the server cluster 100, the performance of the server cluster 100 can be evaluated lightly and quickly; compared with actual execution or simulation evaluation, it can significantly reduce time and overhead, and can reduce the research costs of designers and the actual hardware deployment costs, thereby improving the design efficiency of the server cluster 100.
[0045] The interconnection topology of the server cluster 100 can be as follows: Figure 2 As shown, server cluster 100 includes multiple servers 10 and multiple network switches. Each server 10 may include multiple processing units 11 and corresponding multiple storage units 12. Storage units 12 include memory 121 and cache. Each processing unit 11 is connected to a corresponding memory 121, and the cache can be provided within the processing unit 11.
[0046] The processing unit 11 may be a graphics processing unit (GPU). The memory 121 may store computational data generated by the processing unit 11. A cache is a buffer area for data exchange and is a relatively small, high-speed temporary storage device. Directly reading data from the memory 121 by the processing unit 11 requires a certain waiting period. Therefore, the data to be used may be stored in the cache to reduce the waiting time of the processing unit 11.
[0047] Multiple processing units 11 within the same server 10 are interconnected using high bandwidth within the server 10, such as NVIDIA Virtual Link (NVLINK). Processing units 11 between different servers 10 are interconnected via a network switch.
[0048] Two processing units 11 on different servers 10 are interconnected via a network switch. Each two network switches connected to processing units 11 are also interconnected via a network switch. The network switch can adopt a non-blocking architecture, such as a clos architecture. A non-blocking architecture ensures efficient interconnection between multiple processing units 11.
[0049] For Figure 2 The server cluster 100 shown faces a variety of parameters that need to be considered during construction, such as the bandwidth of the network switch, the computing performance of the processing unit 11, the bandwidth of the memory 121, the performance of the cache, etc. According to the parameters that need to be considered in actual applications, the specific computing power design parameters, storage design parameters and network design parameters that need to be obtained can be determined.
[0050] See also Figure 2 In some embodiments, the server cluster 100 includes multiple servers 10, at least one of which includes at least one processing unit 11. The computing power design parameters include any one or more of the number of servers 10, the number of processing units 11 in the server 10, the computing power of the processing units 11, the average computing power coefficient of the processing units 11, and the variance of the actual computing power of the processing units 11.
[0051] Specifically, computing power design parameters may include:
[0052] y: the number of servers 10;
[0053] x: the number of processing units 11 in the server 10;
[0054] C: computing power of the processing unit 11;
[0055] λ: average computing power coefficient of the processing unit 11;
[0056] : The variance of the actual computing capability of the processing unit 11.
[0057] The computing power of a processing unit 11 refers to the computing power of a single processing unit 11. The average computing power coefficient of a processing unit 11 refers to the computing power that the processing unit 11 can actually exert, and can be, for example, 0.5, 0.6, etc. The variance of the actual computing power of a processing unit 11 can be calculated based on the actual computing power of multiple processing units 11.
[0058] See also Figure 2In some embodiments, the server cluster 100 includes multiple servers 10, at least one server 10 includes at least one storage unit 12, and the storage design parameters include any one or more of a cache consistency access identification parameter, a hit rate of the storage unit 12, an access latency of the storage unit 12, a cache line size of the storage unit 12, and a bandwidth of the storage unit 12.
[0059] Specifically, the storage design parameters may include:
[0060] U: cache consistency access identification parameter;
[0061] : hit rate of storage unit 12;
[0062] : Access delay of storage unit 12;
[0063] : cache line size of storage unit 12;
[0064] : Bandwidth of storage unit 12.
[0065] Cache coherence refers to the ability of the caches of multiple processing units to accurately and consistently reflect the latest state of shared memory data. i = 1, 2, 3, 4, representing the number of levels in storage unit 12, with level 1 being L1 cache, level 2 being L2 cache, level 3 being L3 cache, and level 4 being memory 121.
[0066] Specifically, it can represent the hit rate of the i-th level in the storage unit 12, p1 represents the hit rate of L1 cache; p2 represents the hit rate of L2 cache; p3 represents the hit rate of L3 cache; p4 represents the hit rate of memory 121, and p4=1.
[0067] Specifically, it can represent the access delay of the i-th level in the storage unit 12, Indicates the access latency of L1 cache; Indicates the access latency of L2 cache; Indicates the access latency of L3 cache; Indicates the access latency of the memory 121 .
[0068] Specifically, it can represent the bandwidth of the i-th level in the storage unit 12, Indicates the bandwidth of L1 cache; Indicates the bandwidth of L2 cache; Indicates the bandwidth of L3 cache; Indicates the bandwidth of memory 121.
[0069] See also Figure 2 In some embodiments, the server cluster 100 includes multiple servers 10, and the network design parameters include any one or more of the bandwidth, latency, failure probability, and packet loss probability of the Internet within the server 10 and the bandwidth, latency, failure probability, and packet loss probability of the Internet between the servers 10.
[0070] Specifically, network design parameters may include:
[0071] The bandwidth of the internet in the server 10;
[0072] The latency of the internet within the server 10;
[0073] Failure probability of the interconnection network within the server 10;
[0074] packet loss probability of the internet within the server 10;
[0075] The bandwidth of the Internet network between the 10 servers;
[0076] The latency of the interconnection network between the 10 servers;
[0077] Failure probability of the interconnection network between 10 servers;
[0078] Packet loss probability of the interconnected network between 10 servers.
[0079] See also Figure 2 and Figure 3 In some embodiments, a simulation training task is determined, and based on the computing power design parameters, storage design parameters, and network design parameters, the computing performance of the server cluster 100 in executing the simulation training task is calculated to obtain a performance evaluation result (i.e., 040), including:
[0080] 041: Determine the simulation training task and the total computing amount and total data volume of the simulation training task;
[0081] 042: Calculate the execution time of at least one iteration of the simulation training task by the server cluster 100 based on the computing power design parameters, the storage design parameters, and the network design parameters;
[0082] 043: Determine the performance evaluation results based on the execution time.
[0083] Specifically, the corresponding simulation training task is determined based on the application requirements of the server cluster 100, and the total computing amount required for the simulation training task and the total amount of data to be processed by the simulation training task are determined. For example, if the server cluster 100 is used for model training, the corresponding simulation training task can be determined, and the corresponding total computing amount and total data amount can be determined.
[0084] It is understood that the model training process includes multiple iterations. In the embodiment of the present application, the execution time of at least one iteration of the simulation training task performed by the server cluster 100 is calculated based on the acquired computing power design parameters, storage design parameters, and network design parameters.
[0085] In one example, the execution time of one iteration of the simulation training task performed by the server cluster 100 is calculated. Of course, the total execution time of two, three or more iterations of the simulation training task performed by the server cluster 100 can also be calculated, which is not limited here.
[0086] Afterwards, the performance evaluation result of the server cluster 100 can be determined based on the execution time. For example, a time threshold can be set. If the execution time is greater than or equal to the time threshold, the performance evaluation result of the server cluster 100 is determined to be unqualified, and the designer can adjust the server cluster 100 based on the performance evaluation result to improve the performance evaluation result. If the execution time is less than the time threshold, the performance evaluation result of the server cluster 100 is determined to be qualified, and the designer can actually deploy the server cluster 100 based on the current server cluster 100 design.
[0087] In this way, the performance evaluation of the server cluster 100 can be achieved through lightweight calculations.
[0088] See also Figure 2 and Figure 4 In some embodiments, determining the simulation training task and the total computational amount and total data amount (ie, 041) of the simulation training task includes:
[0089] 0411: Construct simulation training tasks based on neural network models;
[0090] 0412: Determine the total computational effort and total data volume of the simulation training task by analyzing the neural network model.
[0091] Specifically, if the application requirement of server cluster 100 is to train a neural network model, a simulated training task can be constructed based on the neural network model. In some embodiments, the neural network model is a language model, and the server cluster 100 is required to process a distributed AI computing task. Therefore, a simulated training task can be constructed based on any language model, and the simulated training task can be a distributed AI computing task.
[0092] For the selected neural network model, the total computational load and the total data volume can be determined directly using the data in the literature or model configuration file corresponding to the neural network model; the total computational load and the total data volume can also be determined by analyzing the neural network model through a performance testing tool.
[0093] In this way, the performance evaluation result is closer to the evaluation result obtained by the server cluster 100 actually executing the training task, thereby improving the reliability of the performance evaluation result.
[0094] During at least one iteration of the simulation training task executed by the server cluster 100, the execution time includes the computation time of the processing unit 11, the time the processing unit 11 spends reading data from the corresponding storage unit 12, the time the processing unit 11 spends synchronizing data from other storage units 12 in the same server 10, and the time the processing unit 11 spends synchronizing data from the storage units 12 of the processing units 11 in other servers 10. The execution time calculation process is described in detail below.
[0095] See also Figure 2 and Figure 5 In some embodiments, a server cluster 100 includes multiple servers 10, at least one of which includes at least one processing unit 11 and corresponding at least one storage unit 12. Calculating the execution time (i.e., 042) of at least one iteration of a simulation training task performed by the server cluster 100 based on computing power design parameters, storage design parameters, and network design parameters includes:
[0096] 0421: Calculate, based on the computing power design parameters, a first expected value for the server cluster 100 to execute at least one iteration of the simulation training task, where the first expected value is the expected computing time of at least one processing unit 11;
[0097] 0422: Calculate, based on the computing power design parameters and the storage design parameters, a second expected value for the server cluster 100 to execute at least one iteration of the simulation training task, where the second expected value is an expected reading time for at least one processing unit 11 to read a corresponding storage unit 12;
[0098] 0423: Calculate a third expected value of at least one iteration of the simulated training task performed by the server cluster 100 based on the computing power design parameters, the storage design parameters, and the network design parameters, where the third expected value is an expected synchronization time for at least one processing unit 11 to synchronize with other storage units 12 in the same server 10;
[0099] 0424: Calculate, based on the computing power design parameters, the storage design parameters, and the network design parameters, a fourth expected value for the server cluster 100 to perform at least one iteration of the simulation training task, where the fourth expected value is an expected synchronization time for at least one processing unit 11 to synchronize with at least one storage unit 12 in other servers 10;
[0100] 0425: Calculate the execution time according to the first expected value, the second expected value, the third expected value and the fourth expected value.
[0101] Specifically, taking the execution of one iteration of a simulation training task as an example, based on the computing power design parameters, a first expected value of the server cluster 100 executing one iteration of the simulation training task can be calculated. The first expected value is the expected computing time of at least one processing unit 11. The first expected value can represent the computing time of one iteration of the simulation training task.
[0102] It should be noted that when the server 10 includes multiple processing units 11, the multiple processing units 11 can execute computing tasks in parallel. Therefore, the computing time of the multiple processing units 11 depends on the computing time of the slowest processing unit 11, and the expected computing time refers to the expected computing time of the slowest processing unit 11.
[0103] In some embodiments, calculating a first expected value of at least one iteration of the simulated training task by the server cluster 100 based on the computing power design parameter, where the first expected value is an expected computing time of at least one processing unit 11, includes:
[0104] According to the computing power design parameters, the Gumbel extreme value distribution is used to calculate the first expected value of the server cluster 100 executing at least one iteration of the simulation training task. The first expected value is the expected computing time of at least one processing unit 11.
[0105] Because the computational time of processing unit 11 is affected by performance fluctuations such as load imbalance, heat dissipation and frequency reduction, hardware heterogeneity, and communication contention, the computational time of multiple processing units 11 is random. Therefore, the computational time of at least one processing unit 11 is modeled using the Gumbel extreme value distribution to calculate the first expected value. In this way, the randomness of the working performance of processing unit 11 is taken into account, and the first expected value under performance fluctuations can be calculated, making the calculated execution time closer to the actual execution time, thereby making the performance evaluation result more accurate.
[0106] Based on the computing power design parameters and the storage design parameters, a second expected value for the server cluster 100 to execute one iteration of the simulated training task can be calculated. The second expected value is the expected time for at least one processing unit 11 to read the corresponding storage unit 12. Each processing unit 11 corresponds to a storage unit 12, and the storage unit 12 corresponding to the processing unit 11 is called a local storage unit. During the execution of one iteration of the simulated training task, each processing unit 11 can read the local storage unit one or more times.
[0107] According to the computing power design parameters, storage design parameters and network design parameters, the third expected value of the server cluster 100 performing one iteration of the simulation training task is calculated. The third expected value is the expected synchronization time of at least one processing unit 11 synchronizing with other storage units 12 in the same server 10.
[0108] The storage units 12 corresponding to other processing units 11 in the same server 10 may all be referred to as first remote storage units. The third expected value is an expected synchronization time for at least one processing unit 11 to synchronize with the first remote storage unit.
[0109] According to the computing power design parameters, storage design parameters and network design parameters, the fourth expected value of the server cluster 100 performing one iteration of the simulation training task is calculated and analyzed. The fourth expected value is the expected synchronization time of at least one processing unit 11 synchronizing at least one storage unit 12 in other servers 10.
[0110] The storage units 12 corresponding to the other processing units 11 in the other servers 10 may all be referred to as second remote storage units. The fourth expected value is an expected synchronization time for at least one processing unit 11 to synchronize with the second remote storage unit.
[0111] Then, the execution time can be calculated based on the first expected value, the second expected value, the third expected value, and the fourth expected value. In one example, the execution time can be obtained by adding the first expected value, the second expected value, the third expected value, and the fourth expected value.
[0112] In this way, various time consumptions when the server cluster 100 executes the training task are taken into consideration, and the calculated execution time is more reliable.
[0113] See also Figure 2 and Figure 6 In some embodiments, calculating the execution time (i.e., 042) of at least one iteration of the simulation training task by the server cluster 100 based on the computing power design parameters, the storage design parameters, and the network design parameters includes:
[0114] 0426: Determine whether the server cluster 100 supports cache coherent access based on storage design parameters;
[0115] 0427: When the server cluster 100 does not support cache consistency access, the first time calculation formula is used to calculate the execution time of the server cluster 100 to perform at least one iteration of the simulation training task based on the computing power design parameters, storage design parameters and network design parameters.
[0116] Specifically, whether the server cluster 100 supports cache coherence access may be determined according to the cache coherence access identification parameter in the storage design parameters.
[0117] When the server cluster 100 does not support cache consistency access, when the processing unit 11 reads data in the first remote storage unit and the second remote storage unit, it is necessary to move the data in the first remote storage unit and the second remote storage unit to the local storage unit and then read from the local storage unit.
[0118] In the case where the server cluster 100 supports access to the consistent storage unit 12 , the processing unit 11 can directly access and read data in the first remote storage unit and the second remote storage unit.
[0119] When the server cluster 100 does not support cache consistency access, the first time calculation formula can be used to calculate the execution time of the server cluster 100 to perform at least one iteration of the simulation training task based on the computing power design parameters, storage design parameters and network design parameters.
[0120] In some embodiments, the first time calculation formula is:
[0121] ;
[0122] in, is the execution time, is the total computational effort, is the number of processing units 11 in the server 10, is the number of servers 10, is the average computing capacity coefficient of the processing unit 11, is the computing capability of the processing unit 11, is the variance of the actual computing capability of the processing unit 11, is the first distribution coefficient, is the second distribution coefficient, is the total data volume, is the cache line size of storage unit 12, is the probability of occurrence of a hit on the i-th level of storage unit 12, is the bandwidth of the i-th level of storage unit 12, is the access delay of the i-th level of storage unit 12, is the bandwidth of the Internet in the server 10, is the failure probability of the internet in server 10, is the packet loss probability of the Internet in server 10, is the delay of the Internet in server 10, is the bandwidth of the Internet network between the 10 servers, is the failure probability of the interconnected network between the 10 servers, is the packet loss probability of the 10 interconnected networks of servers, is the delay of the Internet among the 10 servers.
[0123] Specifically, Indicates the first expected value. , y, C, λ and The first expected value is calculated based on the Gumbel extreme value distribution. . It can represent the expected computation time of each processing unit 11, , and As follows:
[0124] ;
[0125] ;
[0126] in, Represents the distribution function of the standard normal distribution, where e is a natural constant.
[0127] Represents the second expected value.
[0128] 、 、 Design parameters for storage. represents the transmission time of the processing unit 11 reading data from the local storage unit, Indicates the number of communications by which the processing unit 11 reads data. It represents the additional read delay caused by miss each time data is read from the storage unit 12 .
[0129] , which represents the probability of an L1 cache hit; , represents the probability of L2cache hit; , represents the probability of an L3 cache hit; Indicates the probability of occurrence when memory 121 hit occurs.
[0130] Represents the third expected value.
[0131] 、 、 、 、 、 、 and Design parameters for the network. It indicates the actual transmission time of moving the data of the first remote storage unit to the local storage unit in the case of link failure and packet loss and retransmission. Indicates the transmission time under ideal handling conditions, It represents the probability of zero failure and zero packet loss calculated using the Gumbel extreme value distribution.
[0132] In this way, taking into account the link failures and packet loss retransmissions during the transmission process, as well as the randomness of the occurrence of these situations, the actual transmission time is calculated using the Gumbel extreme value distribution, making the calculated execution time closer to the actual execution time, thereby making the performance evaluation results more accurate.
[0133] It represents the time taken by the processing unit 11 to read data from the local storage unit. The calculation method is similar to the second expected value and will not be repeated here.
[0134] Represents the fourth expected value.
[0135] Indicates the actual transmission time taken to move data from the second remote storage unit to the local storage unit in the event of a link failure and packet loss and retransmission. Indicates the transmission time under ideal handling conditions, It represents the probability of zero failure and zero packet loss calculated using the Gumbel extreme value distribution.
[0136] In this way, taking into account the link failures and packet loss retransmissions during the transmission process, as well as the randomness of the occurrence of these situations, the actual transmission time is calculated using the Gumbel extreme value distribution, making the calculated execution time closer to the actual execution time, thereby making the performance evaluation results more accurate.
[0137] It represents the time taken by the processing unit 11 to read data from the local storage unit. The calculation method is similar to the second expected value and will not be repeated here.
[0138] The execution time can be obtained by adding the first expected value, the second expected value, the third expected value and the fourth expected value.
[0139] See also Figure 2 and Figure 7 In some embodiments, calculating the execution time (i.e., 042) of at least one iteration of the simulation training task by the server cluster 100 based on the computing power design parameters, the storage design parameters, and the network design parameters includes:
[0140] 0428: When the server cluster 100 supports cache consistency access, the second time calculation formula is used to calculate the execution time of the server cluster 100 to perform at least one iteration of the simulation training task based on the computing power design parameters, storage design parameters and network design parameters.
[0141] In some embodiments, the second time calculation formula is:
[0142] .
[0143] Specifically, in the embodiment of the present application, the calculation process of the first expected value and the second expected value is the same as the calculation process of the first expected value and the second expected value in the first time calculation formula, and will not be repeated here.
[0144] In the case of supporting cache coherence access, the processing unit 11 can directly access and read the data in the first remote storage unit, and compared with the first time calculation formula, the actual transmission time of moving the data from the first remote storage unit to the local storage unit can be omitted.
[0145] 、 、 、 、 、 、 and Design parameters for the network. It represents the actual transmission time consumed by the processing unit 11 to read data from the first remote storage unit in the case of link failure and packet loss and retransmission. Indicates the transmission time under ideal reading conditions. It represents the probability of zero failure and zero packet loss calculated using the Gumbel extreme value distribution.
[0146] In this way, taking into account the link failures and packet loss and retransmission caused by remote network instability during the reading process, as well as the randomness of the occurrence of these situations, the actual reading time is calculated using the Gumbel extreme value distribution, making the calculated execution time closer to the actual execution time, thereby making the performance evaluation results more accurate.
[0147] It represents the time taken by the processing unit 11 to read data from the local storage unit. The calculation method is similar to the second expected value and will not be repeated here.
[0148] Represents the fourth expected value.
[0149] It represents the actual transmission time consumed by the processing unit 11 to read data from the second remote storage unit in the case of link failure and packet loss and retransmission. Indicates the transmission time under ideal reading conditions. It represents the probability of zero failure and zero packet loss calculated using the Gumbel extreme value distribution.
[0150] In this way, taking into account the link failures and packet loss and retransmission caused by remote network instability during the reading process, as well as the randomness of the occurrence of these situations, the actual reading time is calculated using the Gumbel extreme value distribution, making the calculated execution time closer to the actual execution time, thereby making the performance evaluation results more accurate.
[0151] It represents the time taken by the processing unit 11 to read data from the local storage unit. The calculation method is similar to the second expected value and will not be repeated here.
[0152] The execution time can be obtained by adding the first expected value, the second expected value, the third expected value and the fourth expected value.
[0153] In the embodiment of the present application, different time calculation formulas are used to calculate the execution time depending on whether the server cluster 100 supports cache coherent access. This takes into account the difference in execution time between the server cluster 100 supporting cache coherent access and the server cluster 100 not supporting cache coherent access, making the calculated execution time closer to the actual execution time, thereby making the performance evaluation results more accurate.
[0154] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0155] See also Figure 8, an embodiment of the present application further provides a computer program product 200, which is applied to the server cluster 100. The computer program product 200 includes a computing power parameter collection module 210, a storage parameter collection module 220, a network parameter collection module 230 and a performance evaluation module 240. The computing power parameter collection module 210 is used to obtain the computing power design parameters of the server cluster 100. The storage parameter collection module 220 is used to obtain the storage design parameters of the server cluster 100. The network parameter collection module 230 is used to obtain the network design parameters of the server cluster 100. The performance evaluation module 240 is used to determine the simulation training task, and calculate the computing performance of the server cluster 100 in executing the simulation training task based on the computing power design parameters, storage design parameters and network design parameters, and obtain a performance evaluation result.
[0156] In the computer program product of the embodiment of the present application, the computing power parameter collection module 210, the storage parameter collection module 220, and the network parameter collection module 230 are used to obtain computing power design parameters, storage design parameters, and network design parameters respectively. The performance evaluation module 240 determines the simulation training task and calculates the computing performance of the server cluster 100 in executing the simulation training task based on the computing power design parameters, storage design parameters, and network design parameters of the server cluster 100 to obtain the performance evaluation result. Therefore, the technical problem of time-consuming and high overhead in cluster performance evaluation in the actual server cluster 100 can be solved, and the performance of the server cluster 100 can be evaluated in a lightweight and fast manner, which significantly reduces time and overhead, and can reduce the research cost of designers and the actual hardware deployment cost, thereby improving the design efficiency of the server cluster 100.
[0157] In certain embodiments, the performance evaluation module 240 is specifically used to determine the simulation training task and the total computing amount and total data amount of the simulation training task; based on the computing power design parameters, storage design parameters and network design parameters, calculate the execution time of the server cluster 100 to perform at least one iteration of the simulation training task; and determine the performance evaluation result based on the execution time.
[0158] In some embodiments, the performance evaluation module 240 is specifically used to construct a simulation training task based on a neural network model; by analyzing the neural network model, the total computational amount and total data amount of the simulation training task are determined.
[0159] In some embodiments, the server cluster 100 includes a plurality of servers 10 , and at least one server 10 includes at least one processing unit 11 and corresponding at least one storage unit 12 . The performance evaluation module 240 is specifically used to calculate the first expected value of the server cluster 100 performing at least one iteration of the simulation training task based on the computing power design parameters, and the first expected value is the calculation expected time of at least one processing unit 11; calculate the second expected value of the server cluster 100 performing at least one iteration of the simulation training task based on the computing power design parameters and the storage design parameters, and the second expected value is the expected reading time of at least one processing unit 11 reading the corresponding storage unit 12; calculate the third expected value of the server cluster 100 performing at least one iteration of the simulation training task based on the computing power design parameters, the storage design parameters and the network design parameters, and the third expected value is the synchronization expected time of at least one processing unit 11 synchronizing with other storage units 12 in the same server 10; calculate the fourth expected value of the server cluster 100 performing at least one iteration of the simulation training task based on the computing power design parameters, the storage design parameters and the network design parameters, and the fourth expected value is the synchronization expected time of at least one processing unit 11 synchronizing with at least one storage unit 12 in other servers 10; calculate the execution time based on the first expected value, the second expected value, the third expected value and the fourth expected value.
[0160] In some embodiments, the performance evaluation module 240 is specifically used to determine whether the server cluster 100 supports cache consistency access based on storage design parameters; if the server cluster 100 does not support cache consistency access, the first time calculation formula is used to calculate the execution time of the server cluster 100 to perform at least one iteration of the simulation training task based on the computing power design parameters, storage design parameters and network design parameters.
[0161] In some embodiments, the performance evaluation module 240 is specifically used to calculate the execution time of the server cluster 100 for at least one iteration of the simulation training task based on computing power design parameters, storage design parameters and network design parameters when the server cluster 100 supports cache consistency access, using a second time calculation formula.
[0162] For descriptions of features in the embodiment corresponding to the computer program product 200, reference may be made to the relevant descriptions of the embodiment corresponding to the performance evaluation method, which will not be detailed here.
[0163] See also Figure 9 An embodiment of the present application further provides an electronic device 300, comprising a memory 310 and a processor 320, wherein the memory 310 stores a computer program, and the processor 320 is configured to run the computer program to execute the steps in any one of the above-mentioned performance evaluation method embodiments.
[0164] See also Figure 10 An embodiment of the present application further provides a computer-readable storage medium 400, in which a computer program 410 is stored, wherein the computer program 410 is configured to execute the steps of any of the above-mentioned performance evaluation method embodiments when running.
[0165] In an exemplary embodiment, the computer-readable storage medium 400 may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0166] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned performance evaluation method embodiments are implemented.
[0167] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned performance evaluation method embodiments are implemented.
[0168] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0169] The above describes in detail a performance evaluation method, a computer program product 200, an electronic device 300, and a computer-readable storage medium 400 provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only intended to help understand the method and core ideas of the present application. It should be noted that, for those skilled in the art, without departing from the principles of the present application, several improvements and modifications may be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A performance evaluation method, characterized in that: Applied to a server cluster, the performance evaluation method includes: Obtaining computing power design parameters of the server cluster; Obtaining storage design parameters of the server cluster; Obtaining network design parameters of the server cluster; Determine a simulation training task, and calculate the computing performance of the server cluster in executing the simulation training task based on the computing power design parameters, the storage design parameters, and the network design parameters to obtain a performance evaluation result; The determining of the simulation training task and calculating the computing performance of the server cluster in executing the simulation training task based on the computing power design parameters, the storage design parameters, and the network design parameters to obtain a performance evaluation result include: Determining the simulation training task and the total computational amount and total data amount of the simulation training task; Calculating, based on the computing power design parameters, the storage design parameters, and the network design parameters, the execution time of the server cluster performing at least one iteration of the simulation training task; Determine the performance evaluation result according to the execution time; The server cluster includes a plurality of servers, at least one of the servers includes at least one processing unit and at least one corresponding storage unit, and calculating, based on the computing power design parameter, the storage design parameter, and the network design parameter, the execution time of the server cluster for executing at least one iteration of the simulation training task, including: Calculating, based on the computing power design parameters, a first expected value for the server cluster to execute at least one iteration of the simulation training task, where the first expected value is an expected computing time of at least one of the processing units; Calculating, based on the computing power design parameter and the storage design parameter, a second expected value for the server cluster to execute at least one iteration of the simulation training task, where the second expected value is an expected reading time for at least one of the processing units to read the corresponding storage unit; Calculate, based on the computing power design parameter, the storage design parameter, and the network design parameter, a third expected value for the server cluster to execute at least one iteration of the simulation training task, where the third expected value is an expected synchronization time of at least one processing unit synchronizing with other storage units in the same server; Calculating, based on the computing power design parameter, the storage design parameter, and the network design parameter, a fourth expected value for the server cluster to perform at least one iteration of the simulation training task, the fourth expected value being an expected synchronization time of at least one of the processing units synchronizing with at least one of the storage units in the other servers; The execution time is calculated according to the first expected value, the second expected value, the third expected value, and the fourth expected value.
2. The performance evaluation method according to claim 1, wherein: The server cluster includes multiple servers, at least one of which includes at least one processing unit, and the computing power design parameters include any one or more of the number of servers, the number of processing units in the servers, the computing power of the processing units, the average computing power coefficient of the processing units, and the variance of the actual computing power of the processing units.
3. The performance evaluation method according to claim 1, wherein: The server cluster includes multiple servers, at least one of the servers includes at least one storage unit, and the storage design parameters include any one or more of a cache consistency access identification parameter, a hit rate of the storage unit, an access latency of the storage unit, a cache line size of the storage unit, and a bandwidth of the storage unit.
4. The performance evaluation method according to claim 1, wherein: The server cluster includes multiple servers, and the network design parameters include any one or more of the bandwidth, delay, failure probability, and packet loss probability of the intra-server Internet and the bandwidth, delay, failure probability, and packet loss probability of the inter-server Internet.
5. The performance evaluation method according to claim 1, wherein: Determining the simulation training task and the total computational amount and total data amount of the simulation training task includes: Constructing the simulation training task based on the neural network model; By analyzing the neural network model, the total computational workload and total data volume of the simulation training task are determined.
6. The performance evaluation method according to claim 5, characterized in that: The neural network model is a language model, and the simulation training task is a distributed artificial intelligence computing task.
7. The performance evaluation method according to claim 1, wherein: The calculating, based on the computing power design parameter, the storage design parameter, and the network design parameter, the execution time of the server cluster for executing at least one iteration of the simulation training task includes: determining whether the server cluster supports consistent storage unit access based on the storage design parameters; In the case that the server cluster does not support consistent storage unit access, the first time calculation formula is used to calculate the execution time of the server cluster performing at least one iteration of the simulation training task based on the computing power design parameters, the storage design parameters and the network design parameters.
8. The performance evaluation method according to claim 7, characterized in that: The first time calculation formula is: ; in, is the execution time, is the total amount of computation, is the number of processing units in the server, is the number of servers, is the average computing capacity coefficient of the processing unit, is the computing power of the processing unit, is the variance of the actual computing power of the processing unit, is the first distribution coefficient, is the second distribution coefficient, is the total data volume, is the cache line size of the storage unit, is the probability of the i-th level hit of the storage unit, is the bandwidth of the i-th level of the storage unit, is the access delay of the i-th level of the storage unit, is the bandwidth of the Internet within the server, is the failure probability of the internet network within the server, is the packet loss probability of the Internet within the server, is the delay of the Internet in the server, is the bandwidth of the Internet between the servers, is the failure probability of the interconnection network between the servers, is the packet loss probability of the inter-server internetwork, is the delay of the Internet network between the servers.
9. The performance evaluation method according to claim 8, characterized in that: The calculating, based on the computing power design parameter, the storage design parameter, and the network design parameter, the execution time of the server cluster for executing at least one iteration of the simulation training task includes: When the server cluster supports consistent storage unit access, a second time calculation formula is used to calculate the execution time of the server cluster performing at least one iteration of the simulation training task based on the computing power design parameters, the storage design parameters and the network design parameters.
10. The performance evaluation method according to claim 9, characterized in that: The second time calculation formula is: 。 11. A computer program product, characterized in that Applied to a server cluster, the computer program product includes: A computing power parameter collection module, used to obtain the computing power design parameters of the server cluster; A storage parameter collection module, used to obtain storage design parameters of the server cluster; A network parameter collection module, used to obtain network design parameters of the server cluster; a performance evaluation module, configured to determine a simulation training task and calculate the computing performance of the server cluster in executing the simulation training task based on the computing power design parameters, the storage design parameters, and the network design parameters, to obtain a performance evaluation result; The determining of the simulation training task and calculating the computing performance of the server cluster in executing the simulation training task based on the computing power design parameters, the storage design parameters, and the network design parameters to obtain a performance evaluation result include: Determining the simulation training task and the total computational amount and total data amount of the simulation training task; Calculating, based on the computing power design parameters, the storage design parameters, and the network design parameters, the execution time of the server cluster performing at least one iteration of the simulation training task; Determine the performance evaluation result according to the execution time; The server cluster includes a plurality of servers, at least one of the servers includes at least one processing unit and at least one corresponding storage unit, and calculating, based on the computing power design parameter, the storage design parameter, and the network design parameter, the execution time of the server cluster for executing at least one iteration of the simulation training task, including: Calculating, based on the computing power design parameters, a first expected value for the server cluster to execute at least one iteration of the simulation training task, where the first expected value is an expected computing time of at least one of the processing units; Calculating, based on the computing power design parameter and the storage design parameter, a second expected value for the server cluster to execute at least one iteration of the simulation training task, where the second expected value is an expected reading time for at least one of the processing units to read the corresponding storage unit; Calculate, based on the computing power design parameter, the storage design parameter, and the network design parameter, a third expected value for the server cluster to execute at least one iteration of the simulation training task, where the third expected value is an expected synchronization time of at least one processing unit synchronizing with other storage units in the same server; Calculating, based on the computing power design parameter, the storage design parameter, and the network design parameter, a fourth expected value for the server cluster to perform at least one iteration of the simulation training task, the fourth expected value being an expected synchronization time of at least one of the processing units synchronizing with at least one of the storage units in the other servers; The execution time is calculated according to the first expected value, the second expected value, the third expected value, and the fourth expected value.
12. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the performance evaluation method according to any one of claims 1 to 10 when executing the computer program.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the performance evaluation method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Time consumption prediction simulation method, device, equipment, medium and system for heterogeneous computing power
CN117827619A
Method and device for predicting training time consumption in heterogeneous computing power based on CXL
CN119204361A