Performance evaluation method, computer program product, electronic device and storage medium
By obtaining the computing power, storage and network design parameters of the server cluster, and calculating the performance evaluation results of the simulated training task, the problem of time-consuming and overhead of server cluster performance evaluation is solved, and fast and low-cost performance evaluation and design is achieved.
Patent Information
- Application Number
- CN202510838342.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-20
AI Technical Summary
Cluster performance evaluation in an actual server cluster is time-consuming and overhead, resulting in high design and deployment costs.
By obtaining the computing power, storage and network design parameters of the server cluster, determining the simulation training task, and computing the computing performance of the server cluster performing the simulation training task, and obtaining the performance evaluation results.
It realizes lightweight and rapid evaluation of server cluster performance, significantly reducing time and overhead, reducing designer research and hardware deployment costs, and improving design efficiency.
Smart Images

Figure CN120371674A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of performance evaluation, and particularly to a performance evaluation method, a computer program product, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the development of Artificial Intelligence (AI) technology, the advantages of neural network models have gradually emerged, and the industry has also started to invest heavily in researching larger neural network models. Constrained by problems such as large computing and storage requirements for training, the training difficulty of neural network models has increased exponentially. To solve the training problem of neural network models, building a large-scale intelligent computing server cluster is taken as the main solution. However, as the scale of the server cluster expands, the time and cost of cluster performance evaluation also increase rapidly. How to reduce the time and cost of cluster performance evaluation is an urgent problem to be solved currently. Summary of the Invention
[0003] This application provides a performance evaluation method, a computer program product, an electronic device, and a computer-readable storage medium to at least solve the problem of large time and cost in cluster performance evaluation in an actual server cluster in related technologies.
[0004] This application provides a performance evaluation method applied to a server cluster. The performance evaluation method includes: Obtaining the computing power design parameters of the server cluster; Obtaining the storage design parameters of the server cluster; Obtaining the network design parameters of the server cluster; Determining a simulation training task, and calculating the computing performance of the server cluster for executing the simulation training task according to the computing power design parameters, storage design parameters, and network design parameters to obtain a performance evaluation result.
[0005] This application also provides a computer program product applied to a server cluster. The computer program product includes: A computing power parameter collection module for obtaining the computing power design parameters of the server cluster; A storage parameter collection module for obtaining the storage design parameters of the server cluster; A network parameter collection module for obtaining the network design parameters of the server cluster; A performance evaluation module for determining a simulation training task, and calculating the computing performance of the server cluster for executing the simulation training task according to the computing power design parameters, storage design parameters, and network design parameters to obtain a performance evaluation result.
[0006] The present application further provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above performance evaluation methods when executing the computer program.
[0007] The present application further provides a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of any of the above performance evaluation methods when executed by a processor.
[0008] Through the present application, by determining a simulation training task and calculating the computing performance of the server cluster for executing the simulation training task based on the computing power design parameters, storage design parameters, and network design parameters of the server cluster to obtain a performance evaluation result, it is possible to solve the technical problems of time-consuming and high overhead in evaluating the cluster performance in an actual server cluster, achieve the technical effects of realizing lightweight and fast evaluation of the performance of the server cluster, significantly reducing the time and overhead, and reducing the research cost and actual hardware deployment cost of designers, and improving the design efficiency of the server cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0010] Figure 1 It is a schematic flowchart of a performance evaluation method provided by an embodiment of the present application; Figure 2 It is a schematic diagram of the interconnection topology of a server cluster provided by an embodiment of the present application; Figure 3 It is a schematic flowchart of a performance evaluation method provided by an embodiment of the present application; Figure 4 It is a schematic flowchart of a performance evaluation method provided by an embodiment of the present application; Figure 5 It is a schematic flowchart of a performance evaluation method provided by an embodiment of the present application; Figure 6 It is a schematic flowchart of a performance evaluation method provided by an embodiment of the present application; Figure 7 It is a schematic flowchart of a performance evaluation method provided by an embodiment of the present application; Figure 8 It is a schematic diagram of the modules of a computer program product provided by an embodiment of the present application; Figure 9 It is a schematic diagram of the modules of an electronic device provided by an embodiment of the present application; Figure 10 Schematic diagram of a module of a computer-readable storage medium provided by an embodiment of the present application. Specific implementation manners
[0011] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0012] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0013] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0014] Please refer to Figure 1 and Figure 2 , an embodiment of the present application provides a performance evaluation method, and the method will be described in detail in combination with the execution process of the performance evaluation method.
[0015] The performance evaluation method is applied to the server cluster 100, and the performance evaluation method includes: 010: Obtain the computing power design parameters of the server cluster 100; 020: Obtain the storage design parameters of the server cluster 100; 030: Obtain the network design parameters of the server cluster 100; 040: Determine a simulation training task, and calculate the computing performance of the server cluster 100 for executing the simulation training task according to the computing power design parameters, storage design parameters, and network design parameters, and obtain a performance evaluation result.
[0016] In the performance evaluation method of the embodiments of this application, by determining a simulation training task and calculating the computing performance of the server cluster 100 for executing the simulation training task according to the computing power design parameters, storage design parameters, and network design parameters of the server cluster 100, a performance evaluation result is obtained. Therefore, the technical problem of time-consuming and high overhead in cluster performance evaluation in an actual server cluster 100 can be solved, achieving the technical effects of realizing lightweight and fast evaluation of the performance of the server cluster 100, significantly reducing time consumption and overhead, and reducing the research cost of designers and the actual hardware deployment cost, and improving the design efficiency of the server cluster 100.
[0017] Specifically, the server cluster 100 can be an intelligent computing server cluster 100. Designers can design the server cluster 100 according to actual application requirements. Before building the actual server cluster 100, obtain the computing power design parameters, storage design parameters, and network design parameters of the server cluster 100.
[0018] The design parameters of the computing power of the server cluster 100 for designers, that is, the computing power design parameters, and the computing power design parameters include parameters related to the calculation of the server cluster 100. The design parameters related to the storage of the server cluster 100 for designers, that is, the storage design parameters. The design parameters of the network of the server cluster 100 for designers, that is, the network design parameters, and the network design parameters include parameters related to the communication of the server cluster 100.
[0019] The computing power design parameters, storage design parameters, and network design parameters can all adopt the corresponding parameters in the product performance manual of the server 10, or the corresponding parameters obtained by performing performance tests on actual server 10 products, or the parameters directly input by designers.
[0020] According to the application requirements of the server cluster 100, for example, according to the training tasks that the server cluster 100 actually needs to execute, determine the corresponding simulation training task. After obtaining the computing power design parameters, storage design parameters, and network design parameters, the computing performance of the server cluster 100 for executing the simulation training task can be calculated according to the computing power design parameters, storage design parameters, and network design parameters, and a performance evaluation result is obtained.
[0021] It should be noted that the performance evaluation result is calculated through the parameters of the server cluster 100, and it is not necessary for the actual server cluster 100 to actually execute the simulation training task.
[0022] After obtaining the performance evaluation results, in the case where the performance evaluation results are unqualified, the designer can adjust the server cluster 100 according to the performance evaluation results to improve the performance evaluation results; in the case where the performance evaluation results are qualified, the actual deployment of the server cluster 100 can be carried out according to the current server cluster 100 design.
[0023] In this way, through the computing power design parameters, storage design parameters, and network design parameters of the server cluster 100, comprehensively considering the factors of computing, storage, and communication, the performance of the server cluster 100 can be evaluated lightly and quickly; compared with actual execution or simulation evaluation, the time consumption and overhead can be significantly reduced, and the research cost and actual hardware deployment cost of the designer can be reduced, and the design efficiency of the server cluster 100 can be improved.
[0024] The interconnection topology diagram of the server cluster 100 can be as Figure 2 shown. The server cluster 100 includes multiple servers 10 and multiple network switches. Each server 10 may include multiple processing units 11 and corresponding multiple storage units 12. The storage unit 12 includes a memory 121 (memory) and a cache. Each processing unit 11 is correspondingly connected to a memory 121, and the cache may be set within the processing unit 11.
[0025] The processing unit 11 may be a Graphics Processing Unit (GPU). The memory 121 may store the operation data in the processing unit 11. The cache is a buffer for data exchange and is a kind of temporary memory with a small capacity and high speed. Since the processing unit 11 needs to wait for a certain period to directly read data from the memory 121, the data to be used can be stored in the cache to reduce the waiting time of the processing unit 11.
[0026] The multiple processing units 11 within the same server 10 are interconnected by high bandwidth within the server 10, for example, by using the NVIDIA Virtual Link (NVLINK) method for interconnection. The processing units 11 between different servers 10 are interconnected through network switches.
[0027] The two processing units 11 between different servers 10 are interconnected through a network switch. Each two network switches connected to the processing unit 11 are also interconnected through a network switch. The network switch can adopt a non-blocking architecture, such as the clos architecture. The non-blocking architecture can ensure efficient interconnection and communication between multiple processing units 11.
[0028] For such as Figure 2The shown server cluster 100 faces various parameters to be considered during construction, such as the bandwidth of the network switch, the computing performance of the processing unit 11, the bandwidth of the memory 121, the performance of the cache, etc. According to the parameters to be considered for actual applications, the specific computing power design parameters, storage design parameters, and network design parameters to be obtained can be determined.
[0029] Please refer to Figure 2 , in some embodiments, the server cluster 100 includes multiple servers 10, and at least one server 10 includes at least one processing unit 11. The computing power design parameters include any one or more of the number of servers 10, the number of processing units 11 in the server 10, the computing power of the processing unit 11, the average computing power coefficient of the processing unit 11, and the variance of the actual computing power of the processing unit 11.
[0030] Specifically, the computing power design parameters may include: y: the number of servers 10; x: the number of processing units 11 in the server 10; C: the computing power of the processing unit 11; λ: the average computing power coefficient of the processing unit 11; : the variance of the actual computing power of the processing unit 11.
[0031] Among them, the computing power of the processing unit 11 refers to the computing power of a single processing unit 11. The average computing power coefficient of the processing unit 11 refers to the computing power that the processing unit 11 can actually exert, for example, it can be 0.5, 0.6, etc. The variance of the actual computing power of the processing unit 11 can be calculated based on the actual computing powers of multiple processing units 11.
[0032] Please refer to Figure 2 , in some embodiments, the server cluster 100 includes multiple servers 10, and at least one server 10 includes at least one storage unit 12. The storage design parameters include any one or more of the cache coherence access identification parameter, the hit rate of the storage unit 12, the access latency of the storage unit 12, the cache line size of the storage unit 12, and the bandwidth of the storage unit 12.
[0033] Specifically, the storage design parameters may include: U: the cache coherence access identification parameter; : the hit rate of the storage unit 12; : the access latency of the storage unit 12; : the cache line size of the storage unit 12; : Bandwidth of storage unit 12.
[0034] Among them, cache coherence means that the caches of multiple processing units can correctly and consistently reflect the latest state of the shared memory data. i = 1, 2, 3, 4 represents the number of levels in storage unit 12. The first level is L1 cache, the second level is L2 cache, the third level is L3 cache, and the fourth level is memory 121.
[0035] Specifically, it can represent the hit rate of the i-th level in storage unit 12. p1 represents the hit rate of L1 cache; p2 represents the hit rate of L2 cache; p3 represents the hit rate of L3 cache; p4 represents the hit rate of memory 121, and p4 = 1.
[0036] Specifically, it can represent the access latency of the i-th level in storage unit 12. represents the access latency of L1 cache; represents the access latency of L2 cache; represents the access latency of L3 cache; represents the access latency of memory 121.
[0037] Specifically, it can represent the bandwidth of the i-th level in storage unit 12. represents the bandwidth of L1 cache; represents the bandwidth of L2 cache; represents the bandwidth of L3 cache; represents the bandwidth of memory 121.
[0038] Please refer to Figure 2 , in some embodiments, the server cluster 100 includes multiple servers 10, and the network design parameters include any one or more of the bandwidth, latency, failure probability, packet loss probability of the internal interconnection network within the server 10, and the bandwidth, latency, failure probability, and packet loss probability of the interconnection network between the servers 10.
[0039] Specifically, the network design parameters may include: Bandwidth of the internal interconnection network within the server 10; Latency of the internal interconnection network within the server 10; Failure probability of the internal interconnection network within the server 10; Packet loss probability of the internal interconnection network within the server 10; The bandwidth of the interconnected network of 10 servers; The latency of the interconnected network of 10 servers; The failure probability of the interconnected network of 10 servers; The packet loss probability of the interconnected network of 10 servers.
[0040] Please refer to Figure 2 and Figure 3 , in some embodiments, determine a simulation training task, and calculate the computing performance of the server cluster 100 for executing the simulation training task according to the computing power design parameters, storage design parameters, and network design parameters, to obtain a performance evaluation result (i.e., 040), including: 041: Determine the simulation training task, as well as the total computing amount and total data amount of the simulation training task; 042: According to the computing power design parameters, storage design parameters, and network design parameters, calculate the execution time for the server cluster 100 to execute the simulation training task for at least one iteration; 043: Determine the performance evaluation result according to the execution time.
[0041] Specifically, determine the corresponding simulation training task according to the application requirements of the server cluster 100, and determine the total computing amount required for the simulation training task, as well as the total data amount to be processed by the simulation training task. For example, if the server cluster 100 is used for model training, the corresponding simulation training task can be determined, and the corresponding total computing amount and total data amount can be determined.
[0042] It can be understood that the model training process includes multiple iterations. In the embodiments of the present application, according to the obtained computing power design parameters, storage design parameters, and network design parameters, calculate the execution time for the server cluster 100 to execute the simulation training task for at least one iteration.
[0043] In one example, what is calculated is the execution time for the server cluster 100 to execute one iteration of the simulation training task. Of course, it is also possible to calculate the total execution time for the server cluster 100 to execute two, three, or more iterations of the simulation training task, which is not limited herein.
[0044] After that, the performance evaluation result of the server cluster 100 can be determined according to the execution time. For example, a time threshold can be set. When the execution time is greater than or equal to the time threshold, it is determined that the performance evaluation result of the server cluster 100 is unqualified, and the designer can adjust the server cluster 100 according to the performance evaluation result to improve the performance evaluation result; when the execution time is less than the time threshold, it is determined that the performance evaluation result of the server cluster 100 is qualified, and the designer can perform the actual deployment of the server cluster 100 according to the current design of the server cluster 100.
[0045] In this way, the performance evaluation of the server cluster 100 can be achieved through lightweight calculations.
[0046] Please refer to Figure 2 and Figure 4 , in some embodiments, determining the simulation training task and the total amount of calculations and total amount of data of the simulation training task (i.e., 041) includes: 0411: Constructing a simulation training task based on a neural network model; 0412: Determining the total amount of calculations and total amount of data of the simulation training task by analyzing the neural network model.
[0047] Specifically, if the application requirement of the server cluster 100 is to perform the training of a neural network model, a simulation training task can be constructed based on the neural network model. In some embodiments, the neural network model is a language model, and the distributed AI computing task needs to be processed by the server cluster 100. Therefore, a simulation training task can be constructed based on any language model, and the simulation training task can be a distributed AI computing task.
[0048] For the selected neural network model, the total amount of calculations and total amount of data can be directly determined using the data in the corresponding literature or model configuration file of the neural network model; the total amount of calculations and total amount of data can also be determined by analyzing the neural network model through a performance testing tool.
[0049] In this way, the performance evaluation result is closer to the evaluation result obtained by the server cluster 100 actually performing the training task, improving the reliability of the performance evaluation result.
[0050] During the process of the server cluster 100 performing at least one iteration of the simulation training task, the execution time includes the calculation time-consuming of the processing unit 11, the time-consuming of the processing unit 11 reading data from the corresponding storage unit 12, the time-consuming of the processing unit 11 synchronizing data from other storage units 12 in the same server 10, and the time-consuming of the processing unit 11 synchronizing data from the storage unit 12 of the processing unit 11 in other servers 10. The calculation process of the execution time is described in detail below.
[0051] Please refer toFigure 2 and Figure 5 ,In some embodiments, the server cluster 100 includes multiple servers 10, and at least one server 10 includes at least one processing unit 11 and a corresponding at least one storage unit 12. According to the computing power design parameters, storage design parameters, and network design parameters, the server cluster 100 calculates the execution time (i.e., 042) for performing at least one iteration of the simulation training task, including: 0421: According to the computing power design parameters, calculate the first expected value for the server cluster 100 to perform at least one iteration of the simulation training task. The first expected value is the computing expected time of at least one processing unit 11; 0422: According to the computing power design parameters and storage design parameters, calculate the second expected value for the server cluster 100 to perform at least one iteration of the simulation training task. The second expected value is the reading expected time for at least one processing unit 11 to read the corresponding storage unit 12; 0423: According to the computing power design parameters, storage design parameters, and network design parameters, calculate the third expected value for the server cluster 100 to perform at least one iteration of the simulation training task. The third expected value is the synchronization expected time for at least one processing unit 11 to synchronize other storage units 12 in the same server 10; 0424: According to the computing power design parameters, storage design parameters, and network design parameters, calculate the fourth expected value for the server cluster 100 to perform at least one iteration of the simulation training task. The fourth expected value is the synchronization expected time for at least one processing unit 11 to synchronize at least one storage unit 12 in other servers 10; 0425: Calculate the execution time according to the first expected value, second expected value, third expected value, and fourth expected value.
[0052] Specifically, taking one iteration of the simulation training task as an example, according to the computing power design parameters, the first expected value for the server cluster 100 to perform one iteration of the simulation training task can be calculated. The first expected value is the computing expected time of at least one processing unit 11. The first expected value can represent the computing time for one iteration of the simulation training task.
[0053] It should be noted that when the server 10 includes multiple processing units 11, the multiple processing units 11 can perform computing tasks in parallel. Therefore, the computing time of the multiple processing units 11 depends on the time-consuming of the slowest processing unit 11, and the computing expected time refers to the expected time-consuming of the slowest processing unit 11.
[0054] In some embodiments, according to the computing power design parameters, calculating the first expected value for the server cluster 100 to perform at least one iteration of the simulation training task, where the first expected value is the computing expected time of at least one processing unit 11, includes: According to the computing power design parameters, the Gumbel extreme value distribution is used to calculate the first expected value of the server cluster 100 to execute the simulation training task for at least one iteration. The first expected value is the computing expected time of at least one processing unit 11.
[0055] Since the computing time consumption of the processing unit 11 is affected by performance fluctuations such as load imbalance, heat dissipation frequency reduction, hardware heterogeneity, and communication competition, the computing time consumption of multiple processing units 11 is random. Therefore, the Gumbel extreme value distribution is used to model the computing time consumption of at least one processing unit 11 to calculate the first expected value. In this way, considering the randomness of the working performance of the processing unit 11, the first expected value in the case of performance fluctuations can be calculated, so that the calculated execution time is closer to the actual execution time, and thus the performance evaluation result is more accurate.
[0056] According to the computing power design parameters and the storage design parameters, the second expected value of the server cluster 100 to execute one iteration of the simulation training task can be calculated. The second expected value is the reading expected time for at least one processing unit 11 to read the corresponding storage unit 12. Each processing unit 11 corresponds to a storage unit 12, and the storage unit 12 corresponding to the processing unit 11 is called the local storage unit. During the process of executing one iteration of the simulation training task, each processing unit 11 can read the local storage unit once or multiple times.
[0057] According to the computing power design parameters, the storage design parameters, and the network design parameters, the third expected value of the server cluster 100 to execute one iteration of the simulation training task is calculated. The third expected value is the synchronization expected time for at least one processing unit 11 to synchronize other storage units 12 in the same server 10.
[0058] The storage units 12 corresponding to other processing units 11 in the same server 10 can all be called the first remote storage units. The third expected value is the synchronization expected time for at least one processing unit 11 to synchronize the first remote storage units.
[0059] According to the computing power design parameters, the storage design parameters, and the network design parameters, the fourth expected value of the server cluster 100 to execute one iteration of the simulation training task is calculated and analyzed. The fourth expected value is the synchronization expected time for at least one processing unit 11 to synchronize at least one storage unit 12 in other servers 10.
[0060] The storage units 12 corresponding to other processing units 11 in other servers 10 can all be called the second remote storage units. The fourth expected value is the synchronization expected time for at least one processing unit 11 to synchronize the second remote storage units.
[0061] After that, the execution time can be calculated based on the first expected value, the second expected value, the third expected value, and the fourth expected value. In one example, the execution time can be obtained by adding the first expected value, the second expected value, the third expected value, and the fourth expected value.
[0062] In this way, considering all aspects of the time consumption when the server cluster 100 executes the training task, the reliability of the calculated execution time is relatively high.
[0063] Please refer to Figure 2 and Figure 6 , in some embodiments, according to the computing power design parameters, the storage design parameters, and the network design parameters, calculate the execution time (i.e., 042) for the server cluster 100 to execute at least one iteration of the simulation training task, including: 0426: Determine whether the server cluster 100 supports cache coherence access according to the storage design parameters; 0427: In the case where the server cluster 100 does not support cache coherence access, according to the computing power design parameters, the storage design parameters, and the network design parameters, use the first time calculation formula to calculate the execution time for the server cluster 100 to execute at least one iteration of the simulation training task.
[0064] Specifically, it can be determined whether the server cluster 100 supports cache coherence access according to the cache coherence access identification parameter in the storage design parameters.
[0065] In the case where the server cluster 100 does not support cache coherence access, when the processing unit 11 reads the data in the first remote storage unit and the second remote storage unit, it is necessary to transfer the data in the first remote storage unit and the second remote storage unit to the local storage unit and then read it from the local storage unit.
[0066] In the case where the server cluster 100 supports the access of the coherent storage unit 12, the processing unit 11 can directly access and read the data in the first remote storage unit and the second remote storage unit.
[0067] In the case where the server cluster 100 does not support cache coherence access, according to the computing power design parameters, the storage design parameters, and the network design parameters, the first time calculation formula can be used to calculate the execution time for the server cluster 100 to execute at least one iteration of the simulation training task.
[0068] In some embodiments, the first time calculation formula is: ; where is the execution time, is the total amount of computation, is the number of processing units 11 in the server 10, is the number of servers 10, is the average computing power coefficient of the processing unit 11, is the computing power of the processing unit 11, is the variance of the actual computing power of the processing unit 11, is the first distribution coefficient, is the second distribution coefficient, is the total data volume, is the cache line size of the storage unit 12, is the probability that occurs when the i-th level of the storage unit 12 is hit, is the bandwidth of the i-th level of the storage unit 12, is the access latency of the i-th level of the storage unit 12, is the bandwidth of the internal interconnection network within the server 10, is the failure probability of the internal interconnection network within the server 10, is the packet loss probability of the internal interconnection network within the server 10, is the latency of the internal interconnection network within the server 10, is the bandwidth of the interconnection network between servers 10, is the failure probability of the interconnection network between servers 10, is the packet loss probability of the interconnection network between servers 10, is the latency of the interconnection network between servers 10.
[0069] Specifically, represents the first expected value. , y, C, λ, and are computing power design parameters. The first expected value is calculated according to the Gumbel extreme value distribution as . can represent the expected computing time consumption of each processing unit 11, , and are as follows: ; ; where, represents the distribution function of the standard normal distribution, and e is the natural constant.
[0070] represents the second expected value.
[0071] , , are storage design parameters. represents the transmission time for the processing unit 11 to read data from the local storage unit, Indicates the number of communication times for the processing unit 11 to read data. Indicates the additional read latency caused by misses each time data is read from the storage unit 12.
[0072] , indicates the probability that occurs when the L1 cache hits; , indicates the probability that occurs when the L2 cache hits; , indicates the probability that occurs when the L3 cache hits; Indicates the probability that occurs when the memory 121 hits.
[0073] Indicates the third expected value.
[0074] 、 、 、 、 、 、 and are network design parameters. Indicates the actual transmission time for transferring the data of the first remote storage unit to the local storage unit in the presence of link failures and packet loss retransmission. Indicates the transmission time in the ideal transfer case, Indicates the probability of no-fault and no-packet-loss calculated using the Gumbel extreme value distribution.
[0075] In this way, considering the situation of link failures and packet loss retransmission during the transmission process, as well as the randomness of the situation, the actual transmission time is calculated through the Gumbel extreme value distribution, making the calculated execution time closer to the actual execution time, and thus making the performance evaluation result more accurate.
[0076] Indicates the time for the processing unit 11 to read data from the local storage unit. The calculation method is similar to the second expected value and will not be elaborated here.
[0077] Indicates the fourth expected value.
[0078] Indicates the actual transmission time for transferring the data of the second remote storage unit to the local storage unit in the presence of link failures and packet loss retransmission. Indicates the transmission time in the ideal transfer case, Indicates the probability of no-fault and no-packet-loss calculated using the Gumbel extreme value distribution.
[0079] In this way, considering the link failures and packet loss retransmission during the transmission process, as well as the randomness of such situations, the actual transmission time is calculated through the Gumbel extreme value distribution, making the calculated execution time closer to the actual execution time, and thus making the performance evaluation results more accurate.
[0080] It represents the time taken for the processing unit 11 to read data from the local storage unit. The calculation method is similar to that of the second expected value and will not be elaborated here.
[0081] The execution time can be obtained by adding the first expected value, the second expected value, the third expected value, and the fourth expected value.
[0082] Please refer to Figure 2 and Figure 7 , in some embodiments, according to the computing power design parameters, storage design parameters, and network design parameters, the execution time (i.e., 042) for the server cluster 100 to execute at least one iteration of the simulation training task is calculated, including: 0428: When the server cluster 100 supports cache coherent access, according to the computing power design parameters, storage design parameters, and network design parameters, the second time calculation formula is used to calculate the execution time for the server cluster 100 to execute at least one iteration of the simulation training task.
[0083] In some embodiments, the second time calculation formula is: .
[0084] Specifically, in the embodiments of the present application, the calculation processes of the first expected value and the second expected value are the same as those in the first time calculation formula and will not be elaborated here.
[0085] It represents the third expected value. When cache coherent access is supported, the processing unit 11 can directly access and read the data in the first remote storage unit. Compared with the first time calculation formula, the actual transmission time for moving the data in the first remote storage unit to the local storage unit can be omitted.
[0086] , , , , , , and are network design parameters. It represents the actual transmission time for the processing unit 11 to read data from the first remote storage unit in the presence of link failures and packet loss retransmission. It represents the transmission time in the ideal reading case. Indicates the probability of no fault and no packet loss calculated using the Gumbel extreme value distribution.
[0087] In this way, the cases of link failure and packet loss retransmission during the reading process caused by unstable remote network and the randomness of the occurrence of the cases are considered. By calculating the actual reading time-consuming through the Gumbel extreme value distribution, the calculated execution time is closer to the actual execution time, and thus the performance evaluation result is more accurate.
[0088] Indicates the time-consuming for the processing unit 11 to read data from the local storage unit. The calculation method is similar to that of the second expected value and will not be elaborated here.
[0089] Indicates the fourth expected value.
[0090] Indicates the actual transmission time-consuming for the processing unit 11 to read data from the second remote storage unit in the case of link failure and packet loss retransmission. Indicates the transmission time-consuming in the ideal reading case. Indicates the probability of no fault and no packet loss calculated using the Gumbel extreme value distribution.
[0091] In this way, the cases of link failure and packet loss retransmission during the reading process caused by unstable remote network and the randomness of the occurrence of the cases are considered. By calculating the actual reading time-consuming through the Gumbel extreme value distribution, the calculated execution time is closer to the actual execution time, and thus the performance evaluation result is more accurate.
[0092] Indicates the time-consuming for the processing unit 11 to read data from the local storage unit. The calculation method is similar to that of the second expected value and will not be elaborated here.
[0093] The first expected value, the second expected value, the third expected value and the fourth expected value are added together to obtain the execution time.
[0094] In the embodiments of the present application, according to whether the server cluster 100 supports cache coherence access, different time calculation formulas are used to calculate the execution time. In this way, the time-consuming differences in the two cases where the server cluster 100 does not support and supports cache coherence access are considered, so that the calculated execution time is closer to the actual execution time, and thus the performance evaluation result is more accurate.
[0095] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0096] Please refer to Figure 8 Figure 8 , an embodiment of the present application further provides a computer program product 200, which is applied to a server cluster 100. The computer program product 200 includes a computing power parameter collection module 210, a storage parameter collection module 220, a network parameter collection module 230, and a performance evaluation module 240. The computing power parameter collection module 210 is used to obtain the computing power design parameters of the server cluster 100. The storage parameter collection module 220 is used to obtain the storage design parameters of the server cluster 100. The network parameter collection module 230 is used to obtain the network design parameters of the server cluster 100. The performance evaluation module 240 is used to determine a simulation training task, and calculate the computing performance of the server cluster 100 for executing the simulation training task based on the computing power design parameters, storage design parameters, and network design parameters, so as to obtain a performance evaluation result.
[0097] In the computer program product of the embodiment of the present application, the computing power parameter collection module 210, the storage parameter collection module 220, and the network parameter collection module 230 are respectively used to obtain the computing power design parameters, storage design parameters, and network design parameters. The performance evaluation module 240 determines a simulation training task, and calculates the computing performance of the server cluster 100 for executing the simulation training task based on the computing power design parameters, storage design parameters, and network design parameters of the server cluster 100, so as to obtain a performance evaluation result. Therefore, the technical problem of time-consuming and high overhead in cluster performance evaluation in the actual server cluster 100 can be solved, and the technical effect of realizing lightweight and fast evaluation of the performance of the server cluster 100, significantly reducing the time consumption and overhead, and reducing the research cost and actual hardware deployment cost of designers, and improving the design efficiency of the server cluster 100 can be achieved.
[0098] In some embodiments, the performance evaluation module 240 is specifically used to determine a simulation training task, the total computing amount and total data amount of the simulation training task; calculate the execution time of the server cluster 100 for executing the simulation training task at least once based on the computing power design parameters, storage design parameters, and network design parameters; and determine the performance evaluation result according to the execution time.
[0099] In some embodiments, the performance evaluation module 240 is specifically used to construct a simulation training task based on a neural network model; determine the total computing amount and total data amount of the simulation training task by analyzing the neural network model.
[0100] In some embodiments, the server cluster 100 includes multiple servers 10, and at least one server 10 includes at least one processing unit 11 and a corresponding at least one storage unit 12. The performance evaluation module 240 is specifically configured to calculate a first expected value of the server cluster 100 for performing at least one iteration of the simulation training task according to the computing power design parameters, where the first expected value is the expected computing time of at least one processing unit 11; calculate a second expected value of the server cluster 100 for performing at least one iteration of the simulation training task according to the computing power design parameters and the storage design parameters, where the second expected value is the expected reading time for at least one processing unit 11 to read the corresponding storage unit 12; calculate a third expected value of the server cluster 100 for performing at least one iteration of the simulation training task according to the computing power design parameters, the storage design parameters, and the network design parameters, where the third expected value is the expected synchronization time for at least one processing unit 11 to synchronize other storage units 12 in the same server 10; calculate and analyze a fourth expected value of the server cluster 100 for performing at least one iteration of the simulation training task according to the computing power design parameters, the storage design parameters, and the network design parameters, where the fourth expected value is the expected synchronization time for at least one processing unit 11 to synchronize at least one storage unit 12 in other servers 10; and calculate the execution time according to the first expected value, the second expected value, the third expected value, and the fourth expected value.
[0101] In some embodiments, the performance evaluation module 240 is specifically configured to determine whether the server cluster 100 supports cache coherence access according to the storage design parameters; and in the case where the server cluster 100 does not support cache coherence access, calculate the execution time of the server cluster 100 for performing at least one iteration of the simulation training task according to the computing power design parameters, the storage design parameters, and the network design parameters by using a first time calculation formula.
[0102] In some embodiments, the performance evaluation module 240 is specifically configured to, in the case where the server cluster 100 supports cache coherence access, calculate the execution time of the server cluster 100 for performing at least one iteration of the simulation training task according to the computing power design parameters, the storage design parameters, and the network design parameters by using a second time calculation formula.
[0103] For the description of the features in the corresponding embodiments of the computer program product 200, reference may be made to the relevant description in the corresponding embodiments of the performance evaluation method, which will not be elaborated here one by one.
[0104] Please refer to Figure 9 , an embodiment of the present application further provides an electronic device 300, including a memory 310 and a processor 320. A computer program is stored in the memory 310, and the processor 320 is configured to run the computer program to execute the steps in any one of the above-mentioned performance evaluation method embodiments.
[0105] Please refer toFigure 10 Embodiments of the present application also provide a computer-readable storage medium 400, in which a computer program 410 is stored. The computer program 410 is configured to execute the steps in any of the above-described embodiments of the performance evaluation method when running.
[0106] In an exemplary embodiment, the above computer-readable storage medium 400 may include, but is not limited to: various media that can store computer programs such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs.
[0107] Embodiments of the present application also provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the performance evaluation method.
[0108] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the performance evaluation method.
[0109] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0110] The above has introduced in detail a performance evaluation method, a computer program product 200, an electronic device 300, and a computer-readable storage medium 400 provided by the present application. Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A performance evaluation method, characterized in that, Applied to a server cluster, the performance evaluation method includes: Obtain the computing power design parameters of the server cluster; Obtain the storage design parameters of the server cluster; Obtain the network design parameters of the server cluster; Determine a simulation training task, and calculate the computing performance of the server cluster for executing the simulation training task according to the computing power design parameters, the storage design parameters, and the network design parameters, to obtain a performance evaluation result.
2. The performance evaluation method according to claim 1, wherein The server cluster includes multiple servers, and at least one of the servers includes at least one processing unit. The computing power design parameters include any one or more of the number of servers, the number of processing units in the server, the computing power of the processing unit, the average computing power coefficient of the processing unit, and the variance of the actual computing power of the processing unit.
3. The performance evaluation method according to claim 1, wherein The server cluster includes multiple servers, and at least one of the servers includes at least one storage unit. The storage design parameters include any one or more of the cache coherence access identification parameter, the hit rate of the storage unit, the access latency of the storage unit, the cache line size of the storage unit, and the bandwidth of the storage unit.
4. The performance evaluation method according to claim 1, wherein The server cluster includes multiple servers. The network design parameters include any one or more of the bandwidth, latency, failure probability, packet loss probability of the interconnection network within the server, and the bandwidth, latency, failure probability, and packet loss probability of the interconnection network between the servers.
5. The performance evaluation method according to claim 1, characterized in that The determining the simulation training task, and calculating the computing performance of the server cluster for executing the simulation training task according to the computing power design parameters, the storage design parameters, and the network design parameters, to obtain a performance evaluation result, includes: Determine the simulation training task and the total computing amount and total data amount of the simulation training task; Calculate the execution time of the server cluster for executing the simulation training task for at least one iteration according to the computing power design parameters, the storage design parameters, and the network design parameters; Determine the performance evaluation result according to the execution time.
6. The performance evaluation method according to claim 5, wherein The determining the simulation training task and the total computing amount and total data amount of the simulation training task includes: Construct the simulation training task based on a neural network model; Determine the total computing amount and total data amount of the simulation training task by analyzing the neural network model.
7. The performance evaluation method according to claim 6, wherein The neural network model is a language model, and the simulation training task is a distributed artificial intelligence computing task.
8. The performance evaluation method according to claim 5, wherein The server cluster includes multiple servers, and at least one of the servers includes at least one processing unit and a corresponding at least one storage unit. The calculating the execution time of the server cluster for executing the simulation training task for at least one iteration according to the computing power design parameters, the storage design parameters, and the network design parameters includes: Calculate a first expected value of the server cluster for executing the simulation training task for at least one iteration according to the computing power design parameters, where the first expected value is the expected computing time of at least one of the processing units; According to the computing power design parameters and the storage design parameters, calculate a second expected value for the server cluster to execute at least one iteration of the simulation training task, where the second expected value is the expected read time for at least one of the processing units to read the corresponding storage unit; According to the computing power design parameters, the storage design parameters, and the network design parameters, calculate a third expected value for the server cluster to execute at least one iteration of the simulation training task, where the third expected value is the expected synchronization time for at least one of the processing units to synchronize other storage units in the same server; According to the computing power design parameters, the storage design parameters, and the network design parameters, calculate a fourth expected value for the server cluster to execute at least one iteration of the simulation training task, where the fourth expected value is the expected synchronization time for at least one of the processing units to synchronize at least one storage unit in other servers; Calculate the execution time according to the first expected value, the second expected value, the third expected value, and the fourth expected value.
9. The performance evaluation method according to claim 8, wherein The method for calculating the execution time for the server cluster to execute at least one iteration of the simulation training task according to the computing power design parameters, the storage design parameters, and the network design parameters includes: Determine whether the server cluster supports consistent storage unit access according to the storage design parameters; In the case where the server cluster does not support consistent storage unit access, calculate the execution time for the server cluster to execute at least one iteration of the simulation training task according to the computing power design parameters, the storage design parameters, and the network design parameters using a first time calculation formula.
10. The performance evaluation method according to claim 9, characterized in that, The first time calculation formula is: ; Among them, is the execution time, is the total amount of calculations, is the number of processing units in the server, is the number of servers, is the average computing power coefficient of the processing unit, is the computing power of the processing unit, is the variance of the actual computing power of the processing unit, is the first distribution coefficient, is the second distribution coefficient, is the total amount of data, is the cache line size of the storage unit, is the probability of the i-th level hit of the storage unit, is the bandwidth of the i-th level of the storage unit, is the access latency of the i-th level of the storage unit, is the bandwidth of the interconnection network within the server, is the failure probability of the interconnection network within the server, is the packet loss probability of the interconnection network within the server, is the latency of the interconnection network within the server, is the bandwidth of the interconnection network between the servers, is the failure probability of the interconnection network between the servers, is the packet loss probability of the interconnection network between the servers, is the latency of the interconnection network between the servers.
11. The performance evaluation method according to claim 10, characterized in that, The method for calculating the execution time for the server cluster to execute at least one iteration of the simulation training task according to the computing power design parameters, the storage design parameters, and the network design parameters includes: In the case where the server cluster supports consistent storage unit access, calculate the execution time for the server cluster to execute at least one iteration of the simulation training task according to the computing power design parameters, the storage design parameters, and the network design parameters using a second time calculation formula.
12. The performance evaluation method according to claim 11, wherein The second time calculation formula is: 。 13. A computer program product, characterized in that, Applied to a server cluster, the computer program product includes: A computing power parameter collection module for obtaining the computing power design parameters of the server cluster; A storage parameter collection module for obtaining the storage design parameters of the server cluster; A network parameter collection module for obtaining the network design parameters of the server cluster; A performance evaluation module for determining a simulation training task, and calculating the computing performance of the server cluster for executing the simulation training task according to the computing power design parameters, the storage design parameters, and the network design parameters, to obtain a performance evaluation result.
14. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for implementing the steps of the performance evaluation method according to any one of claims 1 to 12 when executing the computer program.
15. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein when the computer program is executed by a processor, the steps of the performance evaluation method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Time consumption prediction simulation method, device, equipment, medium and system for heterogeneous computing power
CN117827619A
Heterogeneous computing platform and task simulation and time consumption prediction method, device and equipment thereof
CN117971630A
Method and device for predicting training time consumption in heterogeneous computing power based on CXL
CN119204361A