Performance testing method, device, program product and medium for accelerator system

By conducting benchmark communication, collective communication and computing tests on the accelerator system, combining model training and inference, and adopting scoring rules and weight coefficients, the comprehensiveness and accuracy problems of accelerator performance testing in existing technologies are solved, and accurate performance evaluation in different scenarios is achieved.

CN120547099BActive Publication Date: 2025-09-23INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511029783.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-09-23
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

Existing technologies lack comprehensiveness and accuracy when conducting performance testing of artificial intelligence accelerators, and cannot effectively reflect their performance in different scenarios, especially in terms of communication performance and scenario adaptability.

Method used

The performance testing method of the accelerator system is adopted, including benchmark communication test, collective communication test and computing test, combined with model training and reasoning, through scoring rules and weight coefficients to comprehensively evaluate the performance of the accelerator system in different application scenarios.

Benefits of technology

It achieves a more comprehensive and accurate evaluation of the accelerator system performance, can reflect its communication performance, computing performance and energy efficiency in different scenarios, and provide more accurate performance test results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120547099B_ABST
    Figure CN120547099B_ABST
Patent Text Reader

Abstract

The present application discloses a performance testing method, device, program product, and medium for an accelerator system, which are applied to the field of accelerator technology. The method comprises obtaining basic communication test data, aggregate communication test data, calculation test data, application performance test data, and energy efficiency test data of the accelerator system; determining a score value corresponding to each test data based on a preset scoring rule; obtaining, for each preset application scenario, a weight coefficient configured for each score value in the application scenario, and obtaining a performance test result of the accelerator system in the application scenario based on the weight coefficient and each score value. By applying the solution of the present application, the performance test of the accelerator system can be implemented more comprehensively and accurately, and the performance requirements for the accelerator system in different application scenarios are taken into account.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of accelerator technology, and in particular to a performance testing method, device, program product, and medium for an accelerator system. Background Art

[0002] Artificial intelligence (AI) is experiencing explosive growth. The widespread adoption of GPUs (Graphics Processing Units) has made parallel computing faster, cheaper, and more efficient. Furthermore, the expansion of storage capacity has also led to a massive explosion in image data, text data, transaction data, and mapping data. This has led to a sharp increase in demand for computing performance and has fueled the emergence of a large number of AI accelerators, including GPGPUs (General-Purpose Graphics Processing Units), ASICs (Application-Specific Integrated Circuits), and FPGAs (Field-Programmable Gate Arrays).

[0003] AI accelerators are core components of AI computing and intelligent computing centers, directly impacting computing performance. Current performance testing of AI accelerators typically focuses on a single metric, such as floating-point operations per second (FLOPS) or integer operations per second (IOPS), using standard benchmarks like MLPerf. A few solutions incorporate supplementary metrics, but these remain incomplete and fail to effectively measure capabilities like communication performance and scenario adaptability, leaving room for improvement in the accuracy of the resulting evaluation results. Furthermore, the performance requirements for AI accelerators vary across different use cases, and current evaluation methods are not deeply tied to these scenarios, effectively failing to accurately reflect the performance of AI accelerators in diverse scenarios.

[0004] In summary, how to more comprehensively and accurately implement the performance test of the accelerator system and reflect its performance in different scenarios is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention

[0005] This application provides a performance testing method, device, program product, and medium for an accelerator system to more comprehensively and accurately implement the performance testing of the accelerator system and reflect its performance in different scenarios.

[0006] This application provides a performance testing method for an accelerator system, comprising:

[0007] Performing a baseline communication test, a collective communication test, and a computation test on the accelerator system, respectively, to obtain basic communication test data reflecting the communication performance between different devices in the accelerator system, collective communication test data reflecting the distributed communication operation execution performance of the accelerator system, and computation test data reflecting the computation performance of the accelerator system;

[0008] Performing model training and reasoning on the accelerator system to obtain application performance test data reflecting the model training and reasoning capabilities of the accelerator system, and energy efficiency test data reflecting the energy efficiency of the accelerator system during model training and reasoning;

[0009] Determine, based on preset scoring rules, scoring values ​​corresponding to the basic communication test data, the collective communication test data, the computing test data, the application performance test data, and the energy efficiency test data;

[0010] For each preset application scenario, a weight coefficient configured for each score value in the application scenario is obtained, and based on the weight coefficient and each score value, a performance test result of the accelerator system in the application scenario is obtained.

[0011] The present application provides an electronic device, comprising:

[0012] Memory for storing computer programs;

[0013] A processor is configured to implement the steps of the above-mentioned performance testing method for the accelerator system when executing the computer program.

[0014] The present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the performance testing method of the accelerator system as described above are implemented.

[0015] The present application provides a computer program product, including a computer program, which implements the steps of the performance testing method of the accelerator system as described above when the computer program is executed by a processor.

[0016] By applying the technical solution provided by the embodiment of the present invention, in order to ensure the accuracy of the performance test results obtained, the accelerator system can be tested in multiple aspects respectively, so as to obtain more comprehensive evaluation indicators, which is conducive to more accurate performance testing of the accelerator system. Specifically, the present application solution will perform benchmark communication tests, collective communication tests and computing tests on the accelerator system respectively. Through the benchmark communication test, basic communication test data reflecting the communication performance between different devices in the accelerator system can be obtained. Through the collective communication test, collective communication test data reflecting the distributed communication operation execution performance of the accelerator system can be obtained. And through the computing test, computing test data reflecting the computing performance of the accelerator system can be obtained. It can be seen that in addition to the computing performance of the accelerator system, the present application also takes into account the communication performance between different devices and the distributed communication operation execution performance of the accelerator system, which is conducive to a more comprehensive evaluation of the accelerator system.

[0017] In addition, the present application scheme takes into account that the accelerator system is usually used for model training or reasoning. Therefore, by performing model training and reasoning on the accelerator system, application performance test data reflecting the model training and reasoning capabilities of the accelerator system and energy efficiency test data reflecting the energy efficiency of the accelerator system during model training and reasoning are obtained. This can effectively reflect the performance of the accelerator system during model training and reasoning, thereby ensuring that the present application scheme can more comprehensively and accurately implement the performance test of the accelerator system.

[0018] In addition, the present application also takes into account that the performance requirements for the accelerator system are different in different application scenarios. Therefore, based on the preset scoring rules, after determining the corresponding scoring values ​​of the basic communication test data, the collective communication test data, the calculation test data, the application performance test data, and the energy efficiency test data, in the present application scheme, each preset application scenario will be evaluated separately, that is, for each application scenario, the weight coefficient configured for each scoring value in the application scenario will be obtained, and then based on the weight coefficient and each scoring value, the performance test result of the accelerator system in the application scenario will be obtained. In other words, the present application scheme will determine the performance test results of the accelerator system in different scenarios.

[0019] In summary, the present application solution realizes the performance test of the accelerator system more comprehensively and accurately, and takes into account the differences in performance requirements for the accelerator system in different application scenarios, and thus determines the corresponding performance test results of the accelerator system in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0021] Figure 1 A flowchart of a performance testing method for an accelerator system provided in a specific embodiment of the present invention;

[0022] Figure 2 A schematic diagram of the topological structure of an accelerator system provided in a specific embodiment of the present invention;

[0023] Figure 3 This is a schematic structural diagram of an electronic device according to the present invention;

[0024] Figure 4 This is a schematic structural diagram of a computer-readable storage medium of the present invention. DETAILED DESCRIPTION

[0025] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0026] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0027] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods. Figure 1 , Figure 1 This is a flowchart of an implementation method of an accelerator system performance test method provided in a specific embodiment of the present invention. The accelerator system performance test method may include the following steps:

[0028] Step S101: Perform a baseline communication test, a collective communication test, and a computing test on the accelerator system to obtain basic communication test data reflecting the communication performance between different devices in the accelerator system, collective communication test data reflecting the distributed communication operation execution performance of the accelerator system, and computing test data reflecting the computing performance of the accelerator system.

[0029] The present application scheme takes into account that, for an accelerator system, in addition to computing performance, the communication performance between different devices in the accelerator system and the distributed communication operation execution performance of the accelerator system will also affect the performance of the accelerator system ultimately presented to the user. For example, although some accelerator systems have a high FLOPS (Floating Point Operations Per Second), that is, their computing performance is very strong, due to reasons such as unreasonable topology design, the communication performance between different devices in the accelerator system is low, or the distributed communication operation execution performance is low, which will also lead to the final performance of the accelerator system being unsatisfactory. Therefore, the present application scheme is not based solely on a single indicator such as computing performance, but will comprehensively consider multiple aspects of performance to achieve performance testing of the accelerator system.

[0030] The accelerator system described in this application may include multiple devices, such as one or more CPUs and a GPU array controlled by these CPUs. Different accelerator systems may have different topological structures, communication rules, and other software and hardware settings. Figure 2 In the example, multiple accelerators are deployed in each AI server. These AI servers form an AI server cluster, and communication connections are achieved through a switch network. They can be controlled by the host, which can usually be implemented by a master server.

[0031] In addition to using GPUs (graphics processing units) as accelerators in accelerator systems, ASICs (application-specific integrated circuits), FPGAs (field-programmable gate arrays), TPUs (tensor processing units), NPUs (neural network processing units), and VPUs (visual processing units) can also be used as accelerators. In other words, different accelerator systems can have a variety of device types that implement the acceleration function, and these can be set according to actual needs. In some cases, accelerator systems may also include multiple types of accelerators simultaneously, without affecting the implementation of the present invention.

[0032] This application requires a baseline communication test of the accelerator system to obtain basic communication test data that reflects the communication performance between different devices in the accelerator system. Regarding the baseline communication test, the specific test indicators can be set according to actual needs, but it is understood that they should be able to effectively reflect the communication performance between different devices in the accelerator system. For example, the communication bandwidth and latency between different devices can usually be selected to reflect the basic communication performance of the accelerator system.

[0033] In a specific embodiment of the present invention, the benchmark communication test performed on the accelerator system in step S101 to obtain basic communication test data reflecting the communication performance between different devices in the accelerator system may specifically include:

[0034] Performing a benchmark communication test on the accelerator system to obtain a data transmission rate and an average delay when sending data from a host in the accelerator system to an accelerator in the accelerator system, as a first indicator and a second indicator, respectively; and obtaining a data transmission rate and an average delay when sending data from an accelerator in the accelerator system to a host in the accelerator system, as a third indicator and a fourth indicator, respectively;

[0035] Conducting a benchmark communication test on the accelerator system to obtain the data transmission rate and average latency between accelerators in the accelerator system as the fifth and sixth indicators, respectively.

[0036] The first indicator, the second indicator, the third indicator, the fourth indicator, the fifth indicator, and the sixth indicator are used as basic communication test data obtained to reflect the communication performance between different devices in the accelerator system.

[0037] In this implementation, by performing a benchmark communication test on the accelerator system, the H2D (Host to Device) data transmission rate, also known as the H2D bandwidth, and the average H2D latency can be obtained, serving as the first and second metrics, respectively. Correspondingly, the D2H (Device to Host) data transmission rate, also known as the D2H bandwidth, and the average D2H latency can be obtained, serving as the third and fourth metrics, respectively.

[0038] In practice, bandwidth testing tools can be used to test H2D and D2H bandwidth. Average latency can be tested by calculating the average latency of multiple data block transmissions.

[0039] In this implementation, by performing a benchmark communication test on the accelerator system, the data transmission rate and average latency between accelerators in the accelerator system can be determined. This means the bandwidth and average latency of inter-accelerator communication are used as the fifth and sixth metrics, respectively. In practical applications, the bandwidth and average latency between accelerators can be tested using the p2p Bandwidth Latency Test Tools.

[0040] It should also be noted that, considering that the bandwidth and average latency obtained may vary when the data block sizes being transmitted are different, a more accurate testing method is to test different data block sizes separately and then calculate the average value. In this regard, in one embodiment of the present invention, a benchmark communication test is performed on the accelerator system to obtain the data transmission rate and average latency when sending data from the host in the accelerator system to the accelerator in the accelerator system, which are used as the first indicator and the second indicator, respectively, including:

[0041] Performing a benchmark communication test on the accelerator system to obtain data transmission rates of different sizes of data sent from a host in the accelerator system to an accelerator in the accelerator system, and using an average value as a first indicator;

[0042] A benchmark communication test is performed on the accelerator system to obtain average delays when sending data of different sizes from a host in the accelerator system to an accelerator in the accelerator system, and the average value is used as the second indicator.

[0043] This implementation is explained using the first and second metrics, i.e., H2D bandwidth and average latency, as examples. The third and fourth metrics, i.e., D2H bandwidth and average latency, can also be calculated using the principles of this implementation. Similarly, the fifth and sixth metrics, i.e., bandwidth and average latency between different accelerators, can also be calculated using the principles of this implementation.

[0044] Specifically, in this implementation, the H2D bandwidth for sending data of different sizes is obtained, and the average value is used as the first metric. For example, three commonly used data block sizes, 128MB, 1GB, and 4GB, can be selected to effectively cover both small and large data scenarios. When a 128MB data block size is selected, 128MB data blocks are continuously sent from the host in the accelerator system to the accelerator in the accelerator system. The average bandwidth (GB / s) is measured using a bandwidth test tool, such as Bandwidth Test Tools, to obtain the data transfer rate for sending 128MB data blocks. Similarly, when a 1GB data block size is selected, the data transfer rate for sending 1GB data blocks can be obtained, and when a 4GB data block size is selected, the data transfer rate for sending 4GB data blocks can be obtained. The average of these three data transfer rates can be used as the first metric in this example.

[0045] Similarly, in this example, when a 128MB data block size is selected, 128MB data blocks are continuously sent from the host in the accelerator system to the accelerator in the accelerator system, for example, 1000 times. The latency (in microseconds) of each of these 1000 data blocks is measured, and the average value is used as the average latency when sending 128MB data from the host in the accelerator system to the accelerator in the accelerator system. Accordingly, the average latency when sending 1GB data blocks and the average latency when sending 4GB data blocks from the host in the accelerator system to the accelerator in the accelerator system can be obtained. The average of the average latencies corresponding to these three different data block sizes is then used as the final H2D average latency, which is the second metric.

[0046] In addition, it can be understood that the above selections are three commonly used data block sizes of 128MB, 1GB and 4GB. In other implementations, data blocks of other sizes can be set as needed, and the number is not limited to 3, and can be set according to actual needs.

[0047] As can be seen, this implementation takes into account the impact of data block size on the first and second indicators, thereby facilitating the obtained first and second indicators to more reasonably reflect the H2D bandwidth and average latency, thereby improving the reliability of the present application solution. Furthermore, as described above, the third, fourth, fifth, and sixth indicators can all be calculated according to the principles of this implementation, thereby improving the reliability of the present application solution.

[0048] The present application requires performing a collective communication test on the accelerator system to obtain collective communication test data reflecting the distributed communication operation execution performance of the accelerator system. Regarding the collective communication test, the specific test indicators can be set according to actual needs, but it is understood that they should be able to effectively reflect the distributed communication operation execution performance of the accelerator system. For example, the accelerator system can generally be allowed to perform common distributed communication operations to obtain the required collective communication test data.

[0049] For example, in a specific embodiment of the present invention, performing a collective communication test on the accelerator system as described in step S101 to obtain collective communication test data reflecting the distributed communication operation execution performance of the accelerator system may specifically include:

[0050] Controlling the accelerator system to perform a global reduction operation, and obtaining a data transmission rate and an operation delay when the accelerator system performs the global reduction operation, which are used as the seventh indicator and the eighth indicator respectively;

[0051] Controlling the accelerator system to perform an operation of broadcasting data from a single node to multiple nodes, and obtaining a broadcast time and a bandwidth utilization rate when the accelerator system performs the operation of broadcasting data from a single node to multiple nodes, as a ninth indicator and a tenth indicator respectively;

[0052] Control the accelerator system to perform a multi-node data full collection operation, and obtain the aggregation time and memory usage of the accelerator system when performing the multi-node data full collection operation, which are used as the eleventh and twelfth indicators respectively;

[0053] Controlling the accelerator system to perform a data sharding protocol operation, and obtaining a sharding transmission rate and a load balancing degree when the accelerator system performs the data sharding protocol operation, as the thirteenth indicator and the fourteenth indicator respectively;

[0054] The seventh index, the eighth index, the ninth index, the tenth index, the eleventh index, the twelfth index, the thirteenth index, and the fourteenth index are taken as the obtained collective communication test data for reflecting the distributed communication operation execution performance of the accelerator system.

[0055] In this implementation, the accelerator system will execute the global reduction operation (AllReduce), the operation of broadcasting data from a single node to multiple nodes (Broadcast), the multi-node data full collection operation (AllGather), and the data sharding reduction operation (ReduceScatter). This is because these four operations are required in many models and can effectively reflect the distributed communication operation execution performance of the accelerator system.

[0056] For global reduction operations (AllReduce), for example, calculations such as summation and maximum value can be selected, which are commonly used in distributed training parameter synchronization. In practical applications, pre-defined test tools can be used to control the accelerator system to execute global reduction operations without deploying a model. Similarly, for Broadcast, AllGather, and ReduceScatter operations, pre-defined test tools can also be used to control the accelerator system to execute them without deploying a model.

[0057] By controlling the accelerator system to perform global protocol operations, we can obtain the bandwidth of the accelerator system when performing global protocol operations, that is, the data transmission rate (GB / second) when performing global protocol operations, and the operation delay (milliseconds), which serve as the seventh and eighth indicators respectively.

[0058] Broadcasting data from a single node to multiple nodes (a common operation) is often used to initialize model parameters or configurations. By controlling the accelerator system to execute global reduction operations, we can obtain the broadcast time (milliseconds) and bandwidth utilization (%) when the accelerator system broadcasts data from a single node to multiple nodes, which serve as the ninth and tenth indicators, respectively.

[0059] Multi-node data allgather operations (AllGather) are commonly used to aggregate distributed inference results. By controlling the accelerator system to perform multi-node data allgather operations, we can obtain the aggregation time (milliseconds) and memory usage (MB) of the accelerator system during this operation, which serve as the eleventh and twelfth metrics, respectively.

[0060] The ReduceScatter operation is commonly used in large-scale model sharding training. By controlling the accelerator system to perform the ReduceScatter operation, we can obtain the shard transfer rate (GB / s) and load balance (%) during the operation, which serve as the thirteenth and fourteenth indicators, respectively.

[0061] The seventh to fourteenth indicators obtained by this implementation can effectively reflect the performance of the accelerator system when executing AllReduce, Broadcast, AllGather and ReduceScatter operations, and the performance of the accelerator system when executing AllReduce, Broadcast, AllGather and ReduceScatter operations can effectively represent the distributed communication operation execution performance of the accelerator system. Therefore, the collective communication test data obtained by this implementation can well reflect the distributed communication operation execution performance of the accelerator system, which further ensures the reliability of the solution of this application.

[0062] Furthermore, in a specific embodiment of the present invention, the accelerator system is controlled to perform a global reduction operation, and a data transmission rate when the accelerator system performs the global reduction operation is obtained as a seventh indicator, including:

[0063] Controlling the accelerator system to perform a global protocol operation, and obtaining a detection value of a data transmission rate when the accelerator system performs the global protocol operation through detection;

[0064] Determining whether the topology of the accelerator system is a non-fully connected topology;

[0065] If not, the detected value of the data transmission rate when the accelerator system performs the global reduction operation is used as the seventh indicator;

[0066] If yes, then the detected value of the data transmission rate when the accelerator system performs the global reduction operation is compensated according to the compensation method of A=B×(1+C / D), and the compensation result is used as the seventh indicator;

[0067] Among them, A represents the compensation result, B represents the detection value of the data transmission rate when the accelerator system performs the global reduction operation, D represents the actual number of hops when the accelerator system performs the global reduction operation, and C is the theoretical maximum number of hops.

[0068] This implementation also performs nonlinear compensation, specifically compensating for the data transfer rate during the global reduce operation (AllReduce). Specifically, the data transfer rate of the accelerator system during the global reduce operation is detected. It then determines whether the accelerator system's topology is non-fully connected, such as a tree structure.

[0069] If the topology is not fully connected, compensation is required. Specifically, the data transmission rate (B) measured during AllReduce execution on the accelerator system is compensated according to the formula A = B × (1 + C / D). (1 + C / D) is the compensation coefficient, reflecting the difference between the actual number of hops (D) and the theoretical maximum number of hops (C). In practice, data transmitted from accelerator card 1 to accelerator card 2 may need to pass through multiple switches, with each switch counted as one hop. Therefore, the theoretical maximum hop (C) is determined by the accelerator system topology. Specifically, the maximum number of switches that data must pass through can be determined based on the accelerator system topology. The actual number of hops (D) is the number of switches that a data transmission actually passes through during a particular data transmission. This number is affected by factors such as communication rules and, understandably, varies between different accelerators. Therefore, the actual number of hops (D) can be determined by calculating the average value through extensive testing. The smaller the actual number of hops (D), the better the accelerator system optimizes communication, and therefore, a larger compensation coefficient can be applied to the data transmission rate (B) measured during AllReduce execution.

[0070] Of course, if it is a fully connected topology, there is no need to introduce compensation, and the detection value of the data transmission rate when the accelerator system performs the global reduction operation can be directly used as the seventh indicator.

[0071] This implementation method takes into account that for non-fully connected topologies, the accelerator system with a smaller actual hop number D has better performance when executing AllReduce. Therefore, the data transmission rate of the AllReduce operation is compensated, which is conducive to further ensuring the reliability of the solution of this application.

[0072] In addition, in some implementations, if the accelerator system uses a customized communication library (such as MCCL, CNCL), which is beneficial to improving performance, the score value corresponding to the application performance test data can be increased by additional points, for example, an additional 5 points.

[0073] This application requires computing tests on the accelerator system to obtain computing test data that reflects the computing performance of the accelerator system. Regarding computing tests, specific test indicators can be set according to actual needs, but it is understood that they should be able to effectively reflect the computing performance of the accelerator system. For example, FLOPS (floating-point operations per second) can usually be used to measure the computing performance of the accelerator system.

[0074] In a specific embodiment of the present invention, performing a computing test on the accelerator system in step S101 to obtain computing test data reflecting the computing performance of the accelerator system may specifically include:

[0075] Performing a calculation test on the accelerator system to obtain the number of floating-point operations executed per second when the accelerator system performs calculations as the fifteenth indicator;

[0076] Based on the power consumption of the accelerator system when performing calculations, As the sixteenth indicator;

[0077] The fifteenth index and the sixteenth index are used as the obtained calculation test data for reflecting the calculation performance of the accelerator system;

[0078] Wherein, F is the number of floating-point operations executed per second when the accelerator system performs calculations, and G is the power consumption of the accelerator system when performing calculations.

[0079] In this implementation, FLOPS (floating point operations per second) is not the only indicator used in the calculation test data, but FLOPS is used as the fifteenth indicator, and the FLOPS under power consumption is also taken into account, that is, As the sixteenth indicator.

[0080] Since 1TFLOPS = 1×10 12 FLOPS, so the above formula It can also be expressed as G is the power consumption of the accelerator system during calculation. It can be seen that the lower the value of G, the larger the value of TFLOPS, which means that the number of floating-point operations that can be achieved within the unit power consumption is greater, and the value of the sixteenth indicator is larger at this time.

[0081] In practice, to effectively determine the FLOPS of an accelerator system, install HPL-AI v2.0 and configure it in double-precision (FP64) and mixed-precision (FP16 / FP32) modes. For example, set the matrix size to 100,000 and the block size to 256. After ensuring that the system passes HPL-AI's official validation (residual < 1e-16), proceed to the fifteenth metric. The sixteenth metric can be measured using a power analyzer.

[0082] Step S102: Model training and reasoning are performed through the accelerator system to obtain application performance test data used to reflect the model training and reasoning capabilities of the accelerator system, and energy efficiency test data used to reflect the energy efficiency of the accelerator system during model training and reasoning.

[0083] The present application solution takes into account that accelerator systems are usually used for model training or reasoning. Therefore, in addition to the communication performance between different devices in the accelerator system, the distributed communication operation execution performance, and the computing performance considered above, the model training and reasoning capabilities of the accelerator system should also be determined, and the impact of energy efficiency should be taken into account, so as to more comprehensively and effectively reflect the performance of the accelerator system during model training and reasoning, thereby ensuring that the present application solution can more comprehensively and accurately implement the performance test of the accelerator system.

[0084] Regarding model training and inference, specific test indicators can be set according to actual needs, but it is understandable that they should be able to effectively reflect the model training and inference capabilities of the accelerator system. For example, the throughput, latency, convergence time and other indicators of the accelerator system during model training and inference can usually be used as the detected application performance test data.

[0085] In a specific embodiment of the present invention, the step S102 of performing model training and inference by the accelerator system to obtain application performance test data reflecting the model training and inference capabilities of the accelerator system may specifically include:

[0086] The first model is trained using the accelerator system to obtain the number of images processed per second and the training convergence time of the accelerator system, which are used as the seventeenth and eighteenth indicators respectively;

[0087] The first model is inferred through the accelerator system to obtain the access request delay and the maximum number of access requests processed per second of the accelerator system, which are used as the nineteenth indicator and the twentieth indicator respectively;

[0088] The second model is trained using the accelerator system to obtain the number of prompt words processed per second by the accelerator system as the twenty-first indicator;

[0089] The second model is inferred through the accelerator system to obtain the number of prompt words generated per second and the delay in generating the first prompt word, which are used as the twenty-second and twenty-third indicators respectively;

[0090] The seventeenth indicator, the eighteenth indicator, the nineteenth indicator, the twentieth indicator, the twenty-first indicator, the twenty-second indicator, and the twenty-third indicator are used as application performance test data obtained for reflecting the model training and inference capabilities of the accelerator system;

[0091] Among them, the first model is an image processing model, and the second model is a language model.

[0092] This implementation method takes into account that when measuring the model training and reasoning capabilities of an accelerator system, the training and reasoning capabilities of the same accelerator system for different models are different, and image processing models and language models are currently typical models. Therefore, this implementation method uses image processing models and language models to more comprehensively and effectively measure the model training and reasoning capabilities of the accelerator system.

[0093] Specifically, to measure the accelerator system's model training and inference capabilities, a model deployment is required. The first model is an image processing model. For example, the typical image processing model, ResNet50, can be selected. The dataset, for example, is ImageNet-1K, with a batch size of 256. The testing tool, for example, is PyTorch 2.0 with AMP (Automatic Mixed Precision).

[0094] The accelerator system needs to deploy ResNet50. After training ResNet50 on the accelerator system, the accelerator system's training throughput, or the number of images processed per second (images / s), and the training convergence time, or the training convergence time (hours), can be obtained. These metrics serve as the seventeenth and eighteenth indicators, respectively. When subsequently scoring the seventeenth metric, for example, the score can be calculated using the formula "30 × (actual metric value / theoretical metric value)", where the theoretical metric value is set to 800 images / s. When scoring the eighteenth metric, for example, the score can be calculated using the formula "5 points for every hour faster than the reference time" (with an upper limit of 20 points, for example).

[0095] The accelerator system can also perform inference on the first model. Still taking ResNet50 inference as an example, the test tool can be an inference framework suitable for each accelerator (such as TensorRT), with the batch size set to 1 / 32. When the accelerator system performs ResNet50 inference, it can detect the accelerator system's maximum number of access requests processed per second (QPS). It can also detect the accelerator system's access request latency (milliseconds). Typically, the average latency of multiple access requests is used as the access request latency detected in this implementation. These two metrics serve as the nineteenth and twentieth metrics, respectively. When subsequently scoring the nineteenth metric, for example, the score can be calculated using the method "QPS ≥ benchmark value (1000) = 30 points." When scoring the twenty metrics, for example, the score can be calculated using the method "Access request latency ≤ 20ms = 20 points."

[0096] The second model is a language model. For example, the typical language model Llama-7B can be selected. The data set can be selected as IThe Pile, Batch Size=8, and Deepspeed Zero-3 and FP16 mixed precision configuration can be used.

[0097] The accelerator system requires the deployment of Llama-7B. After training Llama-7B on the accelerator system, the training throughput of the accelerator system, or the number of prompt words processed per second (tokens / s), can be calculated as the 21st metric. When subsequently scoring the 21st metric, for example, the score can be calculated using the formula "30 × (actual metric value / theoretical metric value)", where the theoretical metric value is set to 1500 tokens / s.

[0098] The accelerator system can also perform inference on the second model. Still using Llama-7B inference as an example, the test tool can be the vLLM engine with a batch size of 16. When the accelerator system performs Llama-7B inference, it can detect the accelerator system's prompt word generation rate per second (tokens / s) and the delay in generating the first prompt word (in milliseconds), which serve as the 22nd and 23rd indicators, respectively. When subsequently scoring the 22nd indicator, for example, the score can be calculated as follows: "When prompt word generation rate per second ≥ the theoretical value of the indicator (e.g., 50 tokens / s), the score is 30 points." When scoring the 23rd indicator, for example, the score can be calculated as follows: "When the delay in generating the first prompt word ≤ 200ms, the score is 20 points."

[0099] The design of this embodiment results in indicators 17 to 23 that can effectively reflect the accelerator system's training and reasoning capabilities for image processing models, as well as its training and reasoning capabilities for language models. The corresponding indicators are also set differently based on the respective characteristics of image processing models and language models. Image processing model training requires consideration of throughput and convergence, effectively reflecting the training capability of the image processing model, while image processing model reasoning requires consideration of the maximum number of access requests processed per second and access request latency, effectively reflecting the reasoning capability of the image processing model. Language model training requires consideration of throughput, and based on the characteristics of language model token generation, this embodiment measures the number of tokens processed per second, effectively reflecting the training capability of the language model. Language model reasoning requires consideration of the number of prompt words generated per second and the delay in generating the first prompt word, effectively reflecting the reasoning capability of the language model. In other words, this embodiment effectively demonstrates the performance of the accelerator system in deploying models, training, and reasoning in actual applications.

[0100] This application requires determining energy efficiency test data that reflects the energy efficiency of the accelerator system during model training and inference. Specific energy efficiency indicators can be set based on actual needs, but it is understood that they should be able to effectively reflect the energy efficiency of the accelerator system during model training and inference. For example, indicators such as the energy efficiency ratio of the accelerator system during model training and inference can generally be used as the detected energy efficiency test data.

[0101] In a specific embodiment of the present invention, the step S102 of performing model training and inference by the accelerator system to obtain energy efficiency test data reflecting the energy efficiency of the accelerator system during model training and inference may specifically include:

[0102] Performing model reasoning on the accelerator system to obtain the peak power consumption efficiency ratio of the accelerator system during model reasoning as the twenty-fourth indicator;

[0103] The 25th indicator is the average energy efficiency ratio of the accelerator system during model training and inference, obtained by performing model training and inference on the accelerator system.

[0104] Controlling the accelerator system to cyclically switch between no-load and full-load, and obtaining the power consumption response delay of the accelerator system as the twenty-sixth indicator;

[0105] The twenty-fourth indicator, the twenty-fifth indicator, and the twenty-sixth indicator are used as the energy efficiency test data obtained to reflect the energy efficiency of the accelerator system during model training and inference.

[0106] This implementation takes into account that peak power consumption is generally more likely to occur during model inference. Therefore, in this implementation, the accelerator system performs model inference, and the peak power consumption efficiency ratio of the accelerator system during model inference is obtained. For example, in one scenario, the accelerator system performs ResNet50 inference and obtains the peak power consumption efficiency ratio (QPS / power consumption) during inference. When subsequently scoring the twenty-four indicators, for example, the score can be calculated as "5 points for each 10% increase in the peak power consumption efficiency ratio above the industry average, with a maximum of 20 points."

[0107] In addition, in some implementations, the ratio of the number of floating-point calculations to power consumption (TFLOPS / power consumption) during inference of the accelerator system can also be used as an indicator parameter in the energy efficiency test data, which does not affect the implementation of the present invention.

[0108] The average energy efficiency ratio can be obtained by performing model training and inference on the accelerator system. For example, the accelerator system can be run with a mixed load, that is, some accelerators perform model training, and other accelerators perform model inference. For example, the mixed load can be run continuously for 24 hours, and the average energy efficiency ratio can be calculated. For example, the accelerator system can be trained for 1 hour, then the accelerator system performs model inference for 1 hour, and then the accelerator system performs model training for 1 hour, and so on for 24 hours. The average energy efficiency ratio of the accelerator system during model training and inference can also be calculated as the 25th indicator. When scoring the 25 indicators later, for example, the score can be calculated according to the method of "average energy efficiency ratio 10% + 5 points, upper limit 30 points".

[0109] This implementation also takes power consumption response delay into account, controlling the accelerator system to cycle between no-load and full-load, and obtaining the power consumption response delay (in milliseconds) of the accelerator system. It can be understood that a smaller power consumption response delay indicates that the accelerator system can start and shut down quickly, resulting in stronger energy efficiency. When subsequently scoring the twenty-six indicators, for example, the score could be calculated as follows: "20 points for a power consumption response delay ≤ 100ms, and 5 points deducted for each additional 50ms of delay."

[0110] Step S103: Based on a preset scoring rule, the scoring values ​​corresponding to the basic communication test data, the collective communication test data, the calculation test data, the application performance test data, and the energy efficiency test data are determined.

[0111] After determining the basic communication test data, collective communication test data, calculation test data, application performance test data and energy efficiency test data, the corresponding scoring values ​​can be determined according to the preset scoring rules. The specific content of the scoring rules can be set according to actual needs. Generally speaking, the upper limits of the scoring values ​​of the basic communication test data, collective communication test data, calculation test data, application performance test data and energy efficiency test data will be set to be consistent, for example, all set to 100 points.

[0112] In a specific embodiment of the present invention, determining the score value corresponding to the basic communication test data based on a preset scoring rule may specifically include:

[0113] Determine an indicator score for each of the first indicator, the second indicator, the third indicator, the fourth indicator, the fifth indicator, and the sixth indicator based on a comparison result between each of the first indicator, the second indicator, the third indicator, the fourth indicator, the fifth indicator, and the sixth indicator and the corresponding indicator theoretical value;

[0114] The respective indicator scores of the first indicator, the second indicator, the third indicator, the fourth indicator, the fifth indicator and the sixth indicator are weighted and superimposed, and the obtained result is used as the score value corresponding to the determined basic communication test data.

[0115] This implementation method takes into account that for the first to sixth indicators in the basic communication test data, the corresponding indicator theoretical values ​​can be used to score the first to sixth indicators, and then the weighted superposition of the respective indicator scores can be performed to conveniently determine the score value corresponding to the basic communication test data.

[0116] For example, for H2D bandwidth (the first metric) and D2H bandwidth (the third metric), achieving 80% of the theoretical bandwidth will result in a full score (e.g., 40 points). For every 5% decrease, 5 points will be deducted. For H2D average latency (the second metric) and D2H average latency (the fourth metric), scores can be calculated as 20 × (theoretical value / actual value). For inter-accelerator bandwidth (the fifth metric), achieving 70% of the theoretical bandwidth will result in a full score (e.g., 30 points). For every 5% decrease, 5 points will be deducted. For inter-accelerator average latency (the sixth metric), if the sixth metric is ≤ 1.5 times the theoretical value, 20 points will be awarded; otherwise, no points will be awarded (i.e., 0 points).

[0117] In actual applications, upper limits are usually set for the scores of the first to sixth indicators.

[0118] When weighting and superimposing the scores of the first to sixth indicators, the importance of the first to sixth indicators in measuring communication performance is generally the same. Therefore, in actual applications, the same weighting coefficient can usually be set for the first to sixth indicators. Of course, it can be set if necessary.

[0119] In a specific embodiment of the present invention, determining the score value corresponding to the collective communication test data based on a preset scoring rule may specifically include:

[0120] Determining a score of the accelerator system executing the global reduction operation based on a comparison result between the seventh indicator and the eighth indicator and the corresponding theoretical value of the indicator, as a first collective operation score;

[0121] Determining, based on a comparison result between each of the ninth indicator and the tenth indicator and the corresponding theoretical value of the indicator, a score of the accelerator system performing an operation of broadcasting data from a single node to multiple nodes, as a second set operation score;

[0122] Determining, based on the comparison results between the eleventh indicator and the twelfth indicator and the corresponding theoretical values ​​of the indicators, a score of the accelerator system performing the multi-node data full collection operation as a third set operation score;

[0123] Determine, based on the comparison results between the thirteenth indicator and the fourteenth indicator and the corresponding theoretical value of the indicator, a score of the accelerator system performing the data sharding reduction operation as a fourth set operation score;

[0124] For each preset application scenario, the superposition coefficients assigned to the first set operation score, the second set operation score, the third set operation score and the fourth set operation score in the application scenario are obtained, and weighted superposition is performed based on the superposition coefficients to obtain the score value corresponding to the determined set communication test data in the application scenario.

[0125] This implementation method takes into account that for the seventh to fourteenth indicators in the collective communication test data, the corresponding theoretical values ​​of the indicators can also be used to score the seventh to fourteenth indicators. Specifically, for the seventh indicator (data transmission rate) of the global reduction operation (AllReduce), the score can be calculated according to 30×(actual indicator value / theoretical indicator value), and for the eighth indicator (operation latency) of the global reduction operation (AllReduce), the score can be calculated according to 20×(theoretical indicator value / actual indicator value). The score of the seventh indicator can then be directly superimposed or weighted superimposed with the score of the eighth indicator to obtain the first collective operation score corresponding to the global reduction operation (AllReduce).

[0126] Similarly, for the operation (Broadcast) of broadcasting data from a single node to multiple nodes, the ninth metric (broadcast time) can be scored, for example, using 30 × (theoretical metric value / actual metric value). For the operation (Broadcast) of broadcasting data from a single node to multiple nodes, the tenth metric (bandwidth utilization) can be scored, for example, using 20 × (actual metric value / theoretical metric value). The scores of the ninth and tenth metrics can then be directly or weighted together to obtain the second set of operation scores corresponding to the operation (Broadcast) of broadcasting data from a single node to multiple nodes.

[0127] Similarly, for the 11th metric (aggregation time) of the multi-node data AllGather operation, the score can be calculated as 30 × (theoretical value of the metric / actual value of the metric). For the 12th metric (memory usage) of the multi-node data AllGather operation, the score can be calculated as 20 × (theoretical value of the metric / actual value of the metric). The scores of the 11th and 12th metrics can then be directly added or weighted together to obtain the third set operation score corresponding to the multi-node data AllGather operation.

[0128] Similarly, for the 13th metric (shard transfer rate) of the data sharding protocol operation (ReduceScatter), the score can be calculated as 30 × (actual metric value / theoretical metric value). For the 14th metric (load balance), the score can be calculated as 20 × (actual metric value / theoretical metric value). The scores of the 13th and 14th metrics can then be directly added or weighted together to obtain the score for the fourth set operation corresponding to the data sharding protocol operation (ReduceScatter).

[0129] In practical applications, upper limits of the first set operation score, the second set operation score, the third set operation score, and the fourth set operation score are usually set.

[0130] Furthermore, in this embodiment, for each preset application scenario, a superposition coefficient is assigned to each of the first, second, third, and fourth aggregate operation scores in that application scenario. Weighted superposition is then performed based on the superposition coefficients to obtain a score corresponding to the aggregate communication test data determined in that application scenario. Table 1 shows a table of superposition coefficients assigned to each of the first, second, third, and fourth aggregate operation scores in different scenarios.

[0131] Table 1: Overlay coefficient table under different scenarios

[0132]

[0133] This implementation method takes into account that although in the subsequent steps of the present application scheme, corresponding weight coefficients will be configured for the corresponding scoring values ​​of basic communication test data, collective communication test data, calculation test data, application performance test data and energy efficiency test data for different application scenarios, the above four operations of collective communication test data are also highly related to the application scenarios. Therefore, in this implementation method, after obtaining the collective communication test data and determining the first collective operation score, the second collective operation score, the third collective operation score and the fourth collective operation score representing the AllReduce operation, the Broadcast operation, the AllGather operation and the ReduceScatter operation respectively according to the corresponding indicators, the scoring values ​​corresponding to the collective communication test data in different application scenarios are directly determined according to different application scenarios, which is conducive to making the performance test results obtained in different application scenarios in this application more reasonable and accurate.

[0134] In a specific embodiment of the present invention, determining the score value corresponding to the calculated test data based on a preset scoring rule includes:

[0135] Based on the fifteenth indicator, the score is The calculation method is used to obtain the indicator score of the fifteenth indicator;

[0136] Based on the sixteenth indicator, the score is The calculation method is used to obtain the indicator score of the sixteenth indicator;

[0137] The indicator score of the fifteenth indicator is weighted and superimposed with the score of the sixteenth indicator, and the result obtained is used as the score value corresponding to the determined calculation test data;

[0138] Among them, k2 is the floating-point operation scoring parameter, k3 is a preset coefficient greater than 0 and less than 1, and E represents the theoretical value of the number of floating-point operations executed per second by the accelerator system. Indicates taking k2 and The smaller value of Indicates taking k4 and k1 is the energy efficiency scoring parameter, and k4 is the preset energy efficiency score threshold.

[0139] In this embodiment, the number of floating-point operations executed per second when the accelerator system performs calculations needs to be compared with the theoretical value E of the number of floating-point operations executed per second, and k3 is a preset coefficient greater than 0 and less than 1, so that the formula The numerator can be greater than the denominator.

[0140] In this implementation, a logarithmic function is used to score the fifteenth indicator. This is because the increase in F should increase the score of the fifteenth indicator, but there should be a limit so that when F increases more, the efficiency of increasing the score of the fifteenth indicator will decrease, which is equivalent to making the increase in F contribute to the improvement of the score of the fifteenth indicator. However, as F further increases, the score of the fifteenth indicator will not increase in the same proportion, which is equivalent to reducing the impact of excessive F on the score value of the final calculated test data, so that this nonlinear scoring rule can more accurately reflect the diminishing marginal performance characteristics. In other words, the design of the logarithmic function of this implementation method is conducive to the application scheme being able to more accurately and reasonably measure the score value of the calculated test data. Of course, there is also an upper limit to the score of the fifteenth indicator, namely k2.

[0141] In this implementation, sixteen indicators are used to calculate the score corresponding to the test data, serving as additional energy efficiency points. That is, the larger the F value per unit power consumption, the higher the additional energy efficiency points. Of course, an upper limit, k4, is also set for the additional energy efficiency points.

[0142] Step S104: For each preset application scenario, obtain a weight coefficient configured for each score value in the application scenario, and obtain a performance test result of the accelerator system in the application scenario based on the weight coefficient and each score value.

[0143] In actual applications, application scenarios can generally include at least one of training scenarios, inference scenarios, and edge device scenarios.

[0144] In the present application scheme, for each of the preset multiple application scenarios, it is necessary to obtain the weight coefficients configured for each score value in the application scenario and then perform weighted superposition to obtain the performance test results of the accelerator system in the application scenario. For example, in a specific embodiment of the present invention, in the training scenario, the weight coefficient configuration rule of X2>X3≥X1>X4 is satisfied; in the inference scenario, the weight coefficient configuration rule of X3>X4≥X2>X1 is satisfied; in the edge device scenario, the weight coefficient configuration rule of X4≥X3>X1>X2 is satisfied;

[0145] Among them, X1 represents the weight coefficient configured for the score value corresponding to the basic communication test data, and is the weight coefficient configured for the score value corresponding to the collective communication test data; X2 represents the weight coefficient configured for the score value corresponding to the calculation test data, X3 represents the weight coefficient configured for the score value corresponding to the application performance test data, and X4 represents the weight coefficient configured for the score value corresponding to the energy efficiency test data.

[0146] This implementation takes into account that, for any training scenario, both basic communication test data and aggregate communication test data reflect the overall communication capabilities of the accelerator system. Therefore, for any training scenario, the weight coefficient assigned to the score corresponding to the basic communication test data is consistent with the weight coefficient assigned to the score corresponding to the aggregate communication test data, both being X1. Furthermore, in some implementations, the score for the aggregate communication test data is already calculated based on the scenario, eliminating the need to distinguish between the weight coefficients for basic and aggregate communication test data, simplifying the setup.

[0147] For training scenarios, computational test data is more important, so the weight coefficient X2 assigned to the score corresponding to computational test data is the largest. Application performance test data is equally important as basic communication test data and aggregate communication test data, or may be slightly more important. Therefore, the weight coefficient X3 assigned to the score corresponding to application performance test data is ≥ X1. Energy efficiency test data is less important than other test data in training scenarios, so the weight coefficient X4 assigned to the score corresponding to energy efficiency test data is the smallest.

[0148] For example, in one case, for a training scenario, X2 is set to 40%, X3 is set to 20%, X1 is set to 15%, and X4 is set to 10%.

[0149] In inference scenarios, application performance test data is more important, so the weight coefficient X3 assigned to the score corresponding to application performance test data is the largest. Energy efficiency test data is equally important as computational test data, or even slightly more important, so the weight coefficient X4 assigned to the score corresponding to energy efficiency test data is ≥ X2. Basic communication test data and collective communication test data are less important than other test data in inference scenarios, so the weight coefficient X1 assigned to the score corresponding to basic communication test data and collective communication test data is the smallest.

[0150] For example, in one case, for the inference scenario, X3 is set to 45%, X4 is set to 25%, X2 is set to 20%, and X1 is set to 5%.

[0151] In edge device scenarios, due to their wide distribution and large number, energy efficiency test data is prioritized. Therefore, the weight coefficient X4 assigned to the score corresponding to energy efficiency test data is the largest. Application performance test data is slightly more important than basic communication test data and aggregate communication test data, so the weight coefficient X3 assigned to the score corresponding to application performance test data is greater than X1. Edge device scenarios do not require strong computing performance, so computing test data is not particularly important compared to other test data in edge device scenarios. Therefore, the weight coefficient X2 assigned to the score corresponding to computing test data is the smallest.

[0152] For example, in one scenario, for an edge device scenario, X4 is set to 35%, X3 is set to 25%, X1 is set to 15%, and X2 is set to 10%.

[0153] In actual applications, in addition to training scenarios, inference scenarios, and edge device scenarios, other types of application scenarios can also be set as needed.

[0154] In addition, in some implementations, the application scenario may also include a general scenario. This is to take into account that some accelerator systems may be used in multiple application scenarios simultaneously. For example, an accelerator system may be used in both training and inference scenarios. Therefore, the performance test results of the accelerator system in the general scenario can be obtained according to the weight coefficients configured for each score value in the general scenario, reflecting the comprehensive capabilities of the accelerator system in multiple application scenarios. It is understandable that for the general scenario, the values ​​of X1, X2, X3, and X4 can be relatively balanced, and the values ​​of X2 and X3 can be slightly higher.

[0155] In a specific embodiment of the present invention, it may further include:

[0156] For each preset application scenario, based on the obtained performance test results of the accelerator system in the application scenario, a performance grading result representing the performance level of the accelerator system in the application scenario is determined using preset grading rules.

[0157] In actual applications, for each preset application scenario, the performance test results of the accelerator system in that application scenario are usually presented in the form of a score, which may not be easy for users to understand. In this embodiment, based on the performance test results of the accelerator system in that application scenario, a performance grading result representing the performance level of the accelerator system in that application scenario is determined using preset grading rules, thereby making it easier for users to understand the performance test results.

[0158] The specific content of the preset classification rules can be determined according to actual needs, for example, they can be divided into intelligent computing center level, enterprise level, general level, edge level and development level. Please refer to Table 2 for a schematic diagram of the classification rules in a specific implementation method.

[0159] Table 2: Schematic table of classification rules

[0160]

[0161] By applying the technical solution provided by the embodiment of the present invention, in order to ensure the accuracy of the performance test results obtained, the accelerator system can be tested in multiple aspects respectively, so as to obtain more comprehensive evaluation indicators, which is conducive to more accurate performance testing of the accelerator system. Specifically, the present application solution will perform benchmark communication tests, collective communication tests and computing tests on the accelerator system respectively. Through the benchmark communication test, basic communication test data reflecting the communication performance between different devices in the accelerator system can be obtained. Through the collective communication test, collective communication test data reflecting the distributed communication operation execution performance of the accelerator system can be obtained. And through the computing test, computing test data reflecting the computing performance of the accelerator system can be obtained. It can be seen that in addition to the computing performance of the accelerator system, the present application also takes into account the communication performance between different devices and the distributed communication operation execution performance of the accelerator system, which is conducive to a more comprehensive evaluation of the accelerator system.

[0162] In addition, the present application scheme takes into account that the accelerator system is usually used for model training or reasoning. Therefore, by performing model training and reasoning on the accelerator system, application performance test data reflecting the model training and reasoning capabilities of the accelerator system and energy efficiency test data reflecting the energy efficiency of the accelerator system during model training and reasoning are obtained. This can effectively reflect the performance of the accelerator system during model training and reasoning, thereby ensuring that the present application scheme can more comprehensively and accurately implement the performance test of the accelerator system.

[0163] In addition, the present application also takes into account that the performance requirements for the accelerator system are different in different application scenarios. Therefore, based on the preset scoring rules, after determining the corresponding scoring values ​​of the basic communication test data, the collective communication test data, the calculation test data, the application performance test data, and the energy efficiency test data, in the present application scheme, each preset application scenario will be evaluated separately, that is, for each application scenario, the weight coefficient configured for each scoring value in the application scenario will be obtained, and then based on the weight coefficient and each scoring value, the performance test result of the accelerator system in the application scenario will be obtained. In other words, the present application scheme will determine the performance test results of the accelerator system in different scenarios.

[0164] In summary, the present application solution realizes the performance test of the accelerator system more comprehensively and accurately, and takes into account the differences in performance requirements for the accelerator system in different application scenarios, and thus determines the corresponding performance test results of the accelerator system in different scenarios.

[0165] Corresponding to the above method embodiments, embodiments of the present invention further provide an electronic device, a computer-readable storage medium, and a computer program product, which may refer to each other in correspondence with the above.

[0166] See also Figure 3 As shown, the device may include:

[0167] Memory 301, used for storing computer programs;

[0168] The processor 302 is configured to execute a computer program to implement the steps of the performance testing method for the accelerator system in any of the above embodiments.

[0169] The computer program product includes a computer program / instruction, which implements the steps of the performance testing method of the accelerator system in any of the above embodiments when executed by a processor.

[0170] See Figure 4 The computer-readable storage medium 40 stores a computer program 41. When executed by a processor, the computer program 41 implements the steps of the accelerator system performance testing method described in any of the above embodiments. The computer-readable storage medium 40 herein includes random access memory (RAM), internal memory, read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, or any other form of storage medium known in the art.

[0171] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0172] The above describes in detail the performance testing method, device, program product, and medium for the accelerator system provided by this application. Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is only intended to help understand the method and core concept of this application. It should be noted that, for those skilled in the art, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of this application.

Claims

1. A performance testing method for an accelerator system, characterized in that: include: Performing a baseline communication test, a collective communication test, and a computation test on the accelerator system, respectively, to obtain basic communication test data reflecting the communication performance between different devices in the accelerator system, collective communication test data reflecting the distributed communication operation execution performance of the accelerator system, and computation test data reflecting the computation performance of the accelerator system; Performing model training and reasoning on the accelerator system to obtain application performance test data reflecting the model training and reasoning capabilities of the accelerator system, and energy efficiency test data reflecting the energy efficiency of the accelerator system during model training and reasoning; Determine, based on preset scoring rules, scoring values ​​corresponding to the basic communication test data, the collective communication test data, the computing test data, the application performance test data, and the energy efficiency test data; For each preset application scenario, obtaining a weight coefficient configured for each of the score values ​​in the application scenario, and obtaining a performance test result of the accelerator system in the application scenario based on the weight coefficient and each of the score values; Performing a collective communication test on the accelerator system to obtain collective communication test data reflecting the distributed communication operation execution performance of the accelerator system includes: Controlling the accelerator system to perform a global reduction operation, and obtaining a data transmission rate and an operation delay when the accelerator system performs the global reduction operation, as a seventh indicator and an eighth indicator, respectively; Controlling the accelerator system to perform an operation of broadcasting data from a single node to multiple nodes, and obtaining a broadcast time and a bandwidth utilization rate when the accelerator system performs the operation of broadcasting data from a single node to multiple nodes, as a ninth indicator and a tenth indicator, respectively; Controlling the accelerator system to perform a multi-node data full collection operation, and obtaining an aggregation time and a memory usage size when the accelerator system performs the multi-node data full collection operation, which are used as an eleventh indicator and a twelfth indicator respectively; Controlling the accelerator system to perform a data sharding reduction operation, and obtaining a sharding transmission rate and a load balancing degree when the accelerator system performs the data sharding reduction operation, as a thirteenth indicator and a fourteenth indicator, respectively; The seventh indicator, the eighth indicator, the ninth indicator, the tenth indicator, the eleventh indicator, the twelfth indicator, the thirteenth indicator and the fourteenth indicator are used as the obtained collective communication test data for reflecting the distributed communication operation execution performance of the accelerator system.

2. The performance testing method of the accelerator system according to claim 1, characterized in that: Performing a baseline communication test on the accelerator system to obtain basic communication test data reflecting the communication performance between different devices in the accelerator system, including: Performing a benchmark communication test on the accelerator system to obtain a data transmission rate and an average delay when sending data from a host in the accelerator system to an accelerator in the accelerator system, as a first indicator and a second indicator, respectively; and obtaining a data transmission rate and an average delay when sending data from an accelerator in the accelerator system to a host in the accelerator system, as a third indicator and a fourth indicator, respectively; Performing a benchmark communication test on the accelerator system to obtain a data transmission rate and an average delay when data is transmitted between accelerators in the accelerator system, as a fifth indicator and a sixth indicator, respectively; The first indicator, the second indicator, the third indicator, the fourth indicator, the fifth indicator, and the sixth indicator are used as basic communication test data obtained to reflect the communication performance between different devices in the accelerator system.

3. The performance testing method of the accelerator system according to claim 2, characterized in that: A benchmark communication test is performed on the accelerator system to obtain a data transmission rate and an average delay when sending data from a host in the accelerator system to an accelerator in the accelerator system, as a first indicator and a second indicator, respectively, including: performing a benchmark communication test on the accelerator system to obtain data transmission rates when sending data of different sizes from a host in the accelerator system to an accelerator in the accelerator system, and using an average value as the first indicator; A benchmark communication test is performed on the accelerator system to obtain average delays when sending data of different sizes from a host in the accelerator system to an accelerator in the accelerator system, and the average value is used as the second indicator.

4. The performance testing method of the accelerator system according to claim 2, characterized in that: Determining a score value corresponding to the basic communication test data based on a preset scoring rule includes: Determine an indicator score for each of the first indicator, the second indicator, the third indicator, the fourth indicator, the fifth indicator, and the sixth indicator based on a comparison result between each of the first indicator, the second indicator, the third indicator, the fourth indicator, the fifth indicator, and the sixth indicator and the corresponding theoretical indicator value; The indicator scores of the first indicator, the second indicator, the third indicator, the fourth indicator, the fifth indicator and the sixth indicator are weighted and superimposed, and the obtained result is used as the score value corresponding to the determined basic communication test data.

5. The performance testing method of the accelerator system according to claim 1, characterized in that: Determining a score value corresponding to the aggregate communication test data based on a preset scoring rule includes: Determining a score of the accelerator system performing the global reduction operation based on a comparison result between the seventh indicator and the eighth indicator and the corresponding theoretical value of the indicator, as a first collective operation score; Determining, based on a comparison result between each of the ninth and tenth indicators and the corresponding theoretical indicator values, a score of the accelerator system performing an operation of broadcasting data from a single node to multiple nodes, as a second set operation score; Determining, based on a comparison result between the eleventh indicator and the twelfth indicator and the corresponding theoretical value of the indicator, a score of the accelerator system performing the multi-node data full collection operation as a third set operation score; Determining, based on a comparison result between the thirteenth indicator and the fourteenth indicator and the corresponding theoretical value of the indicator, a score of the accelerator system performing the data sharding reduction operation as a fourth set operation score; For each preset application scenario, the superposition coefficients assigned to the first set operation score, the second set operation score, the third set operation score and the fourth set operation score in the application scenario are obtained, and weighted superposition is performed based on the superposition coefficients to obtain the score value corresponding to the determined set communication test data in the application scenario.

6. The performance testing method of the accelerator system according to claim 1, characterized in that: Controlling the accelerator system to perform a global reduction operation, and obtaining a data transmission rate when the accelerator system performs the global reduction operation as a seventh indicator, including: controlling the accelerator system to perform a global protocol operation, and obtaining, through detection, a detection value of a data transmission rate when the accelerator system performs the global protocol operation; Determining whether the topology of the accelerator system is a non-fully connected topology; If not, using the detected value of the data transmission rate when the accelerator system performs the global reduction operation as the seventh indicator; If yes, then the detected value of the data transmission rate when the accelerator system performs the global reduction operation is compensated according to a compensation method of A=B×(1+C / D), and the compensation result is used as the seventh indicator; Among them, A represents the compensation result, B represents the detection value of the data transmission rate when the accelerator system performs the global reduction operation, D represents the actual number of hops when the accelerator system performs the global reduction operation, and C is the theoretical maximum number of hops.

7. The performance testing method of the accelerator system according to claim 1, characterized in that: Performing a computational test on the accelerator system to obtain computational test data reflecting the computational performance of the accelerator system includes: Performing a calculation test on the accelerator system to obtain the number of floating-point operations executed per second when the accelerator system performs calculations as the fifteenth indicator; Based on the power consumption of the accelerator system when performing calculations, As the sixteenth indicator; using the fifteenth indicator and the sixteenth indicator as the obtained computing test data for reflecting the computing performance of the accelerator system; Wherein, F is the number of floating-point operations executed per second when the accelerator system performs calculations, and G is the power consumption of the accelerator system when performing calculations.

8. The performance testing method of the accelerator system according to claim 7, characterized in that: Determining the score value corresponding to the calculated test data based on a preset scoring rule includes: Based on the fifteenth indicator, the score is The calculation method is used to obtain the indicator score of the fifteenth indicator; Based on the sixteenth indicator, the score is The calculation method is used to obtain the indicator score of the sixteenth indicator; Performing a weighted superposition of the indicator score of the fifteenth indicator and the sixteenth indicator, and obtaining a result as the score value corresponding to the determined calculated test data; Among them, k2 is the floating-point operation scoring parameter, k3 is a preset coefficient greater than 0 and less than 1, and E represents the theoretical value of the number of floating-point operations executed per second of the accelerator system. Indicates taking k2 and The smaller value of Indicates taking k4 and k1 is the energy efficiency scoring parameter, and k4 is the preset energy efficiency score threshold.

9. The performance testing method of the accelerator system according to claim 1, characterized in that: Performing model training and reasoning with the accelerator system to obtain energy efficiency test data reflecting the energy efficiency of the accelerator system during the model training and reasoning, including: Performing model reasoning using the accelerator system to obtain a peak power consumption efficiency ratio when the accelerator system performs model reasoning as a twenty-fourth indicator; Performing model training and inference using the accelerator system to obtain an average energy efficiency ratio of the accelerator system during model training and inference as the twenty-fifth indicator; controlling the accelerator system to cyclically switch between no-load and full-load, and obtaining a power consumption response delay of the accelerator system as a twenty-sixth indicator; The 24th indicator, the 25th indicator, and the 26th indicator are used as the energy efficiency test data obtained to reflect the energy efficiency of the accelerator system when performing model training and inference.

10. The performance testing method of the accelerator system according to claim 1, characterized in that: The application scenario includes at least one of a training scenario, an inference scenario, and an edge device scenario. In the training scenario, the weight coefficient configuration rule of X2>X3≥X1>X4 is satisfied; in the inference scenario, the weight coefficient configuration rule of X3>X4≥X2>X1 is satisfied; in the edge device scenario, the weight coefficient configuration rule of X4≥X3>X1>X2 is satisfied; Among them, X1 represents the weight coefficient configured for the score value corresponding to the basic communication test data, and is the weight coefficient configured for the score value corresponding to the collective communication test data; X2 represents the weight coefficient configured for the score value corresponding to the calculation test data, X3 represents the weight coefficient configured for the score value corresponding to the application performance test data, and X4 represents the weight coefficient configured for the score value corresponding to the energy efficiency test data.

11. The performance testing method of an accelerator system according to any one of claims 1 to 10, characterized in that: Model training and reasoning are performed by the accelerator system to obtain application performance test data reflecting the model training and reasoning capabilities of the accelerator system, including: Training the first model through the accelerator system to obtain the number of images processed per second and the training convergence time of the accelerator system as the seventeenth indicator and the eighteenth indicator respectively; Performing reasoning on the first model through the accelerator system to obtain an access request delay and a maximum number of access requests processed per second of the accelerator system as a nineteenth indicator and a twentieth indicator, respectively; Training the second model through the accelerator system to obtain the number of prompt words processed per second by the accelerator system as the twenty-first indicator; Performing inference on the second model through the accelerator system to obtain the number of prompt words generated per second and the delay in generating the first prompt word of the accelerator system, which serve as the twenty-second and twenty-third indicators respectively; The seventeenth indicator, the eighteenth indicator, the nineteenth indicator, the twentieth indicator, the twenty-first indicator, the twenty-second indicator, and the twenty-third indicator are used as the obtained application performance test data for reflecting the model training and inference capability of the accelerator system; The first model is an image processing model, and the second model is a language model.

12. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the performance testing method for an accelerator system according to any one of claims 1 to 11 when executing the computer program.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the performance testing method of the accelerator system according to any one of claims 1 to 11 are implemented.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the performance testing method of the accelerator system according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Method for testing training and reasoning performance of artificial intelligence acceleration card product

    CN116090552A