AI accelerator performance prediction method and device, computer equipment, readable storage medium and program product
By running AI models on benchmark AI accelerators, obtaining and calculating performance difference ratios, and predicting the performance of target AI accelerators, the high cost and long time of traditional evaluation methods are solved, achieving low-cost, fast and accurate performance evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-28
- Publication Date
- 2026-04-07
AI Technical Summary
Traditional methods for evaluating AI accelerator performance require high hardware procurement costs and are time-consuming, resulting in low testing efficiency.
By running AI models on benchmark AI accelerators, various underlying performance data are obtained, performance difference ratios and ratios are calculated, and the performance of target AI accelerators is predicted.
It reduces evaluation costs and testing time, improves testing efficiency and accuracy, and the prediction results are close to real application scenarios. The method has strong versatility.
Smart Images

Figure CN121807671A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence hardware performance evaluation, in particular to an AI accelerator performance prediction method and device, computer equipment, computer readable storage medium and computer program product. BACKGROUND
[0002] With the development of artificial intelligence technology, the scale and complexity of AI models have higher requirements for the computing power of underlying computing hardware. In order to meet these requirements, various special AI accelerators are designed and updated rapidly.
[0003] In order to evaluate the actual performance of a new model of AI accelerator, the traditional method is to obtain the physical hardware of the new model of AI accelerator, and deploy and run the AI model on the hardware for actual measurement.
[0004] However, the performance of the AI accelerator is evaluated by the actual measurement method, which not only needs high hardware procurement cost, but also consumes a lot of time and has low test efficiency. SUMMARY
[0005] Therefore, it is necessary to provide an AI accelerator performance prediction method, device, computer equipment, computer readable storage medium and computer program product which can quickly and accurately predict the actual performance of the AI accelerator at low cost.
[0006] In a first aspect, the present application provides an AI accelerator performance prediction method, comprising:
[0007] running an AI model on a benchmark AI accelerator to obtain a plurality of underlying performance data generated by the AI model when running;
[0008] obtaining the time length consumed by the underlying processing process corresponding to each underlying performance data; calculating the sum of the time length consumed by the underlying processing process corresponding to all underlying performance data to obtain the total time length consumed by the underlying processing;
[0009] calculating a first ratio between the time length consumed by the underlying processing process corresponding to each underlying performance data and the total time length consumed by the underlying processing;
[0010] obtaining a performance difference ratio between the target AI accelerator and the benchmark AI accelerator;
[0011] According to the performance difference ratio and the first ratio, a second ratio between the time length consumed by the target AI accelerator when performing the same underlying processing process as the benchmark AI accelerator and the total time length consumed by the underlying processing is calculated;
[0012] According to the first ratio and the second ratio, the performance multiple of the target AI accelerator compared with the benchmark AI accelerator is obtained.
[0013] In one embodiment, the types of underlying performance data include data computation, data transfer, and data communication.
[0014] In one embodiment, obtaining the performance difference ratio between the target AI accelerator and the benchmark AI accelerator includes:
[0015] The ratios between the theoretical peak computing power of the target AI accelerator and the theoretical peak computing power of the benchmark AI accelerator, the ratios between the theoretical memory bandwidth of the target AI accelerator and the theoretical memory bandwidth of the benchmark AI accelerator, and the ratios between the theoretical internal interconnect bandwidth of the target AI accelerator and the theoretical internal interconnect bandwidth of the benchmark AI accelerator are used as the performance difference ratios for data computing, data transfer, and data communication, respectively.
[0016] In one embodiment, a second ratio is calculated based on the performance difference ratio and a first ratio, between the time taken by the target AI accelerator to perform the same underlying processing as the benchmark AI accelerator and the total time taken for underlying processing, including:
[0017] The ratio between the first ratio corresponding to the data computation class and the performance difference ratio corresponding to the theoretical peak computing power is used to obtain the second ratio between the time consumed by the target AI accelerator in executing the underlying processing of the data computation class and the total time consumed by the underlying processing.
[0018] The ratio between the first ratio corresponding to the data transfer class and the performance difference ratio corresponding to the theoretical memory bandwidth is calculated to obtain the second ratio between the time spent by the target AI accelerator in executing the underlying processing of the data transfer class and the total time spent in underlying processing.
[0019] The ratio between the first ratio corresponding to data communication and the performance difference ratio corresponding to theoretical internal interconnection bandwidth is calculated to obtain the second ratio between the time consumed by the target AI accelerator in executing the underlying processing of data communication and the total time consumed by the underlying processing.
[0020] In one embodiment, obtaining the performance multiple of the target AI accelerator relative to the benchmark AI accelerator based on a first ratio and a second ratio includes:
[0021] The first sum of the first ratios corresponding to the data calculation class, the first ratios corresponding to the data transfer class, and the first ratios corresponding to the data communication class;
[0022] The second sum of the second ratios corresponding to the data calculation class, the data transfer class, and the data communication class;
[0023] Calculate the ratio between the first sum and the second sum to obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator.
[0024] In one embodiment, the underlying performance data for the data computation class is the performance data generated when the AI model performs mathematical operations; the underlying performance data for the data transfer class is the performance data generated when the AI model moves data between different storage levels; and the underlying performance data for the data communication class is the performance data generated when the AI model communicates between different processing kernels.
[0025] Secondly, this application also provides an AI accelerator performance prediction device, comprising:
[0026] The acquisition module is used to run AI models on benchmark AI accelerators and acquire various low-level performance data generated during the runtime of AI models.
[0027] The acquisition module is also used to acquire the underlying processing time corresponding to each type of underlying performance data; calculate the sum of the underlying processing time corresponding to all underlying performance data to obtain the total underlying processing time; and calculate the first ratio between the underlying processing time corresponding to each type of underlying performance data and the total underlying processing time.
[0028] The data processing module is used to obtain the performance difference ratio between the target AI accelerator and the benchmark AI accelerator;
[0029] The data processing module is also used to calculate a second ratio between the time taken by the target AI accelerator to perform the same underlying processing as the benchmark AI accelerator and the total time taken for underlying processing, based on the performance difference ratio and the first ratio.
[0030] The data processing module is also used to obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator based on the first ratio and the second ratio.
[0031] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0032] Run AI models on benchmark AI accelerators to obtain various low-level performance data generated during AI model runtime;
[0033] Obtain the processing time of the underlying process corresponding to each type of underlying performance data; calculate the sum of the processing times of the underlying processes corresponding to all underlying performance data to obtain the total processing time of the underlying process;
[0034] Calculate the first ratio between the time taken for the underlying processing of each type of underlying performance data and the total time taken for underlying processing;
[0035] Obtain the performance difference ratio between the target AI accelerator and the benchmark AI accelerator;
[0036] Based on the performance difference ratio and the first ratio, calculate the second ratio between the time taken by the target AI accelerator to perform the same underlying processing as the benchmark AI accelerator and the total time taken for underlying processing.
[0037] Based on the first ratio and the second ratio, obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator.
[0038] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0039] Run AI models on benchmark AI accelerators to obtain various low-level performance data generated during AI model runtime;
[0040] Obtain the processing time of the underlying process corresponding to each type of underlying performance data; calculate the sum of the processing times of the underlying processes corresponding to all underlying performance data to obtain the total processing time of the underlying process;
[0041] Calculate the first ratio between the time taken for the underlying processing of each type of underlying performance data and the total time taken for underlying processing;
[0042] Obtain the performance difference ratio between the target AI accelerator and the benchmark AI accelerator;
[0043] Based on the performance difference ratio and the first ratio, calculate the second ratio between the time taken by the target AI accelerator to perform the same underlying processing as the benchmark AI accelerator and the total time taken for underlying processing.
[0044] Based on the first ratio and the second ratio, obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator.
[0045] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0046] Run AI models on benchmark AI accelerators to obtain various low-level performance data generated during AI model runtime;
[0047] Obtain the processing time of the underlying process corresponding to each type of underlying performance data; calculate the sum of the processing times of the underlying processes corresponding to all underlying performance data to obtain the total processing time of the underlying process;
[0048] Calculate the first ratio between the time taken for the underlying processing of each type of underlying performance data and the total time taken for underlying processing;
[0049] Obtain the performance difference ratio between the target AI accelerator and the benchmark AI accelerator;
[0050] Based on the performance difference ratio and the first ratio, calculate the second ratio between the time taken by the target AI accelerator to perform the same underlying processing as the benchmark AI accelerator and the total time taken for underlying processing.
[0051] Based on the first ratio and the second ratio, obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator.
[0052] The aforementioned AI accelerator performance prediction method, apparatus, computer equipment, computer-readable storage medium, and computer program product first obtain various underlying performance data generated during the execution of the AI model by running an AI model on a benchmark AI accelerator. Second, they obtain the underlying processing time corresponding to each type of underlying performance data; calculate the sum of the underlying processing times corresponding to all underlying performance data to obtain the total underlying processing time; calculate a first ratio between the underlying processing time corresponding to each type of underlying performance data and the total underlying processing time; and obtain the performance difference ratio between the target AI accelerator and the benchmark AI accelerator. Third, based on the performance difference ratio and the first ratio, they calculate a second ratio between the time consumed by the target AI accelerator when performing the same underlying processing as the benchmark AI accelerator and the total underlying processing time. Finally, based on the first and second ratios, they obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator. Compared to traditional methods that evaluate AI accelerator performance through actual testing, which results in high testing costs and long processing times, this application uses performance data from benchmark AI accelerators and publicly available specifications of target AI accelerators to predict the performance of the target AI accelerator without obtaining its physical form. This reduces evaluation costs and testing time, and improves testing efficiency. Furthermore, predicting the target AI accelerator's performance based on the measured performance of benchmark AI accelerators makes the prediction results closer to real-world application scenarios, improving prediction accuracy. Moreover, the benchmark AI accelerators and AI models in this application are not limited to specific hardware and models, thus making the prediction method more versatile. Attached Figure Description
[0053] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0054] Figure 1 This is a diagram illustrating the application environment of an AI accelerator performance prediction method in one embodiment.
[0055] Figure 2 This is a flowchart illustrating an AI accelerator performance prediction method in one embodiment;
[0056] Figure 3 This is a flowchart illustrating the AI accelerator performance prediction method in another embodiment;
[0057] Figure 4 This is a structural block diagram of an AI accelerator performance prediction device in one embodiment;
[0058] Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0059] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0060] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0061] The AI accelerator performance prediction method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located on a cloud or other network server. Specifically, terminal 102 or server 104 completes an AI accelerator performance prediction method, which includes:
[0062] Run AI models on benchmark AI accelerators to obtain various low-level performance data generated during AI model runtime;
[0063] Obtain the processing time of the underlying process corresponding to each type of underlying performance data; calculate the sum of the processing times of the underlying processes corresponding to all underlying performance data to obtain the total processing time of the underlying process;
[0064] Calculate the first ratio between the time taken for the underlying processing of each type of underlying performance data and the total time taken for underlying processing;
[0065] Obtain the performance difference ratio between the target AI accelerator and the benchmark AI accelerator;
[0066] Based on the performance difference ratio and the first ratio, calculate the second ratio between the time taken by the target AI accelerator to perform the same underlying processing as the benchmark AI accelerator and the total time taken for underlying processing.
[0067] Based on the first ratio and the second ratio, obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator.
[0068] Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, and projection equipment. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0069] In one exemplary embodiment, such as Figure 2 As shown, a method for predicting the performance of an AI accelerator is provided, and this method is applied to... Figure 1 Taking the terminal in the example, the explanation includes the following steps 202 to 210. Wherein:
[0070] Step 202: Run the AI model on the benchmark AI accelerator to obtain various underlying performance data generated during the AI model's runtime.
[0071] Among them, underlying performance data refers to the time consumed by the AI model to execute each underlying operation.
[0072] For example, running an AI model on a benchmark AI accelerator generates multiple low-level operations. These operations can be categorized based on their functional attributes, primarily into data computation, data transfer, and data communication. Consequently, the underlying performance data corresponding to these operations can also be divided into three categories. Specifically, the types of underlying performance data include data computation, data transfer, and data communication.
[0073] For example, the underlying performance data of the data computation class is the performance data generated when the AI model performs mathematical operations; the underlying performance data of the data transfer class is the performance data generated when the AI model moves data between different storage levels; and the underlying performance data of the data communication class is the performance data generated when the AI model communicates between different processing kernels.
[0074] Among them, the operations performed by the AI model to perform mathematical operations include matrix multiplication, convolution, etc.; the operations of the AI model to move data between different storage levels include data copying from global memory slices to on-chip caches; the operations of the AI model to communicate between different processing kernels include synchronization and waiting operations between AI processing kernels or between control modules and computing modules.
[0075] Step 204: Obtain the time consumed by the underlying processing process corresponding to each type of underlying performance data; calculate the sum of the time consumed by the underlying processing process corresponding to all underlying performance data to obtain the total time consumed by the underlying processing; calculate the first ratio between the time consumed by the underlying processing process corresponding to each type of underlying performance data and the total time consumed by the underlying processing.
[0076] The total processing time at the lower level refers to the sum of the time spent executing each type of lower-level operation.
[0077] For example, running an AI model on a benchmark AI accelerator generates multiple low-level operations, and the execution time of each low-level operation during AI model training is recorded. It is understood that different low-level operations may correspond to the same type of low-level operation. Therefore, by summing the execution times of the same type of low-level operations, the execution time of the corresponding low-level processing for each type of low-level performance data can be obtained, namely, the execution time of low-level operations related to data computation, data transfer, and data communication.
[0078] For example, the execution time of data computation operations, data transfer operations, and data communication operations are calculated as a percentage of the total time spent on underlying processing, respectively, to obtain their respective first ratios. Based on these first ratios, a performance profile of the benchmark AI accelerator is obtained.
[0079] Step 206: Obtain the performance difference ratio between the target AI accelerator and the benchmark AI accelerator.
[0080] The performance of the target AI accelerator and the benchmark AI accelerator refers to the publicly disclosed hardware specifications, including FP32 computing power, HBM memory bandwidth, and chip characteristics. The performance difference ratio between the target AI accelerator and the benchmark AI accelerator includes the difference ratio in data computation, data transfer, and data communication.
[0081] Among them, the FP32 computing power of the accelerator is the core indicator for measuring the underlying operation capability of the accelerator in data computing, the HBM memory bandwidth of the accelerator is the core indicator for measuring the underlying operation capability of the accelerator in data transfer, and the chip characteristics of the accelerator are the core indicator for measuring the underlying operation capability of the accelerator in data communication.
[0082] For example, the ratios between the performance of the target AI accelerator and the corresponding performance of the benchmark AI accelerator are calculated to obtain the performance difference ratios between the target AI accelerator and the benchmark AI accelerator.
[0083] Step 208: Based on the performance difference ratio and the first ratio, calculate the second ratio between the time taken by the target AI accelerator to perform the same underlying processing as the benchmark AI accelerator and the total time taken for underlying processing.
[0084] The second ratio refers to the proportion of the execution time of data computation, data transfer, and data communication operations to the total time spent on underlying processing when running an AI model on the target AI accelerator.
[0085] For example, running an AI model on a target AI accelerator generates multiple low-level operations, which are the same as those generated when a benchmark AI accelerator runs an AI model. These operations also include low-level data computation operations, low-level data transfer operations, and low-level data communication operations. By calculating the ratio between a first ratio of each type of low-level operation and its corresponding performance difference ratio, a second ratio can be obtained between the time taken by the target AI accelerator to perform the same low-level processing and the total time taken for low-level processing.
[0086] Step 210: Based on the first ratio and the second ratio, obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator.
[0087] For example, for a benchmark AI accelerator, the first ratios of the acquired data computation, data transfer, and data communication classes are summed to obtain the total working time of the benchmark AI accelerator under a certain workload. For a target AI accelerator, the second ratios of the acquired data computation, data transfer, and data communication classes are summed to obtain the total working time of the target AI accelerator under the same workload. Based on the ratio between the total working time of the benchmark AI accelerator and the total working time of the target AI accelerator, the performance multiple of the target AI accelerator compared to the benchmark AI accelerator can be obtained.
[0088] In the aforementioned AI accelerator performance prediction method, firstly, the underlying performance data generated during the AI model's runtime is obtained by running an AI model on a benchmark AI accelerator. Secondly, a first ratio is obtained between the time consumed by the underlying processing for each type of underlying performance data and the total time consumed by the underlying processing; and the performance difference ratio between the target AI accelerator and the benchmark AI accelerator is also obtained. Thirdly, based on the performance difference ratio and the first ratio, a second ratio is calculated between the time consumed by the target AI accelerator when performing the same underlying processing and the total time consumed by the underlying processing. Finally, based on the first and second ratios, the performance multiple of the target AI accelerator compared to the benchmark AI accelerator is obtained. Compared to the traditional method of evaluating AI accelerator performance through actual testing, which results in high testing costs and long testing times, this application, using the performance data of the benchmark AI accelerator and the publicly disclosed specifications of the target AI accelerator, can predict the performance of the target AI accelerator without obtaining the physical entity of the target AI accelerator, thereby reducing evaluation costs and testing time and improving testing efficiency. Furthermore, predicting the performance of the target AI accelerator based on the measured performance of the benchmark AI accelerator makes the prediction results closer to real application scenarios, thus improving the accuracy of the prediction. Moreover, the benchmark AI accelerator and AI model in this application are not limited to specific hardware and models, thereby making the prediction method more universal.
[0089] In one embodiment, obtaining the performance difference ratio between the target AI accelerator and the benchmark AI accelerator includes: calculating the ratio between the theoretical peak computing power of the target AI accelerator and the theoretical peak computing power of the benchmark AI accelerator, the ratio between the theoretical memory bandwidth of the target AI accelerator and the theoretical memory bandwidth of the benchmark AI accelerator, and the ratio between the theoretical internal interconnect bandwidth of the target AI accelerator and the theoretical internal interconnect bandwidth of the benchmark AI accelerator, respectively, as the performance difference ratios for data computing, data transfer, and data communication.
[0090] For example, the publicly disclosed hardware specifications of the target AI accelerator and the benchmark AI accelerator include FP32 computing power, HBM memory bandwidth, and chip characteristics. The ratio between the theoretical peak computing power of the target AI accelerator and the theoretical peak computing power of the benchmark AI accelerator is calculated based on the FP32 computing power to obtain the performance difference ratio for data computation corresponding to the FP32 computing power. The ratio between the theoretical memory bandwidth of the target AI accelerator and the theoretical memory bandwidth of the benchmark AI accelerator is calculated based on the HBM memory bandwidth to obtain the performance difference ratio for data transfer corresponding to the HBM memory bandwidth. The ratio between the theoretical internal interconnect bandwidth of the target AI accelerator and the theoretical internal interconnect bandwidth of the benchmark AI accelerator is calculated based on the chip characteristics to obtain the performance difference ratio for data communication corresponding to the chip characteristics.
[0091] For example, since the target AI accelerator and the benchmark AI accelerator use chips from the same series and their internal architecture remains largely unchanged, the performance difference ratio between the theoretical internal interconnect bandwidth of the target AI accelerator and the theoretical internal interconnect bandwidth of the benchmark AI accelerator is set to 1 based on the chip characteristics. This yields a performance difference ratio of 1 for data communication classes corresponding to the chip characteristics.
[0092] In this embodiment, by using the publicly available hardware specifications of the target AI accelerator, the performance difference ratio between the target AI accelerator and the benchmark AI accelerator can be obtained without acquiring the physical entity of the target AI accelerator, thereby reducing the procurement cost of the target AI accelerator.
[0093] In one embodiment, a second ratio is calculated based on the performance difference ratio and a first ratio, between the time taken by the target AI accelerator to perform the same underlying processing as the benchmark AI accelerator and the total time taken for underlying processing, including:
[0094] The ratio between the first ratio corresponding to the data computation class and the performance difference ratio corresponding to the theoretical peak computing power is used to obtain the second ratio between the time consumed by the target AI accelerator in executing the underlying processing of the data computation class and the total time consumed by the underlying processing.
[0095] The ratio between the first ratio corresponding to the data transfer class and the performance difference ratio corresponding to the theoretical memory bandwidth is calculated to obtain the second ratio between the time spent by the target AI accelerator in executing the underlying processing of the data transfer class and the total time spent in underlying processing.
[0096] The ratio between the first ratio corresponding to data communication and the performance difference ratio corresponding to theoretical internal interconnection bandwidth is calculated to obtain the second ratio between the time consumed by the target AI accelerator in executing the underlying processing of data communication and the total time consumed by the underlying processing.
[0097] For example, the underlying processing of the data computation class performed by the target AI accelerator and the benchmark AI accelerator is the same, that is, the workload of the target AI accelerator and the benchmark AI accelerator in performing data computation operations is the same. Therefore, the ratio between the first ratio corresponding to the data computation class and the performance difference ratio corresponding to the theoretical peak computing power can be calculated to obtain the second ratio between the time spent by the target AI accelerator in performing the underlying processing of the data computation class and the total time spent in underlying processing.
[0098] For example, the underlying processing of the data transfer class performed by the target AI accelerator and the benchmark AI accelerator is the same, that is, the workload of the target AI accelerator and the benchmark AI accelerator in performing the data transfer class operation is the same. Therefore, the ratio between the first ratio corresponding to the data transfer class and the performance difference ratio corresponding to the theoretical memory bandwidth can be calculated to obtain the second ratio between the time spent by the target AI accelerator in performing the underlying processing of the data transfer class and the total time spent in the underlying processing.
[0099] For example, the underlying processing procedures for data communication classes performed by the target AI accelerator and the benchmark AI accelerator are the same, that is, the workload of data communication operations performed by the target AI accelerator and the benchmark AI accelerator is the same. Therefore, by calculating the ratio between the first ratio corresponding to data communication classes and the performance difference ratio corresponding to theoretical internal interconnection bandwidth, a second ratio can be obtained between the time spent by the target AI accelerator in performing the underlying processing procedures for data computation classes and the total time spent in underlying processing.
[0100] For example, the greater the performance difference between the target AI accelerator and the benchmark AI accelerator, the stronger the processing capability of the target AI accelerator when performing the underlying processing corresponding to that performance, and the less time it takes to complete the same amount of work.
[0101] In this embodiment, by using a first ratio of the time consumed by the benchmark AI accelerator in performing various low-level operations to the total processing time, and the performance difference ratio between the target AI accelerator and the benchmark AI accelerator, a second ratio of the time consumed by the target AI accelerator in performing various low-level operations to the total processing time can be obtained without actual testing of the target AI accelerator. This reduces the evaluation cost and testing time of the target AI accelerator and improves testing efficiency. Furthermore, since the first ratio is obtained after actual testing of the benchmark AI accelerator, the performance of the target AI accelerator can be predicted based on this first ratio, improving the accuracy of the prediction results.
[0102] In one embodiment, obtaining the performance multiple of the target AI accelerator relative to a benchmark AI accelerator based on a first ratio and a second ratio includes:
[0103] The first sum of the first ratios corresponding to the data calculation class, the first ratios corresponding to the data transfer class, and the first ratios corresponding to the data communication class;
[0104] The second sum of the second ratios corresponding to the data calculation class, the data transfer class, and the data communication class;
[0105] Calculate the ratio between the first sum and the second sum to obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator.
[0106] For example, a first sum is calculated between the first ratio corresponding to the data computation class, the first ratio corresponding to the data transfer class, and the first ratio corresponding to the data communication class, thereby obtaining the total time required for a benchmark AI accelerator to complete a certain workload. A second sum is calculated between the second ratio corresponding to the data computation class, the second ratio corresponding to the data transfer class, and the second ratio corresponding to the data communication class, thereby obtaining the total time required for a target AI accelerator to complete the same workload.
[0107] For example, the ratio between the first sum and the second sum is calculated to obtain the ratio of the time taken by the benchmark AI accelerator to the time taken by the target AI accelerator when completing the same amount of work, thereby obtaining the overall performance speedup ratio of the target AI accelerator compared to the benchmark AI accelerator.
[0108] In this embodiment, by using the first sum and the second sum, the overall performance speedup ratio of the target AI accelerator compared to the benchmark AI accelerator can be obtained without acquiring the physical hardware of the target AI accelerator. This allows for the prediction of the time consumed by the target AI accelerator when running the AI model, thereby reducing the prediction cost and prediction time for the target AI accelerator and improving efficiency.
[0109] like Figure 3 As shown, a specific embodiment illustrates the AI accelerator performance prediction method, including steps 302 to 312. Wherein,
[0110] Step 302: Run the AI model on the benchmark AI accelerator and obtain the underlying performance data generated during the AI model's runtime.
[0111] Specifically, when running an AI model on a benchmark AI accelerator, a performance analysis tool is integrated onto the AI model. This tool monitors and records the execution time of each underlying operation during the AI model training process, thereby obtaining the underlying performance data of the benchmark AI accelerator. The benchmark AI accelerator can be an Atlas 300T, the AI model can be a BERT (Bidirectional Encoder Representations from Transformers) deep learning model, and the performance analysis tool can be a Profiler tool; no specific restrictions are imposed here.
[0112] Step 304: Classify the underlying operations according to their functional attributes, and calculate the execution time of each type of underlying operation.
[0113] Specifically, based on functional attributes, the underlying operations can be divided into three categories: data calculation, data transfer, and data communication.
[0114] Among these, data computation refers to operations performed by the AI model to conduct mathematical calculations, including matrix multiplication and convolution. Data transfer refers to operations where the AI model moves data between different storage levels, including copying data from global memory to on-chip cache. Data communication refers to operations where the AI model communicates between different processing kernels, including synchronization and waiting operations between AI processing kernels or between the control module and the computing module.
[0115] The execution times of each underlying operation are summed up according to the category to obtain the execution time of each category of underlying operations, namely, the execution times of data calculation underlying operations, data transfer underlying operations, and data communication underlying operations.
[0116] Step 306: Calculate the first ratio between the execution time of each type of underlying operation of the benchmark AI accelerator and the total time spent running the AI model.
[0117] Specifically, the execution time of underlying data computation operations is divided by the total time spent running the AI model to obtain the data computation percentage. The data transfer percentage is obtained by dividing the execution time of the underlying data transfer operations by the total time spent running the AI model. The execution time of low-level data communication operations is divided by the total time spent running the AI model to obtain the data communication percentage. .
[0118] For example, when running the BERT model on the benchmark AI model Atlas 300T, the percentages of data computation, data transfer, and data communication are respectively... , as well as ,Right now , , .
[0119] Step 308: Calculate the performance difference ratio between the target AI accelerator and the benchmark AI accelerator.
[0120] Specifically, based on the publicly available hardware specifications of the target AI accelerator and the benchmark AI accelerator, including FP32 computing power, HBM memory bandwidth and chip characteristics, the performance ratio of each hardware specification parameter is calculated to obtain the performance difference ratio between the target AI accelerator and the benchmark AI accelerator.
[0121] Specifically, the FP32 computing power of the target AI accelerator is divided by the FP32 computing power of the benchmark AI accelerator to obtain the data computation difference ratio between the accelerators. The HBM memory bandwidth value of the target AI accelerator is divided by the HBM memory bandwidth value of the benchmark AI accelerator to obtain the data transfer difference ratio of the accelerators. The data communication difference ratio is obtained by dividing the chip characteristic value of the target AI accelerator by the chip characteristic value of the benchmark AI accelerator. .
[0122] Since the target AI accelerator and the benchmark AI accelerator use the same series of chips and their internal architecture remains largely unchanged, the chip characteristics of the target AI accelerator remain the same, thus yielding the data communication difference ratio. The value is 1.
[0123] For example, the benchmark AI accelerator could be the Atlas 300T, and the target AI accelerator could be the Atlas 300TA2. Based on the publicly available hardware specifications of the Atlas 300T and Atlas 300T A2, the data computation difference ratio can be calculated. Data transfer difference ratio Data communication difference ratio .
[0124] Step 310: Based on the first ratio and the performance difference ratio, obtain the second ratio between the execution time of each type of underlying operation of the target AI accelerator and the total time spent running the AI model.
[0125] Specifically, the data computation ratio of the benchmark AI accelerator will be... Divide by the data to calculate the difference ratio In order to obtain the data computing ratio of the target AI accelerator. The percentage of data transferred from benchmark AI accelerators. Divide by data transfer difference ratio In order to obtain the data transfer rate of the target AI accelerator. The proportion of data communication in benchmark AI accelerators. Divide by data communication difference ratio In order to obtain the data communication ratio of the target AI accelerator. .
[0126] For example, when the baseline AI model is Atlas 300T, the predicted percentage of data computation on the target AI accelerator Atlas 300T A2 is obtained when the BERT model is run on the target AI accelerator Atlas 300T A2. Data transfer ratio Data communication ratio They are respectively , , .Right now,
[0127]
[0128]
[0129]
[0130] Step 312: Based on the first ratio and the second ratio, obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator.
[0131] Specifically, the first sum of the first ratios corresponding to data computation, data transfer, and data communication is the data computation ratio of the benchmark AI accelerator. Data transfer ratio Data communication ratio The sum is accumulated to obtain the first sum.
[0132] The second sum of the second ratios corresponding to data computation, data transfer, and data communication is the data computation ratio of the target AI accelerator. Data transfer ratio Data communication ratio The sum is accumulated to obtain the second sum.
[0133] Divide the obtained first sum by the second sum to obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator.
[0134] For example, when the baseline AI model is Atlas 300T, the target AI accelerator is Atlas 300T A2, and the AI model is BERT, the first sum and the second sum are calculated respectively, that is,
[0135]
[0136]
[0137] For example, based on the ratio between the first sum and the second sum This indicates that the target AI accelerator Atlas 300T A2 has a performance factor of 3.0 times compared to the benchmark AI accelerator Atlas 300T. In other words, when running a BERT model on the target AI accelerator Atlas 300T A2, its actual performance is predicted to be approximately 3.0 times that of the benchmark AI accelerator Atlas 300T.
[0138] For example, a real-world test was conducted on the physical hardware of the target AI accelerator, Atlas 300T A2. Specifically, a BERT model was run on the Atlas 300T A2, and the execution time of each low-level operation during AI model training was monitored and recorded using the performance analysis tool Profiler, thereby obtaining the low-level performance data of the target AI accelerator. The measured low-level performance data of the target AI accelerator was compared with the measured low-level performance data of a benchmark AI accelerator, resulting in a measured performance multiple of 3.26x. Since the measured performance multiple of 3.26x is very close to the predicted performance multiple of 3.0x, this indicates that the AI accelerator performance prediction method proposed in this application has high accuracy and effectiveness.
[0139] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0140] Based on the same inventive concept, this application also provides an AI accelerator performance prediction device for implementing the AI accelerator performance prediction method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more AI accelerator performance prediction device embodiments provided below can be found in the limitations of the AI accelerator performance prediction method described above, and will not be repeated here.
[0141] In one exemplary embodiment, such as Figure 4 As shown, an AI accelerator performance prediction device 400 is provided, including: an acquisition module 402 and a data processing module 404, wherein:
[0142] The acquisition module is used to run AI models on benchmark AI accelerators and acquire various low-level performance data generated during the execution of AI models.
[0143] The acquisition module is also used to acquire the underlying processing time corresponding to each type of underlying performance data; calculate the sum of the underlying processing time corresponding to all underlying performance data to obtain the total underlying processing time; and calculate the first ratio between the underlying processing time corresponding to each type of underlying performance data and the total underlying processing time.
[0144] The data processing module is used to obtain the performance difference ratio between the target AI accelerator and the benchmark AI accelerator;
[0145] The data processing module is also used to calculate a second ratio between the time taken by the target AI accelerator to perform the same underlying processing as the benchmark AI accelerator and the total time taken for underlying processing, based on the performance difference ratio and the first ratio.
[0146] The data processing module is also used to obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator based on the first ratio and the second ratio.
[0147] In one embodiment, the acquisition module is further configured to acquire underlying performance data, wherein the types of underlying performance data include data calculation type, data transfer type, and data communication type.
[0148] In one embodiment, the data processing module is further configured to obtain the performance difference ratio between the target AI accelerator and the benchmark AI accelerator, including: calculating the ratio between the theoretical peak computing power of the target AI accelerator and the theoretical peak computing power of the benchmark AI accelerator, the ratio between the theoretical memory bandwidth of the target AI accelerator and the theoretical memory bandwidth of the benchmark AI accelerator, and the ratio between the theoretical internal interconnect bandwidth of the target AI accelerator and the theoretical internal interconnect bandwidth of the benchmark AI accelerator, respectively, as the performance difference ratios for data computing, data transfer, and data communication.
[0149] In one embodiment, the data processing module is further configured to calculate a second ratio between the time taken by the target AI accelerator to perform the same underlying processing procedure as the benchmark AI accelerator and the total time taken for underlying processing, based on the performance difference ratio and the first ratio, including:
[0150] The ratio between the first ratio corresponding to the data computation class and the performance difference ratio corresponding to the theoretical peak computing power is used to obtain the second ratio between the time consumed by the target AI accelerator in executing the underlying processing of the data computation class and the total time consumed by the underlying processing.
[0151] The ratio between the first ratio corresponding to the data transfer class and the performance difference ratio corresponding to the theoretical memory bandwidth is calculated to obtain the second ratio between the time spent by the target AI accelerator in executing the underlying processing of the data transfer class and the total time spent in underlying processing.
[0152] The ratio between the first ratio corresponding to data communication and the performance difference ratio corresponding to theoretical internal interconnection bandwidth is calculated to obtain the second ratio between the time consumed by the target AI accelerator in executing the underlying processing of data communication and the total time consumed by the underlying processing.
[0153] In one embodiment, the data processing module is further configured to obtain the performance multiple of the target AI accelerator relative to the benchmark AI accelerator based on a first ratio and a second ratio, including:
[0154] The first sum of the first ratios corresponding to the data calculation class, the first ratios corresponding to the data transfer class, and the first ratios corresponding to the data communication class;
[0155] The second sum of the second ratios corresponding to the data calculation class, the data transfer class, and the data communication class;
[0156] Calculate the ratio between the first sum and the second sum to obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator.
[0157] In one embodiment, the acquisition module is further configured to acquire underlying performance data of data computation, data transfer, and data communication classes. The underlying performance data of the data computation class is the performance data generated when the AI model performs mathematical operations; the underlying performance data of the data transfer class is the performance data generated when the AI model moves data between different storage levels; and the underlying performance data of the data communication class is the performance data generated when the AI model communicates between different processing kernels.
[0158] Each module in the aforementioned AI accelerator performance prediction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0159] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When executed by the processor, the computer program implements an AI accelerator performance prediction method.
[0160] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0161] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0162] Run AI models on benchmark AI accelerators to obtain various low-level performance data generated during AI model runtime;
[0163] Obtain the processing time of the underlying process corresponding to each type of underlying performance data; calculate the sum of the processing times of the underlying processes corresponding to all underlying performance data to obtain the total processing time of the underlying process;
[0164] Calculate the first ratio between the time taken for the underlying processing of each type of underlying performance data and the total time taken for underlying processing;
[0165] Obtain the performance difference ratio between the target AI accelerator and the benchmark AI accelerator;
[0166] Based on the performance difference ratio and the first ratio, calculate the second ratio between the time taken by the target AI accelerator to perform the same underlying processing as the benchmark AI accelerator and the total time taken for underlying processing.
[0167] Based on the first ratio and the second ratio, obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator.
[0168] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0169] The types of underlying performance data include data computation, data transfer, and data communication.
[0170] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0171] Obtain the performance difference ratio between the target AI accelerator and the benchmark AI accelerator, including:
[0172] The ratios between the theoretical peak computing power of the target AI accelerator and the theoretical peak computing power of the benchmark AI accelerator, the ratios between the theoretical memory bandwidth of the target AI accelerator and the theoretical memory bandwidth of the benchmark AI accelerator, and the ratios between the theoretical internal interconnect bandwidth of the target AI accelerator and the theoretical internal interconnect bandwidth of the benchmark AI accelerator are used as the performance difference ratios for data computing, data transfer, and data communication, respectively.
[0173] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0174] Based on the performance difference ratio and the first ratio, a second ratio is calculated between the time taken by the target AI accelerator to perform the same underlying processing as the benchmark AI accelerator and the total time taken for underlying processing, including:
[0175] The ratio between the first ratio corresponding to the data computation class and the performance difference ratio corresponding to the theoretical peak computing power is used to obtain the second ratio between the time consumed by the target AI accelerator in executing the underlying processing of the data computation class and the total time consumed by the underlying processing.
[0176] The ratio between the first ratio corresponding to the data transfer class and the performance difference ratio corresponding to the theoretical memory bandwidth is calculated to obtain the second ratio between the time spent by the target AI accelerator in executing the underlying processing of the data transfer class and the total time spent in underlying processing.
[0177] The ratio between the first ratio corresponding to data communication and the performance difference ratio corresponding to theoretical internal interconnection bandwidth is calculated to obtain the second ratio between the time consumed by the target AI accelerator in executing the underlying processing of data communication and the total time consumed by the underlying processing.
[0178] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0179] Based on the first ratio and the second ratio, the performance multiple of the target AI accelerator relative to the benchmark AI accelerator is obtained, including:
[0180] The first sum of the first ratios corresponding to the data calculation class, the first ratios corresponding to the data transfer class, and the first ratios corresponding to the data communication class;
[0181] The second sum of the second ratios corresponding to the data calculation class, the data transfer class, and the data communication class;
[0182] Calculate the ratio between the first sum and the second sum to obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator.
[0183] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0184] The underlying performance data for data computation is the performance data generated when the AI model performs mathematical operations; the underlying performance data for data transfer is the performance data generated when the AI model moves data between different storage levels; and the underlying performance data for data communication is the performance data generated when the AI model communicates between different processing kernels.
[0185] The implementation principle and technical effects of the above embodiments are similar to those of the above method embodiments, and will not be repeated here.
[0186] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0187] Run AI models on benchmark AI accelerators to obtain various low-level performance data generated during AI model runtime;
[0188] Obtain the processing time of the underlying process corresponding to each type of underlying performance data; calculate the sum of the processing times of the underlying processes corresponding to all underlying performance data to obtain the total processing time of the underlying process;
[0189] Calculate the first ratio between the time taken for the underlying processing of each type of underlying performance data and the total time taken for underlying processing;
[0190] Obtain the performance difference ratio between the target AI accelerator and the benchmark AI accelerator;
[0191] Based on the performance difference ratio and the first ratio, calculate the second ratio between the time taken by the target AI accelerator to perform the same underlying processing as the benchmark AI accelerator and the total time taken for underlying processing.
[0192] Based on the first ratio and the second ratio, obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator.
[0193] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0194] The types of underlying performance data include data computation, data transfer, and data communication.
[0195] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0196] Obtain the performance difference ratio between the target AI accelerator and the benchmark AI accelerator, including:
[0197] The ratios between the theoretical peak computing power of the target AI accelerator and the theoretical peak computing power of the benchmark AI accelerator, the ratios between the theoretical memory bandwidth of the target AI accelerator and the theoretical memory bandwidth of the benchmark AI accelerator, and the ratios between the theoretical internal interconnect bandwidth of the target AI accelerator and the theoretical internal interconnect bandwidth of the benchmark AI accelerator are used as the performance difference ratios for data computing, data transfer, and data communication, respectively.
[0198] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0199] Based on the performance difference ratio and the first ratio, a second ratio is calculated between the time taken by the target AI accelerator to perform the same underlying processing as the benchmark AI accelerator and the total time taken for underlying processing, including:
[0200] The ratio between the first ratio corresponding to the data computation class and the performance difference ratio corresponding to the theoretical peak computing power is used to obtain the second ratio between the time consumed by the target AI accelerator in executing the underlying processing of the data computation class and the total time consumed by the underlying processing.
[0201] The ratio between the first ratio corresponding to the data transfer class and the performance difference ratio corresponding to the theoretical memory bandwidth is calculated to obtain the second ratio between the time spent by the target AI accelerator in executing the underlying processing of the data transfer class and the total time spent in underlying processing.
[0202] The ratio between the first ratio corresponding to data communication and the performance difference ratio corresponding to theoretical internal interconnection bandwidth is calculated to obtain the second ratio between the time consumed by the target AI accelerator in executing the underlying processing of data communication and the total time consumed by the underlying processing.
[0203] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0204] Based on the first ratio and the second ratio, the performance multiple of the target AI accelerator relative to the benchmark AI accelerator is obtained, including:
[0205] The first sum of the first ratios corresponding to the data calculation class, the first ratios corresponding to the data transfer class, and the first ratios corresponding to the data communication class;
[0206] The second sum of the second ratios corresponding to the data calculation class, the data transfer class, and the data communication class;
[0207] Calculate the ratio between the first sum and the second sum to obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator.
[0208] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0209] The underlying performance data for data computation is the performance data generated when the AI model performs mathematical operations; the underlying performance data for data transfer is the performance data generated when the AI model moves data between different storage levels; and the underlying performance data for data communication is the performance data generated when the AI model communicates between different processing kernels.
[0210] The implementation principle and technical effects of the above embodiments are similar to those of the above method embodiments, and will not be repeated here.
[0211] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:
[0212] Run AI models on benchmark AI accelerators to obtain various low-level performance data generated during AI model runtime;
[0213] Obtain the processing time of the underlying process corresponding to each type of underlying performance data; calculate the sum of the processing times of the underlying processes corresponding to all underlying performance data to obtain the total processing time of the underlying process;
[0214] Calculate the first ratio between the time taken for the underlying processing of each type of underlying performance data and the total time taken for underlying processing;
[0215] Obtain the performance difference ratio between the target AI accelerator and the benchmark AI accelerator;
[0216] Based on the performance difference ratio and the first ratio, calculate the second ratio between the time taken by the target AI accelerator to perform the same underlying processing as the benchmark AI accelerator and the total time taken for underlying processing.
[0217] Based on the first ratio and the second ratio, obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator.
[0218] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0219] The types of underlying performance data include data computation, data transfer, and data communication.
[0220] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0221] Obtain the performance difference ratio between the target AI accelerator and the benchmark AI accelerator, including:
[0222] The ratios between the theoretical peak computing power of the target AI accelerator and the theoretical peak computing power of the benchmark AI accelerator, the ratios between the theoretical memory bandwidth of the target AI accelerator and the theoretical memory bandwidth of the benchmark AI accelerator, and the ratios between the theoretical internal interconnect bandwidth of the target AI accelerator and the theoretical internal interconnect bandwidth of the benchmark AI accelerator are used as the performance difference ratios for data computing, data transfer, and data communication, respectively.
[0223] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0224] Based on the performance difference ratio and the first ratio, a second ratio is calculated between the time taken by the target AI accelerator to perform the same underlying processing as the benchmark AI accelerator and the total time taken for underlying processing, including:
[0225] The ratio between the first ratio corresponding to the data computation class and the performance difference ratio corresponding to the theoretical peak computing power is used to obtain the second ratio between the time consumed by the target AI accelerator in executing the underlying processing of the data computation class and the total time consumed by the underlying processing.
[0226] The ratio between the first ratio corresponding to the data transfer class and the performance difference ratio corresponding to the theoretical memory bandwidth is calculated to obtain the second ratio between the time spent by the target AI accelerator in executing the underlying processing of the data transfer class and the total time spent in underlying processing.
[0227] The ratio between the first ratio corresponding to data communication and the performance difference ratio corresponding to theoretical internal interconnection bandwidth is calculated to obtain the second ratio between the time consumed by the target AI accelerator in executing the underlying processing of data communication and the total time consumed by the underlying processing.
[0228] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0229] Based on the first ratio and the second ratio, the performance multiple of the target AI accelerator relative to the benchmark AI accelerator is obtained, including:
[0230] The first sum of the first ratios corresponding to the data calculation class, the first ratios corresponding to the data transfer class, and the first ratios corresponding to the data communication class;
[0231] The second sum of the second ratios corresponding to the data calculation class, the data transfer class, and the data communication class;
[0232] Calculate the ratio between the first sum and the second sum to obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator.
[0233] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:
[0234] The underlying performance data for data computation is the performance data generated when the AI model performs mathematical operations; the underlying performance data for data transfer is the performance data generated when the AI model moves data between different storage levels; and the underlying performance data for data communication is the performance data generated when the AI model communicates between different processing kernels.
[0235] The implementation principle and technical effects of the above embodiments are similar to those of the above method embodiments, and will not be repeated here.
[0236] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0237] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0238] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0239] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for predicting the performance of an AI accelerator, characterized in that, The method includes: Run an AI model on a benchmark AI accelerator and obtain various underlying performance data generated during the runtime of the AI model. Obtain the processing time of the underlying process corresponding to each type of underlying performance data; calculate the sum of the processing times of the underlying processes corresponding to all underlying performance data to obtain the total processing time of the underlying process; Calculate a first ratio between the time taken for the underlying processing of each type of underlying performance data and the total time taken for the underlying processing; Obtain the performance difference ratio between the target AI accelerator and the benchmark AI accelerator; Based on the performance difference ratio and the first ratio, a second ratio is calculated between the time taken by the target AI accelerator to perform the same underlying processing as the benchmark AI accelerator and the total time taken for the underlying processing. Based on the first ratio and the second ratio, the performance multiple of the target AI accelerator compared to the benchmark AI accelerator is obtained.
2. The method according to claim 1, characterized in that, The types of underlying performance data include data calculation, data transfer, and data communication.
3. The method according to claim 2, characterized in that, The acquisition of the performance difference ratio between the target AI accelerator and the benchmark AI accelerator includes: The ratios between the theoretical peak computing power of the target AI accelerator and the theoretical peak computing power of the benchmark AI accelerator, the ratios between the theoretical memory bandwidth of the target AI accelerator and the theoretical memory bandwidth of the benchmark AI accelerator, and the ratios between the theoretical internal interconnect bandwidth of the target AI accelerator and the theoretical internal interconnect bandwidth of the benchmark AI accelerator are calculated and used as the performance difference ratios for data computing, data transfer, and data communication, respectively.
4. The method according to claim 3, characterized in that, The step of calculating a second ratio based on the performance difference ratio and the first ratio, between the time taken by the target AI accelerator to perform the same underlying processing as the benchmark AI accelerator and the total time taken for the underlying processing, includes: The ratio between the first ratio corresponding to the data computation class and the performance difference ratio corresponding to the theoretical peak computing power is used to obtain the second ratio between the time consumed by the target AI accelerator in executing the underlying processing of the data computation class and the total time consumed by the underlying processing. The ratio between the first ratio corresponding to the data transfer class and the performance difference ratio corresponding to the theoretical memory bandwidth is calculated to obtain the second ratio between the time consumed by the target AI accelerator in executing the underlying processing of the data transfer class and the total time consumed by the underlying processing. The ratio between the first ratio corresponding to the data communication category and the performance difference ratio corresponding to the theoretical internal interconnection bandwidth is calculated to obtain the second ratio between the time consumed by the target AI accelerator in executing the underlying processing of the data communication category and the total time consumed by the underlying processing.
5. The method according to claim 4, characterized in that, The step of obtaining the performance multiple of the target AI accelerator relative to the benchmark AI accelerator based on the first ratio and the second ratio includes: The first sum of the first ratios corresponding to the data calculation class, the first ratios corresponding to the data transfer class, and the first ratios corresponding to the data communication class; The second sum of the second ratios corresponding to the data calculation class, the data transfer class, and the data communication class; Calculate the ratio between the first sum and the second sum to obtain the performance multiple of the target AI accelerator compared to the benchmark AI accelerator.
6. The method according to claim 2, characterized in that, The underlying performance data for the data computation class is the performance data generated when the AI model performs mathematical operations; the underlying performance data for the data transfer class is the performance data generated when the AI model moves data between different storage levels; and the underlying performance data for the data communication class is the performance data generated when the AI model communicates between different processing kernels.
7. An AI accelerator performance prediction device, characterized in that, The device includes: The acquisition module is used to run an AI model on a benchmark AI accelerator and acquire various underlying performance data generated during the runtime of the AI model. The acquisition module is further configured to acquire the underlying processing time corresponding to each type of underlying performance data; calculate the sum of the underlying processing time corresponding to all underlying performance data to obtain the total underlying processing time; and calculate a first ratio between the underlying processing time corresponding to each type of underlying performance data and the total underlying processing time. The data processing module is used to obtain the performance difference ratio between the target AI accelerator and the benchmark AI accelerator; The data processing module is further configured to calculate a second ratio between the time consumed by the target AI accelerator when performing the same underlying processing procedure as the benchmark AI accelerator and the total time consumed by the underlying processing, based on the performance difference ratio and the first ratio. The data processing module is further configured to obtain the performance ratio of the target AI accelerator relative to the benchmark AI accelerator based on the first ratio and the second ratio.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.