Methods, computing devices, media, and program products for predicting performance of a computing system on which an application is run
By using tree-like structures and aggregation computing methods, the problem of traditional methods being unable to accurately predict performance bottlenecks in complex computing systems is solved, enabling performance prediction for systems lacking real hardware support.
Patent Information
- Application Number
- CN202511127782.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-08-12
AI Technical Summary
Traditional methods struggle to accurately predict performance bottlenecks in complex computing systems, especially when real-world hardware testing is lacking.
A tree structure is used to represent the hierarchical relationship of the computing system. The execution time of components is determined by aggregated computation, and the computing performance bottleneck is identified by combining the workload mapping relationship.
It can accurately predict the performance bottlenecks and throughput of complex computing systems without actual execution, and supports modeling of hardware components and complex situations at various granularities.
Smart Images

Figure CN120631705B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application generally relate to the field of artificial intelligence, and more particularly to a method, a computing device, a computer readable storage medium and a computer program product for predicting performance of a computing system on which an application is running. BACKGROUND
[0002] Conventional approaches for predicting performance of a computing system on which an application is running mainly include two kinds, the first one is based on Roofline or Gables model, and the second one is based on measurement.
[0003] In the first approach, the Roofline model is usually used to compare the operation intensity of an application with the theoretical peak performance of a computing system, so as to predict the performance bottleneck of the application. Specifically, the first approach includes, for example, selecting the minimum value among the peak computing performance, the operation intensity and the product of the peak memory bandwidth to represent the achievable performance of the application. In this approach, when the operation intensity is low, the application is limited by the memory bandwidth; when the operation intensity is high, the application is limited by the computing capability. Since the Roofline model is too simplified, it cannot handle complex situations such as multi-processor and multi-level memory. The Gables model is an extension of the Roofline model, which can handle multi-processor and multi-level storage situations, but still cannot handle bandwidth competition and other problems. For example, for artificial intelligence chips such as, but not limited to, Graphics Processing Units (GPUs) which are configured with tensor computing cores and vector computing cores and are configured with multi-level storage, the conventional prediction scheme based on the Gables model or the Roofline model is difficult to handle. In actual products, more complex situations may be faced. For example, there are different types of hardware in the computing system, and different types of workloads, different hardware design choices, etc., so it is difficult to use a simple model to accurately predict the performance bottleneck of complex actual applications.
[0004] In the second approach, the performance counter data is mainly collected by running benchmark programs or actual applications on real hardware. In this approach, various processors provide hardware performance counters (Performance Counters, PFC) and their analysis tools. However, this second approach can only be used to detect the actual performance under the current configuration, and cannot guide the design space exploration and optimization direction selection; and this method needs to be performed on real hardware, and cannot be applied to scenarios where hardware support is lacking in the early design stage.
[0005] In summary, the conventional scheme for predicting the performance of a computing system on which an application is running has the disadvantage that it is difficult to make accurate performance prediction for complex applications that lack real hardware running measurement support. SUMMARY
[0006] The present application provides a method, a computing device, a computer readable storage medium and a computer program product for predicting the performance of a computing system on which an application is running, which can make accurate performance prediction even for complex applications that lack real hardware running measurement support.
[0007] According to a first aspect of the present application, there is provided a method for predicting the performance of a computing system on which an application is running. The method comprises: characterizing a hierarchical relationship of the computing system on which the application is running using a tree structure for aggregated computation, the computing system comprising a plurality of components; computing a workload of the application; determining a mapping relationship of the components comprised by the computing system to the workload; and based on the mapping relationship and the tree structure, computing execution time of the components via aggregated computation so as to determine a performance bottleneck of the computing system on which the application is running.
[0008] In some embodiments, the characterizing the hierarchical relationship of the computing system on which the application is running using the tree structure comprises: configuring a hardware model of the tree structure such that a root node of the tree structure characterizes the computing system on which the application is running, and a plurality of leaf nodes at one or more levels below the root node respectively characterize a plurality of components comprised by the computing system; and configuring, for the components corresponding to the leaf nodes, at least one of a component identifier, and a throughput value, an aggregation method and a sub-component set.
[0009] In some embodiments, the configuring, for the components corresponding to the leaf nodes, at least one of the component identifier, and the throughput value, the aggregation method and the sub-component set comprises: configuring, for the components corresponding to the leaf nodes at a bottom level of a component path, the throughput value, the component path being used to indicate the components and associated paths of sub-components comprised by the components in the tree structure; and configuring, for the components corresponding to the leaf nodes at a non-bottom level of the component path, the aggregation method.
[0010] In some embodiments, the computing the workload of the application comprises any one of: computing the workload via an equation; computing the workload via an intermediate representation of a compiler; and dynamically computing the workload via runtime performance metrics.
[0011] In some embodiments, the aggregation method comprises: a maximum value aggregation mode used for aggregated computation of execution time of components for exclusive hardware resources; and an accumulation aggregation mode used for aggregated computation of execution time of components for shared hardware resources.
[0012] In some embodiments, the calculating the execution time of the component based on the mapping relationship and the tree structure via the aggregated calculation comprises: in response to determining that the aggregation method of the component corresponding to the leaf node of the previous level of the tree structure is the maximum value aggregation mode, selecting the maximum value of the execution time of the component in the current level as the execution time of the component corresponding to the leaf node of the previous level; and in response to determining that the aggregation method of the component corresponding to the leaf node of the previous level of the tree structure is the cumulative aggregation mode, cumulating the execution time of the component corresponding to all the leaf nodes of the current level, so as to take the cumulation result as the execution time of the component corresponding to the leaf node of the previous level.
[0013] In some embodiments, the calculating the workload of the application comprises: generating a workload model based on a plurality of workload elements, each workload element indicating an operation quantity of a predetermined type of operation or a data transmission quantity.
[0014] In some embodiments, the application comprises a loop structure, and the calculating the workload of the application comprises: for each level of loop from the innermost loop to the outer loop of the loop structure, sequentially calculating the flow required for executing each component of the application.
[0015] In some embodiments, the determining the performance bottleneck of the computing system running the application comprises: calculating the execution time of the component corresponding to the leaf node of the bottom layer of the tree structure, so as to sequentially calculate the execution time of the component from the bottom layer of the tree structure along the component path upwards; and based on the calculated execution time of the component, determining the performance bottleneck of the computing system running the application and the total execution time of the computing system.
[0016] In some embodiments, the determining the performance bottleneck of the computing system running the application comprises: comparing the execution time of each component in the current level, so as to take the component path where the component with the longest execution time is located as the performance bottleneck.
[0017] In some embodiments, the determining the mapping relationship of the component included in the computing system to the workload comprises: in response to determining that the types of the components configured in the computing system are different, making the mapping manners of the same workload to different types of components different.
[0018] In some embodiments, the determining the performance bottleneck of the computing system running the application comprises: comparing the execution time of each component in the current level, so as to take the component path where the component with the longest execution time is located as the performance bottleneck.
[0019] According to a second aspect of the present application, there is also provided a method for predicting performance of a computing system running an application. The method comprises: representing hierarchical relationships of a first computing system and a second computing system respectively executing the application using a tree structure for aggregate computation, the first computing system and the second computing system respectively comprising a plurality of components; computing a workload of the application; determining mapping relationships of the components comprised in the first computing system and the second computing system to the workload respectively; and determining performance bottlenecks and total execution times of the first computing system and the second computing system respectively, so as to compare performance of the first computing system and the second computing system, the performance bottlenecks and the total execution times of the first computing system and the second computing system being derived via the mapping relationships and the tree structure respectively and via the aggregate computation.
[0020] According to a third aspect of the present application, there is also provided a computing device. The computing device comprises: at least one processor; and a memory communicatively connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the computing device to perform the method of the first aspect or the second aspect of the present application.
[0021] According to a fourth aspect of the present application, there is also provided a computer readable storage medium. The computer readable storage medium stores a computer program, the computer program being executed by a machine to perform the method of the first aspect or the second aspect of the present application.
[0022] According to a fifth aspect of the present application, there is also provided a computer program product comprising a computer program, the computer program being executed by a machine to perform the method of the first aspect or the second aspect of the present application.
[0023] The present application can support modeling of hardware components of various granularities, and complex situations such as multi-processor, multi-level memory, etc. by representing hierarchical relationships of a computing system comprising a plurality of components running an application using a tree structure for aggregate computation. In addition, the present application can accurately predict performance of a throughput-limited computing system without actually executing by computing a workload of the application; determining mapping relationships of the components comprised in the computing system to the workload; and converting the workload to execution times of each component for aggregate computation, so as to determine performance bottlenecks of the system. Thus, the present application can make accurate performance prediction even for complex applications lacking real hardware running measurement support.
[0024] It should be understood that nothing in this section is intended to limit the scope of the embodiments of the present application. Other aspects of the present application will become apparent to those skilled in the art upon reading the following specification and / or upon inspection of the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0025] The above and other features, aspects and advantages of embodiments of the present application will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings. In the drawings similar or common elements of the drawings are denoted by like reference numbers.
[0026] Figure 1 A schematic diagram of a computing device implementing a method for predicting performance of a computing system on which an application is running according to an embodiment of the present application is shown schematically.
[0027] Figure 2 A flowchart of a method for predicting performance of a computing system on which an application is running according to some embodiments of the present application is shown.
[0028] Figure 3 A schematic diagram of a tree structure for characterizing hierarchical relationships of a computing system on which an application is running according to an embodiment of the present application is shown.
[0029] Figure 4 A flowchart of a method for predicting performance of a computing system on which an application is running according to some other embodiments of the present application is shown.
[0030] In the various drawings, like or corresponding elements are denoted by like or corresponding reference numbers. DETAILED DESCRIPTION
[0031] Preferred embodiments of the present application will be described herein below with reference to the accompanying drawings. While preferred embodiments of the present application are shown in the drawings, it is understood that the present application can be embodied in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present application to those skilled in the art.
[0032] The term "comprising" and variations thereof as used herein are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to those elements, but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. The term "based on" means "based, at least in part, on." The terms "one example embodiment" and "an example embodiment" mean "at least one example embodiment." The term "another embodiment" means "at least one additional embodiment." The terms "a first," "a second," etc. do not require a strict numbering of their objects but are used merely as labels for distinct objects.
[0033] As described previously, conventional approaches for predicting performance of a computing system on which an application is running suffer from the drawback that it is difficult to make accurate performance predictions for complex applications that lack real hardware run-time support.
[0034] To at least partially solve one or more of the above problems and other potential problems, example embodiments of the present application propose a scheme for predicting the performance of a computing system on which an application is running. In this scheme, by using a tree structure to represent the hierarchical relationship of a computing system on which an application is running, including a plurality of components, for aggregated computation, the present application can support hardware components of multiple granularities, and modeling of complex situations such as multi-processors, multi-level memories, etc. In addition, by computing a workload of the application; determining a mapping relationship of components included in the computing system to the workload; and based on the mapping relationship and the tree structure, computing execution times of the components via aggregated computation, in order to determine a performance bottleneck of the computing system on which the application is running, the present application can accurately predict the performance of a throughput-limited computing system without actually executing. Thus, the present application can make accurate performance prediction even for complex applications that lack real hardware running measurement support.
[0035] Figure 1 A schematic diagram of a computing device 100 implementing a method for predicting the performance of a computing system on which an application is running according to an embodiment of the present application is schematically shown. As shown in FIG. 1, the computing device 100 includes a processor 101, a memory 102, and a bus 103. The processor 101 is connected to the memory 102 via the bus 103. The memory 102 stores a program 104 for predicting the performance of a computing system on which an application is running. The program 104 includes a program code for implementing the method for predicting the performance of a computing system on which an application is running according to an embodiment of the present application. Figure 1As shown, the computing device 100 can have one or more processing units, including special-purpose processing units such as a Graphics Processing Unit (GPU), a Field Programmable Gate Array (FPGA), an Application Specific Integrated Circuit (ASIC), a General-purpose computing on graphics processing units (GPGPU), or a tensor processing unit (TPU), and general-purpose processing units such as a CPU. The computing device 100 further includes at least a tree structure characterization unit 102, a workload computation unit 104, a component-to-workload mapping relationship determination unit 106, and a performance bottleneck determination unit 108. It should be appreciated that the above-mentioned tree structure characterization unit 102, workload computation unit 104, component-to-workload mapping relationship determination unit 106, and performance bottleneck determination unit 108 can be software modules, which are executed on one or more processing units configured in the computing device 100. The one or more processing units can be integrated with or separate from the plurality of computing units used for predicting the performance of the computing system on which the application is running.
[0036] Regarding the tree structure characterization unit 102, it is configured to characterize the hierarchical relationship of the computing system on which the application is running using a tree structure, for aggregated computation, the computing system including a plurality of components.
[0037] Regarding the workload computation unit 104, it is configured to compute the workload of the application.
[0038] Regarding the component-to-workload mapping relationship determination unit 106, it is configured to determine the mapping relationship of the components included in the computing system to the workload.
[0039] Regarding the performance bottleneck determination unit 108, it is configured to compute the execution time of the components based on the mapping relationship and the tree structure, via aggregated computation, so as to determine the performance bottleneck of the computing system on which the application is running.
[0040] The method 200 for predicting the performance of the computing system on which the application is running will be described below in conjunction with Figure 2 Figure 3 an embodiment of the present application.Figure 2 A flowchart of a method 200 for predicting performance of a computing system on which an application is running according to some embodiments of the present application is shown. It should be understood that the method 200 can be performed, for example, at the computing device 100 Figure 1 described. The method 200 can also include additional actions not shown and / or can omit actions shown, without limitation in this regard.
[0041] At step 202, the computing device 100 characterizes a hierarchical relationship of a computing system on which an application is running using a tree structure for aggregating computation, the computing system including a plurality of components.
[0042] As to the method of characterizing the hierarchical relationship of the computing system on which the application is running using the tree structure, it includes, for example: the computing device 100 configures a hardware model of the tree structure such that a root node of the tree structure characterizes the computing system on which the application is running, and a plurality of leaf nodes of one or more levels under the root node respectively characterize a plurality of components included by the computing system; and the computing device 100 configures, for a component corresponding to a leaf node, at least one of a component identifier, a throughput value, an aggregation method, and a sub-component set. For example, the computing device 100 configures, for a component corresponding to a leaf node at a bottom level of a component path, a throughput value, the component path being used to indicate an associated path of the component and sub-components included by the component in the tree structure; and the computing device 100 configures, for a component corresponding to a leaf node at a non-bottom level of the component path, an aggregation method. It should be understood that the computing device 100 determines a number of levels included by the tree structure and a number of leaf nodes included by each level based on a type of the computing device and a granularity of components included by the computing system.
[0043] For example, Figure 3 A schematic diagram of a tree structure 300 for characterizing a hierarchical relationship of a computing system on which an application is running according to an embodiment of the present application is shown. As Figure 3 shown, the tree structure 300 characterizes a hierarchical relationship of components included by a computing system (the computing system being, for example, a GPU). The tree structure 300 includes a plurality of levels. A node of a topmost level is a root node. A plurality of leaf nodes of each level under the root node, each leaf node characterizing a component included by the computing system. As Figure 3As shown, the root node indicated by label 310 represents, for example, a computing system (its component is identified as "system"). The next level down from the root node includes three leaf nodes: the first leaf node, the second leaf node, and the third leaf node. The first leaf node indicated by label 320 represents a computing unit (its component is identified as "compute"), the second leaf node indicated by label 322 represents memory (its component is identified as "memory"), and the third leaf node indicated by label 324 represents the internal interconnect (its component is identified as "interconnect"). It should be understood that each component represented by a non-lower-level leaf node may further include components (or "sub-components") represented by one or more lower-level leaf nodes. For example, the computational unit corresponding to the first leaf node further includes: a tensor kernel represented by the fourth leaf node indicated by label 330 (whose component is identified as "tensor_cores"), a SIMT (Compute Unified Device Architecture) kernel represented by the fifth leaf node indicated by label 332 (whose component is identified as "SIMT_cores"), and a special function unit represented by the sixth leaf node indicated by label 334 (whose component is identified as "sfu"), the special function unit being used to predetermine a transcendental function.
[0044] like Figure 3 As shown, the association paths between a compute unit and all its further included components (or "sub-components") constitute a component path about the compute unit. It should be understood that a component path indicates the association path between a component and its included sub-components in a tree structure. The component corresponding to the leaf node at the bottom of the component path is configured with a throughput value. For example, in the leaf node at the bottom of the component path about the compute unit (such as...) Figure 3 The component corresponding to the bottom-level leaf node indicated by marker 340 (e.g., fp8 floating-point precision type) is configured with a throughput value. This throughput value is, for example, "2000 TFLOPS".
[0045] It should be understood that Figure 3 This is merely an example of a tree structure. It should be understood that a tree structure can have more or fewer levels. Each level can include more or fewer leaf nodes. The types of components indicated by leaf nodes are not limited to those specified in the example. Figure 3 The component type shown in the example.
[0046] Regarding component identifiers, they are used to uniquely identify the corresponding component or sub-component. For example, the component identifier "tensor_cores.fp16" is used to identify... Figure 3 Tensor kernel of medium half-precision floating-point (fp16) type.
[0047] As to the throughput value, it is used to indicate the processing capability of the corresponding component or sub-component. The throughput value can be a direct numerical value, a formula or a range. Table 1 below schematically illustrates the code for specifying the throughput value for the tensor core of the half-precision floating point (fp16) type (“tensor_cores.fp16”).
[0048] Table 1
[0049]
[0050] It should be understood that the hardware model supports parameterized expressions.
[0051] For example, as to the method for configuring the throughput value for the component corresponding to the leaf node at the bottom of the component path, it for example comprises: the computing device 100 causes the throughput value configured for the component corresponding to the leaf node at the bottom of the component path to be dynamically calculated via an expression. By adopting the above means, the present application can make the computing system of various different types of hardware characterized by changing parameters based on one basic hardware model.
[0052] For example, Table 2 below schematically illustrates the code for dynamically calculating the throughput values of the entire component “l2” and the entire component “tensor_core” via an expression.
[0053] Table 2
[0054]
[0055] As to the aggregation method, it is used to define the calculation mode of the execution time of the combined component. In some embodiments, the aggregation method comprises two aggregation methods: the maximum value aggregation mode (or “MAX aggregation”) and the cumulative aggregation mode (or “SUM aggregation”).
[0056] As to the maximum value aggregation mode, it is used to perform aggregation calculation on the execution time of the component of the exclusive hardware resource. The maximum value aggregation mode is configured to take the maximum value of the execution time of the component to be aggregated as the aggregated execution time. The component of the exclusive hardware resource is for example but not limited to: the memory, such as the read operation and the write operation, which can be performed simultaneously. The following formula (1) schematically illustrates the calculation method of the maximum value aggregation mode.
[0057] execution_time = max(read_time, write_time) (1)
[0058] In the above formula (1), read_time represents the execution time of the memory for the read operation. write_time represents the execution time of the memory for the write operation. execution_time represents the aggregated execution time.
[0059] Regarding the accumulation aggregation mode, it is suitable for aggregating calculation of the execution time of components sharing hardware resources. The accumulation aggregation mode is configured to accumulate the execution time of the components to be executed for the aggregation calculation, so as to take the accumulation result as the aggregated execution time. Among them, the components sharing hardware resources are, for example but not limited to: calculation units of different precisions. Because, the components sharing hardware resources usually occupy the same resource, for example, the calculation units of different precisions occupy the same bandwidth, therefore, it is necessary to calculate the aggregated execution time in the accumulation aggregation mode. The following formula (2) exemplarily illustrates the calculation method of the accumulation aggregation mode.
[0060] execution_time = sum(fp16_time, fp32_time, fp64_time) (2)
[0061] In the above formula (2), fp16_time represents the execution time of the calculation unit with the precision of fp16. fp32_time represents the execution time of the calculation unit with the precision of fp32. fp64_time represents the execution time of the calculation unit with the precision of fp64. execution_time represents the aggregated execution time.
[0062] Regarding the sub-component set, it indicates the set of sub-components included by the component.
[0063] At step 204, the computing device 100 calculates the workload of the application.
[0064] Regarding the method of calculating the workload of the application, it for example includes: the computing device 100 generates a workload model based on a plurality of workload elements, each workload element indicating the operation quantity of a predetermined type of operation or the transmission data quantity. The following table 3 exemplarily illustrates the code for generating the workload model taking matrix multiplication as an example.
[0065] Table 3
[0066]
[0067] As shown in Table 3, the name of the workload model is "matrix_multiply_fp16". The workload model is used to describe a workload of a 16*16 matrix multiplication of a half-precision floating point (fp16) type (as indicated by "16x16 FP16 matrix multiplication" in Table 3). The workload model includes a plurality of workload elements. Each workload element indicates an operation number of a predetermined type of operation or a data amount of transferring data related to the matrix multiplication of the half-precision floating point type of the 16*16 matrix dimension. For example, a first workload element (as indicated by "tensor_core_fp16_ops" in Table 3) indicates an operation number of a tensor core of the half-precision floating point (fp16) type. The meaning indicated by "value": 5.0E+15" in Table 3 is that a workload value configured for the first workload element is "5.0E+15". That is, the workload value corresponding to the operation number of the operation related to the matrix multiplication of the half-precision floating point type of the 16*16 matrix dimension is configured as "5.0E+15". A second workload element (as indicated by "hbm_read_bytes" in Table 3) indicates a read byte of a high bandwidth memory (HBM). The meaning indicated by "value": 1.0E+12" in Table 3 is that a workload value configured for the second workload element is "1.0E+12". That is, the workload value corresponding to the read byte of the high bandwidth memory is configured as "1.0E+12". For another example, a third workload element (as indicated by "hbm_write_bytes" in Table 3) indicates a write byte of the high bandwidth memory. The meaning indicated by "value": 5.0E+11" in Table 3 is that a workload value configured for the third workload element is "5.0E+11". That is, the workload value corresponding to the write byte of the high bandwidth memory is configured as "5.0E+11". It should be understood that the code of Table 3 is exemplary. It is not a limitation on the method of calculating the workload of the application program.
[0068] It should be understood that, with respect to the workload model, it is a multi-source workload model. In some embodiments, calculating the workload of the application program includes any one of: calculating a workload via an equation; calculating a workload via a compiler intermediate representation; and dynamically calculating a workload via a runtime performance metric.
[0069] With respect to the method of calculating a workload via an equation, it includes, for example: calculating an operation number of a computing operation and a data amount of a memory data transfer required by the application program to be used as a workload. Thereby, the present application can make a quick estimation for the workload by using an equation even in an early design stage.
[0070] The following illustrates the method for calculating the workload by way of an example of matrix multiplication. The matrix multiplication is represented as C = A * B, for example. Wherein, A represents a first input matrix with a size of M x K. B represents a second input matrix with a size of K x N. Wherein, M, K, N are all natural numbers.
[0071] The following Table 4 illustrates the code of the equation for estimating the workload of the matrix multiplication.
[0072] Table 4
[0073]
[0074] As shown in Table 4, M, N, K represent the dimension parameters of the matrix multiplication, respectively. The "element_size" represents the element size, for example, the element size is "2" for the element of the half-precision floating point type. Since the matrix multiplication is performed, each element needs K times of multiplication and K-1 times of addition. Therefore, the number of calculations (or "operation number") of the matrix multiplication calculated via the instruction "compute_ops = M * N * K * 2" is about 215 million times. For the matrix multiplication, the first input matrix A and the second input matrix B need to be read, and the output matrix C needs to be written. Therefore, the data amount of the memory transfer calculated via the instruction "compute_ops = M * N * K * 2" is about 6.29 MB.
[0075] In some embodiments, the method for calculating the workload via the compiler intermediate representation, for example, includes: analyzing the loop structure and the instructions, and accumulating from the innermost loop to the entire program, so as to obtain the workload, which at least includes: the operation number of the calculation operation and the data amount of the memory transfer data, in some embodiments, the data amount of the memory transfer data includes: the global memory read data amount, the global memory write data amount, the shared memory read data amount, and the shared memory write data amount. Thus, the present application can provide a more accurate workload.
[0076] In some embodiments, the method for calculating the workload via the compiler intermediate representation, for example, includes: for the case that the application program includes the loop structure, the computing device 100 calculates the traffic required for executing each component of the application program in turn for each layer of loop from the innermost loop of the loop structure to the outer loop.
[0077] The following illustrates the method of calculating the workload via the compiler intermediate representation by way of an example of matrix multiplication. For example, the computing device 100 calculates, for each inner loop, the shared memory read data amount required for loading the sub-block processed by the inner loop from the shared memory to the register, and the operation amount of the computation operation required for performing the sub-block multiplication on the register; calculates, based on the inner loop factor, the outer loop factor, and the operation amount of the computation operation and the shared memory read data amount of each multiply-add instruction execution, the operation amount and the shared memory read data amount required for the matrix multiplication; calculates, for each outer loop, the global memory read data amount required for loading the sub-block processed by the outer loop from the global memory to the shared memory, and the global memory write data amount required for writing the result back to the global memory from the register; and calculates, based on the outer loop factor, the global memory read data amount and the global memory write data amount for each outer loop, the global memory read data amount and the global memory write data amount required for the matrix multiplication. The method of calculating the workload via the compiler intermediate representation will be described in detail below, and will not be described herein again.
[0078] In some embodiments, the method of dynamically calculating the workload via the runtime performance indicator, for example, includes: obtaining, for a computing system running a kernel function, a performance indicator of a runtime, and calculating a workload of a workload model. By using the above means, the present application can accurately capture the computation and memory access amount in the actual execution process of the computing device. The following Table 5 illustrates the code of the method of dynamically calculating the workload via the runtime performance indicator.
[0079] Table 5
[0080]
[0081] In Table 5 above, for example, the operation number of the tensor core operation of the predetermined data type is calculated by the instruction “tensor_core_ops_f16= (tensor_core_bf16 + tensor_core_fp16) * (XXX) * 2”. The instruction “hbm_read_bytes = sum_all_hbm_bank(hbm_bank_read_bytes)” represents the read data amount of the global memory of the computing unit calculated by the accumulation manner. The instruction “hbm_write_bytes = sum_all_hbm_bank(hbm_bank_write_bytes)” represents the write data amount of the global memory of the computing system calculated by the accumulation manner. The instruction “shared_memory_read_bytes = XXX * (shared_memory_simt_read_bytes + shared_memory_tensor_core_read_bytes)” represents the read data amount of the shared memory of the computing system for calculation. The instruction “shared_memory_write_bytes = XXX * (shared_memory_simt_write_bytes + shared_memory_tensor_core_write_bytes)” represents the write data amount of the shared memory of the computing system for calculation. In the above instructions, “XXX” represents the hardware width; the identifier on the right side of the equal sign represents the original signal of the hardware component. It should be understood that the traffic of the components of the computing system is not limited to the operation number, read or write data amount of the components.
[0082] At step 206, the computing device 100 determines the mapping relationship of the components included in the computing system to the workloads.
[0083] For example, the computing device 100 establishes the correspondence between the components included in the computing system and the workload elements via the mapping model. By using the above means, the present application can quickly evaluate the workload situation of different designs by modifying the mapping relationship without changing the workload model. Table 6 below schematically illustrates the code for establishing the mapping relationship of the components to the workloads regarding the tensor operation.
[0084] Table 6
[0085]
[0086] In the above Table 6, "system.compute.tensor_cores.fp16" represents the component of the half-precision floating-point type of tensor core in the computing unit of the computing system. "tensor_core_fp16_ops" represents the number of operations of the computing operation of the half-precision floating-point type of tensor core. "system.memory.hbm3.read" represents the component of the high-bandwidth memory for reading in the memory of the computing system. "hbm_read_bytes" represents the amount of data read by the high-bandwidth memory. "system.memory.hbm3.write" represents the component of the high-bandwidth memory for writing in the memory of the computing system. "hbm_write_bytes" represents the amount of data written by the high-bandwidth memory.
[0087] At step 208, the computing device 100 calculates the execution time of the component via aggregated computation based on the mapping relationship and the tree structure, so as to determine the performance bottleneck of the computing system running the application.
[0088] As to the method of calculating the execution time of the component via aggregated computation, it for example comprises: calculating the execution time of the component corresponding to the leaf node at the bottom of the component path based on the throughput value of the component corresponding to the leaf node at the bottom of the component path and the workload load value. For example, the corresponding workload load value is divided by the throughput value to obtain the execution time of the component corresponding to the leaf node at the bottom of the component path.
[0089] The following Table 7 schematically illustrates the pseudo code of calculating the execution time of the component corresponding to the leaf node.
[0090] Table 7
[0091]
[0092] For example, the execution time of each component calculated respectively comprises: the execution time of the half-precision floating-point type of tensor core = 5.0E+15 / 1.0E+15 = 5.0 seconds. The execution time of HBM reading = 1.0E+12 / 3.0E+12 = 0.33 seconds. The execution time of HBM writing = 5.0E+11 / 3.0E+12 = 0.17 seconds. The execution time of the cross-chip link = 12.0E+11 / 9.0E+11 = 0.22 seconds.
[0093] Regarding the method of calculating the execution time of a component via aggregated computation, in some embodiments, it comprises, for example: calculating the execution time of all components on a component path from the bottom layer upwards along the component path. Specifically, in response to the computing device 100 determining that the aggregation method of the component corresponding to the leaf node of the upper level of the tree structure is the maximum value aggregation mode (e.g., indicated by the instruction "if parent.aggregation_method == "max"" in Table 8), the maximum value of the execution time of the component in the current level is selected as the execution time of the component corresponding to the leaf node of the upper level; and in response to determining that the aggregation method of the component corresponding to the leaf node of the upper level of the tree structure is the cumulative aggregation mode (e.g., indicated by the instruction "else:" in Table 8), the execution times of all components corresponding to the leaf nodes of the current level are cumulatively added so as to take the cumulative result as the execution time of the component corresponding to the leaf node of the upper level.
[0094] The following Table 8 schematically illustrates the pseudo code of the method of calculating the execution time of all components on a component path from the bottom layer of the hierarchical relationship upwards along the component path.
[0095] Table 8
[0096]
[0097] Regarding the method of determining the performance bottleneck of the computing system on which the application is running, it comprises, for example: comparing the execution times of all components within the same level so as to determine the component corresponding to the maximum execution time as the performance bottleneck. The following Table 9 schematically illustrates the pseudo code of calculating the component of the performance bottleneck and the total execution time.
[0098] Table 9
[0099]
[0100] As shown in Table 9, the component corresponding to the maximum execution time is determined as the performance bottleneck by the instruction "bottleneck_path = max(all_times.items(), key=lambda x: x[1])[0]". The total execution time is calculated by the instruction "execution_time = all_times[bottleneck_path]". Then, the execution time of the current component is divided by the calculated total execution time to obtain the utilization rate of the current component.
[0101] For example, the execution time of HBM calculated by applying the aggregation rule = max(0.33, 0.17) = 0.33 seconds. The execution time of the computing unit = max(5.0, 0) = 5.0 seconds. The total execution time of the computing system = max(5.0, 0.33, 0.22) = 5.0 seconds.
[0102] In some embodiments, the method for determining the performance bottleneck of the computing system on which the application is running, for example, comprises: calculating the execution time of the component corresponding to the leaf node at the bottom layer of the tree structure, so as to calculate the execution time of the component successively upwards along the component path from the bottom layer of the tree structure; and determining the performance bottleneck of the computing system on which the application is running and the total execution time of the computing system based on the calculated execution time of the component.
[0103] For example, the following workload is calculated based on the performance on the GPU. For example, the workload comprises: the operation number of the tensor core operation of the half-precision floating point type = 5.0E+15. The HBM read data volume = 1.0E+12 bytes. The HBM write data volume = 5.0E+11 bytes. The cross-chip link data volume = 2.0E+11 bytes.
[0104] In some embodiments, the calculated execution time of the component corresponding to the leaf node at the bottom layer is, for example: the execution time of the tensor core of the half-precision floating point type = 5.0E+15 / 1.0E+15 = 5.0 seconds. The execution time of the HBM read = 1.0E+12 / 3.0E+12 = 0.33 seconds. The execution time of the HBM write = 5.0E+11 / 3.0E+12 = 0.17 seconds. The execution time of the cross-chip link = 2.0E+11 / 9.0E+11 = 0.22 seconds.
[0105] For example, based on the maximum aggregation mode, the maximum value of the execution time of the HBM memory read of the component corresponding to the leaf node at the bottom layer, i.e., 0.33 seconds, and the execution time of the HBM write, i.e., 0.17 seconds, is taken as the execution time of the HBM of the component corresponding to the leaf node at the upper level, i.e., the execution time of the HBM = max(0.33, 0.17) = 0.33 seconds.
[0106] Similarly, based on the maximum aggregation mode, the execution time of the computing unit corresponding to the leaf node at the upper level is calculated = max(5.0, 0) = 5.0 seconds. Wherein, “5.0” is, for example, the execution time of the tensor core of the half-precision floating point type; and “0” is, for example, the execution time of the SIMT core, for example, the SIMT core in the current system is not configured in the general tree structure, so the execution time of the SIMT core is “0”).
[0107] Based on the maximum aggregation mode, the execution time of the component memory corresponding to the last level leaf node is calculated = max(0.22, 0, 0) = 0.22 seconds.
[0108] Regarding the method for determining the total execution time of the computing system, for example, it includes: based on the maximum aggregation mode, the maximum value (i.e., 5.0 seconds) of the execution time 5.0 seconds of the computing unit corresponding to the aforementioned last level leaf node, the execution time 0.33 seconds of the HBM memory read, and the execution time 0.22 seconds of the memory, as the total execution time of the computing system corresponding to the root node of the uppermost level (i.e., the total execution time of the system = max(5.0, 0.33, 0.22) = 5.0 seconds).
[0109] Regarding the method for determining the performance bottleneck, in some embodiments, for example, it includes: comparing the execution times of the components in the current hierarchy, so as to take the component path where the component with the longest execution time is located as the performance bottleneck.
[0110] For example, comparing the execution times of the components corresponding to the leaf nodes in the current hierarchy, i.e., comparing the execution time 0.33 seconds of the HBM, the execution time 0.22 seconds of the memory, and the execution time 5.0 seconds of the computing unit, taking the component path "system.compute.tensor_cores.fp16" where the computing unit with the longest execution time is located as the performance bottleneck. The bottleneck time is 5.0 seconds.
[0111] In some embodiments, the method 200 further includes: calculating the component utilization of each component; and determining a system optimization strategy based on the calculated component utilization. The following formula (3) illustrates the calculation method of the component utilization.
[0112] Component utilization = execution time of component / total execution time of system (3)
[0113] For example, the component utilization of the tensor core of the half-precision floating-point type (tensor_cores.fp16) is 100%; the component utilization of the third-generation high-bandwidth memory for reading (hbm3.read) is 6.6%; the component utilization of the third-generation high-bandwidth memory for writing (hbm3.write) is 3.4%; and the component utilization of the fourth-generation cross-chip interconnection (cross-chip-link4) is 4.4%.
[0114] Based on the calculated component utilization, a system optimization strategy is determined, for example, for the tensor core of the half-precision floating-point type with the highest component utilization, the number of tensor cores is increased or the algorithm is optimized to reduce the computing demand.
[0115] In the above scheme, the hierarchical relationship of the computing system including multiple components running the application program is characterized by using a tree structure for aggregated computation; the present application can support hardware components of various granularities, and modeling of complex conditions such as multi-processor, multi-level memory, etc. In addition, the workload of the application program is calculated; the mapping relationship of the components included in the computing system to the workload is determined; and based on the mapping relationship and the tree structure, the execution time of the components is calculated via aggregated computation, so as to determine the performance bottleneck of the computing system running the application program, the present application can accurately predict the performance of the throughput-limited computing system without actual execution. Thus, the present application can make accurate performance prediction even for complex applications that lack real hardware running test support.
[0116] Further, the present application can quickly predict the performance for different designs by simply changing the mapping relationship of the components to the workload. In addition, the present application can automatically identify the performance bottleneck of the computing system and quantify the optimization space.
[0117] In some embodiments, for the case where the application program includes a loop structure, the method for calculating the workload of the application program, for example, includes that the computing device 100 calculates the traffic of each component required for executing the application program in turn for each loop from the innermost loop to the outer loop of the loop structure. It should be understood that the traffic, for example, but not limited to, the memory access amount or the computation amount of a certain component. In some embodiments, each component required for executing the application program, for example, includes different types of storage units, computing units and communication units.
[0118] The following takes matrix multiplication as an example to illustrate the method for calculating the workload of the application program.
[0119] Specifically, the method for calculating the workload of the application program, for example, includes the following steps. First, the computing device 100 calculates, for each inner loop, the shared memory read data amount required for loading the sub-block processed by the inner loop from the shared memory to the register, and the operation amount of the computation operation required for executing the sub-block multiplication on the register.
[0120] For the inner loop shown in Table 10 below (for example, ti indicates the row loop of the inner loop, tj indicates the column loop of the inner loop, and tk indicates the inner product dimension loop of the inner loop), the variation range of ti, tj and tk is: 0 to 64, with a step of 8. The loop times of each dimension (or "inner loop factor") is "8". The inner loop iterates 8*8*8 = 512 times in total.
[0121] In each inner loop iteration, two 8x8 sub-blocks (e.g., AS[ti:ti+8, tk:tk+8] and BS[tk:tk+8, tj:tj+8]) are loaded from shared memory to registers (AR and BR), respectively. Each sub-block, for example, includes 8x8 = 64 elements, each of which is 2 bytes in size. Thus, each sub-block is 128B in size. For example, as shown in Table 10, each of the 8x8 sub-blocks (AS[ti:ti+8, tk:tk+8]) is loaded from shared memory to register (AR) via the instruction "LdMatrix(AR, AS[ti:ti+8, tk:tk+8])", which is 8x8x2B = 128B in data volume; and each of the 8x8 sub-blocks (BS[tk:tk+8, tj:tj+8]) is loaded from shared memory to register (BR) via the instruction "LdMatrix(BR, BS[tk:tk+8, tj:tj+8])", which is 8x8x2B = 128B in data volume.
[0122] In the inner loop iteration (ti, tj, tk), as shown in Table 10 by the instruction "Mma(CR, AR, BR)", each Mma instruction performs a multiply-add computation operation of the half-precision floating-point type (or simply "FP16 operation") in an operation quantity of 1024 (i.e., 8x8x8x2 = 1024).
[0123] Second, the computing device 100 calculates the operation quantity and the shared memory read data volume required for the matrix multiplication based on the inner loop factor, the outer loop factor, and the operation quantity of the computation operation performed by each multiply-add instruction and the data volume of the shared memory read.
[0124] For example, the operation quantity of the multiply-add computation operation of the half-precision floating-point type performed by each multiply-add (Mma) instruction is calculated to be 1024. The inner loop factor is 8x8x8 = 512. The outer loop factor is 16x16x16 = 4096. Based on the inner loop factor, the outer loop factor, and the operation quantity of each Mma instruction, the operation quantity of the multiply-add computation operation of the half-precision floating-point type is calculated to be 1024x512x4096 = 2.15x10^9.
[0125] As shown in Table 10, the single instruction "LdMatrix(BR, BS[tk:tk+8, tj:tj+8])" or the instruction "LdMatrix(AR, AS[ti:ti+8, tk:tk+8])" calculates the shared memory read data amount required for each multiply-accumulate (Mma) to be 8x8x2B = 128B at step 402. The outer loop factor is 16x16x16 = 4096. Based on the inner loop factor, the outer loop factor, and the shared memory read data amount required for the single instruction "LdMatrix", the shared memory read data amount required for the matrix multiplication is calculated. For example, the shared memory read data amount required for the matrix multiplication is 128Bx2x512x4096 = 5.37x10^8 B.
[0126] Again, the computing device 100 calculates, for each outer loop, the global memory read data amount required to load the sub-blocks processed by the outer loop from the global memory to the shared memory, and the global memory write data amount required to write the results from the registers back to the global memory.
[0127] As shown in Table 10, the outer loop includes three nested loops (e.g., i indicates the row loop of the outer loop, j indicates the column loop of the outer loop, and k indicates the inner product dimension loop of the outer loop), each loop has a range of 0 to 1024 with a step of 64, thus, the number of loops (or "outer loop factor") for each dimension is "16". The total iterations of the outer loop is, for example, 16*16*16 = 4096.
[0128] For each outer loop, two 64x64 blocks are loaded from the global memory, for example, in each inner k loop. One is the block of the first input matrix A (e.g., the block is located at A[i:i+64, k:k+64]) and the other is the block of the second input matrix B (e.g., the block is located at B[k:k+64, j:j+64]) to the shared memory (e.g., the shared memory is AS and BS, respectively). Each of which is loaded with 64*64 = 4096 elements, each element is 2 bytes, thus, each block is 8KB. As shown in Table 10, the data amount of each global memory load 64x64 block to the shared memory AS via "LdMatrix(AR, AS[ti:ti+8, tk:tk+8])" is 64x64x2B = 8KB; and the data amount of each global memory load 64x64 block to the shared memory BS via "LdMatrix(BS, B[k:k+64, j:j+64])" is 64x64x2B = 8KB.
[0129] Further, the computing device 100 calculates the global memory read data amount and the global memory write data amount required for the matrix multiplication based on the outer loop factor, the global memory read data amount and the global memory write data amount for each outer loop.
[0130] For example, as shown in Table 10, for each outer loop matrix block, the global memory read data amount is (64x64x2B)x2 = 16KB (first input matrix A matrix block and second input matrix B matrix block); the global memory write data amount is 64x64x2B = 8KB (output matrix C matrix block). The outer loop factor is 16x16x16 = 4096. Then the global memory read data amount required for the matrix multiplication is 16KBx4096 = 6.71x10^7 B. The global memory write data amount required for the matrix multiplication is 8KBx4096 = 3.35x10^7 B.
[0131] By employing the above-mentioned hierarchical workload analysis method, the workload of the entire program, such as the operation number of the calculation operation and the memory access amount at each level, can be accurately calculated from the inner loop instruction by multiplying each level loop factor.
[0132] The following Table 10 schematically illustrates the code of the equation for estimating the workload of the matrix multiplication. The code shown in Table 10 utilizes the multi-layer nested loop to perform the matrix multiplication operation.
[0133] Table 10
[0134]
[0135] The following will be described in conjunction with Figure 4 A method 400 for predicting the performance of a computing system on which an application program is run according to another embodiment of the present application will be described. Figure 4 A flowchart of the method 400 for predicting the performance of a computing system on which an application program is run according to another embodiment of the present application is shown. It should be understood that the method 400 can be performed, for example, at the computing device 100 described above. Figure 1 The method 400 can also include additional actions not shown and / or can omit actions shown, without limitation in this regard.
[0136] At step 402, the computing device 100 uses a tree structure to respectively represent the hierarchical relationship of a first computing system and a second computing system performing an application program for aggregated calculation, the first computing system and the second computing system respectively including a plurality of components.
[0137] As to the executed application program, it is, for example, but not limited to, a matrix multiplication.
[0138] As to the first computing system, it is, for example, a high-end GPU of some type (referred to as "GPU_A" for short).
[0139] As to the second computing system, it is, for example, a mid-end GPU of some type (referred to as "GPU_B" for short).
[0140] Table 11 below schematically illustrates code for characterizing the hierarchical relationship of the first computing system using a tree structure. Table 12 below schematically illustrates code for characterizing the hierarchical relationship of the second computing system using a tree structure. The tree structure is explained in detail below in connection with Table 11. In the tree structure of the first computing system, the system at the root node, the next level of leaf nodes under the root node correspond to the components of: compute and memory, respectively. As to the component of compute, the associated aggregation manner is the maximum value aggregation manner, the next level of leaf nodes under the component path correspond to the component of tensor cores, and the next level of leaf nodes under the component path correspond to the component of fp16. Table 11 configures the throughput value of fp16 through the code ("throughput":{"value": 3.12E+14}). As to the component of memory, the associated aggregation manner is the maximum value aggregation manner, the next level of leaf nodes under the component path correspond to the component of global memory and shared memory. The next level of leaf nodes under the component path correspond to the component of read and write for global memory, and the next level of leaf nodes under the component path correspond to the component of read and write for shared memory. Table 11 configures the throughput value of read and write, respectively, through the code.
[0141] Table 11
[0142]
[0143] Table 12
[0144]
[0145] At step 404, the computing device 100 calculates the workload of the application.
[0146] For example, the computing device 100 determines the workloads of the application via static analysis of the compiler intermediate representation. Table 13 below schematically illustrates the code for determining the workloads. As shown in Table 13, the workload elements, i.e., the number of tensor core operations of half-precision floating point type (tensor_core_fp16_ops), the amount of data of global memory read (global_memory_read_bytes), the amount of data of global memory write (global_memory_write_bytes), the amount of data of shared memory read (shared_memory_read_bytes), and the amount of data of shared memory write (shared_memory_write_bytes) are respectively configured with corresponding workload values.
[0147] Table 13
[0148]
[0149] At step 406, the computing device 100 respectively determines the mapping relationship of the components included in the first computing system and the second computing system to the workloads.
[0150] For example, the computing device 100 respectively maps the components included in the first computing system and the second computing system to the workloads via the mapping model. Table 14 below schematically illustrates the code of the mapping model.
[0151] Table 14
[0152]
[0153] It should be understood that the manner of determining the workloads and the manner of determining the mapping relationship of the components to the workloads are different for different applications, and the method for determining the workloads and the mapping relationship of the components to the workloads for an application of scientific computing containing a large number of special function calls will be described in detail below. Here, no further description is given.
[0154] At step 408, the computing device 100 respectively determines the performance bottlenecks and the total execution times of the first computing system and the second computing system in order to compare the performance of the first computing system and the second computing system, the performance bottlenecks and the total execution times of the first computing system and the second computing system being obtained via aggregation calculation based on the mapping relationship and the tree structure respectively.
[0155] For example, for the first computing system, the execution time of the tensor cores of the half-precision floating point type (tensor_cores.fp16) is 6.89E-6 seconds, the execution time of the global memory read (global_memory.read) is 2.79E-6 seconds, the execution time of the global memory write (global_memory.write) is 1.40E-6 seconds, the execution time of the shared memory read (shared_memory.read) is 3.53E-6 seconds, and the execution time of the shared memory write (shared_memory.write) is 3.53E-6 seconds. Via the aggregated computation of the execution times of the global memory and the shared memory along the component path from the bottom layer upwards in the tree structure, the computed execution time of the global memory is 2.79E-6 seconds (as indicated by, for example, “global_memory: max(2.79E-6, 1.40E-6) = 2.79E-6 seconds”). The execution time of the shared memory is 3.53E-6 seconds (as indicated by, for example, “shared_memory: max(3.53E-6, 3.53E-6) = 3.53E-6 seconds”).
[0156] Further based on the maximum value aggregation mode for the aggregated computation, the execution time of the memory is 3.53E-6 seconds (as indicated by, for example, “memory: max(2.79E-6, 3.53E-6) = 3.53E-6 seconds”). The execution time of the computing unit is 6.89E-6 seconds.
[0157] Further based on the maximum value aggregation mode for the aggregated computation, the total execution time of the first computing system of the root node of the final computation is 6.89E-6 seconds (as indicated by, for example, “system: max(6.89E-6, 3.53E-6) = 6.89E-6 seconds”). That is, the total execution time of the first computing system is 6.89 microseconds.
[0158] In addition, via the comparison of the execution times of the components, the component path in which the component with the longest execution time, the tensor cores of the half-precision floating point type (tensor_cores.fp16), is located is identified as the performance bottleneck of the first computing system.
[0159] For example, for the second computing system, the execution time of the tensor core of the half-precision floating point type (tensor_cores.fp16) is 1.38E-5 seconds, the execution time of the global memory read (global_memory.read) is 4.66E-6 seconds, the execution time of the global memory write (global_memory.write) is 2.33E-6 seconds, the execution time of the shared memory read (shared_memory.read) is 6.10E-6 seconds, and the execution time of the shared memory write (shared_memory.write) is 6.10E-6 seconds.
[0160] Similarly to the manner of calculating the first computing system, the aggregated calculation is further performed based on the maximum value aggregation mode, and the following results are obtained, respectively: the execution time of the global memory is 4.66E-6 seconds (as indicated by "global_memory: max(4.66E-6, 2.33E-6) = 4.66E-6 seconds", for example). The execution time of the shared memory is 6.10E-6 seconds (as indicated by "shared_memory: max(6.10E-6, 6.10E-6) = 6.10E-6 seconds", for example). The execution time of the memory is 6.10E-6 seconds (as indicated by "memory: max(4.66E-6, 6.10E-6) = 6.10E-6 seconds", for example). The execution time of the computing unit is 1.38E-5 seconds. The total execution time of the second computing system is 1.38E-5 seconds (as indicated by "system: max(1.38E-5, 6.10E-6) = 1.38E-5 seconds", for example). That is, the total execution time of the second computing system is 13.8 microseconds.
[0161] In addition, via comparison of the execution times of the components, the component path in which the component with the longest execution time, the tensor core of the half-precision floating point type (tensor_cores.fp16), is located is taken as the bottleneck of the performance of the second computing system.
[0162] The performance bottlenecks on the first computing system (GPU A) and the second computing system (GPU B) are both the tensor core of the half-precision floating point type and its computing capability, and the memory system is not a bottleneck under the workload of matrix multiplication. In addition, based on comparison of the total execution times of the systems, it can be known that the execution speed of the first computing system (GPU A) is about 2 times faster than that of the second computing system (GPU B), and the throughput proportion of the tensor core of the first computing system (GPU A) is consistent with that of the second computing system (GPU B).
[0163] In the above scheme, the application can evaluate the performance bottleneck and total execution time of different computing systems for a predetermined application, and then conveniently and accurately select a suitable computing system matching the predetermined application.
[0164] In some embodiments, the method of determining the workloads and determining the mapping relationship of the components to the workloads, for example, comprises: if the computing device 100 determines that the types of the components configured in the computing system are different, the mapping manner of the same workloads to the different types of components is different. Specifically, if the types of the computing components or the memory components configured in the computing system are different, the mapping manner of the same workloads to the different types of computing components or memory components is different. For example, for the two different types of computing components, vector cores and SIMT cores, the mapping manner of the workloads to the computing components is different.
[0165] The following illustrates the different mapping manners of the workloads to the different types of components by taking whether a special unit component is configured in the computing system as an example. For example, the computing device 100 determines different mapping relationships of the components to the workloads according to whether a special unit component is configured in the computing system, the special unit being, for example, for a predetermined transcendental function.
[0166] For example, first, the computing device 100 determines the workloads of an application in which the number of times of calling a special function exceeds a predetermined number threshold.
[0167] If the computing device 100 determines that a special unit component is configured in the computing system, the special unit component is respectively mapped with: a standard workload element for the special unit, the standard workload element including: the number of operations of single-precision floating-point operations, the number of operations of 32-bit integer type operations, the number of operations of sine operations, the number of operations of cosine operations, and the number of operations of exponential operations.
[0168] It should be understood that some hardware components are configured with special units (or “sfus”), and some hardware components are not configured with sfus. The following Table 15 illustratively shows the code of the mapping relationship of the components configured with special units.
[0169] Table 15
[0170]
[0171] If the computing device 100 determines that a special unit component is not configured in the computing system, the special function mapped to the special unit component is implemented by simulating a general-purpose operator.
[0172] The following Table 16 illustratively shows the code of the mapping relationship of the components not configured with special units.
[0173] Table 16
[0174]
[0175] As shown in Table 16, for the case where the hardware component sfu is not configured, the special functions originally mapped to sfu need to be implemented through general-purpose arithmetic logic units (ALUs), and each special function needs multiple standard operations (e.g., the number of operations of single-precision floating-point operations indicated by "fp32_ops", the number of operations of 32-bit integer type operations indicated by "int32_ops", the number of operations of SFU sine operations indicated by "sfu_sin_ops", the number of operations of SFU cosine operations indicated by "sfu_cos_ops", and the number of operations of SFU exponential operations indicated by "sfu_exp_ops"). To support such mapping between different hardware architectures, conversion coefficients (e.g., coefficients 15, 20, etc.) are included in the expressions. The conversion coefficients represent how many standard operations are equivalent to one special function operation.
[0176] Such a mapping mechanism enables the system to calculate the equivalent workloads required to implement the same functions on different hardware, thereby making accurate performance comparisons. In addition, the present application can also automatically apply appropriate conversion coefficients by changing the mapping of hardware components to workload elements without changing the definition of the original workload.
[0177] For example, the method of determining the workload and determining the mapping relationship between the components and the workload further includes: further comparing the performance of the computing system configured with sfu and the computing system not configured with sfu.
[0178] For example, for the computing system configured with sfu, the execution time of the SFU sine operation is 0.5 seconds (as indicated by "sfu sin time: 5.0E+9 / 1.0E+10 = 0.5 seconds"). The execution time of the SFU cosine operation is 0.3 seconds (as indicated by "sfu cos time: 3.0E+9 / 1.0E+10 = 0.3 seconds"). The execution time of the SFU exponential operation is 0.25 seconds (as indicated by "sfu exp time: 2.0E+9 / 8.0E+9 = 0.25 seconds"). The execution time of the SIMT of single-precision floating-point type is 1.0 seconds (as indicated by "SIMT fp32 time: 1.0E+11 / 1.0E+11 = 1.0 seconds"). The execution time of the SIMT of 32-bit integer type is 0.42 seconds (as indicated by "SIMT int32 time: 5.0E+10 / 1.2E+11 = 0.42 seconds").
[0179] The execution time of the SIMT compute unit is 1.0 seconds via the aggregate computation (as indicated by "SIMT total time = max(1.0, 0.42) = 1.0 seconds"). The total execution time of the compute system is 1.05 seconds (as indicated by "Compute total time = max(1.05, 1.0) = 1.05 seconds").
[0180] For example, for the compute system without the sfu configured, its equivalent fp32 operation = 1.0E+11 + (5.0E+9 x 15) + (3.0E+9 x 15) + (2.0E+9 x 20) = 2.6E+11. The execution time of the SIMT of the single-precision floating-point type is 2.17 seconds (e.g., calculated via "SIMT fp32 time (including emulated special functions): 2.6E+11 / 1.2E+11 = 2.17 seconds"). The execution time of the SIMT of the 32-bit integer type is 0.33 seconds (e.g., calculated via "SIMT int32 time: 5.0E+10 / 1.5E+11 = 0.33 seconds").
[0181] The total execution time of the compute system is 2.17 seconds (e.g., calculated via "Compute total time = max(2.17, 0.33) = 2.17 seconds").
[0182] The execution time using the compute system configured with the sfu is 1.05 seconds; the execution time using the compute system without the sfu configured is 2.17 seconds. The latter has a performance improvement of about 2.07 times. Thus, the special function unit and the sfu can provide significant performance improvement under such workloads, demonstrating the value of hardware customization for predetermined application domains.
[0183] The various processes and processes described above, such as the methods 200, 400, can be executed at a computing device. The computing device includes, for example, at least one processor (at least one graphics processor and at least one central processor) and a memory communicatively connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor. In some embodiments, the methods 200, 400 can be implemented as a computer software program or program product, which is tangibly contained in a machine-readable medium. In some embodiments, part or all of the computer program can be loaded and / or installed on the computing device via a Read-Only Memory (ROM) and / or a communication unit. When the computer program is loaded into a Random-access memory (RAM) and executed by the GPU and the CPU, one or more actions of the methods 200, 400 described above can be executed.
[0184] The present application can be a method, an apparatus, a system, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for performing various aspects of the present application. The computer readable storage medium can be a tangible device that can retain and store instructions for execution by a processor. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A computer readable storage medium can be any available medium or
[0185] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The computer readable program instructions can be embodied on one or more computer readable storage media. The computer readable program instructions described herein can be implemented in a variety of operating environments, which have various computers such as client computers, server computers, handheld computers, mobile computers, netbooks, set-top boxes, music players, personal data assistants, applications, calculators, computer system, computers based on general or special purpose computers, and so on. It is understood that any one or combination of the following can be utilized: a computer program product, a program, a program of instructions, an instruction set, an application, a software application, a software package, a file, a computer, a processor, a processor implementation, a processor based system, a computer based system, an article of manufacture, an
[0186] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can be coupled to a computer or other programmable data processing apparatus, which can be used to program a computer, to cause the computer to perform a process to implement the functions / acts specified in the flowchart and / or block diagram block or blocks. The computer readable storage medium can be, for example, but is not limited to, a floppy disk, a compact disc, a tape, a hard disk drive, an electronic memory, a solid state memory, or a hybrid memory (e.g., a memory including a combination of different types of memory such as magnetic, optical, and / or semiconductor storage).
[0187] The computer program product of the present application can be a computer program product which implements the methods of the present application on a data carrier, such as a diskette or hard disk. The implementation can also take place on a network, such as the Internet, using a server or the like.
[0188] It should be understood that the various forms of flow shown in the figures can be re-ordered, added to, or have steps deleted. For example, the steps recited in the present application can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present application are achieved. This is not limited herein.
[0189] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and scope of the disclosure.
Claims
1. A method for predicting the performance of a computing system on which an application runs, characterized in that, include: A tree structure is used to represent the hierarchical relationship of the computing system on which the application runs, for the purpose of aggregating computations, the computing system comprising multiple components; Calculate the workload of the application; Determine the mapping relationship between the components of the computing system and the workload; as well as Based on the mapping relationship and tree structure, the execution time of the computational components is calculated via aggregation to identify the performance bottleneck of the computing system on which the application runs. The workload of the computing application includes: generating a workload model based on multiple workload elements, each workload element indicating the number of operations or the amount of data transferred for a specific type of operation; Determining the mapping relationship between the components included in the computing system and the workload includes: in response to determining that the same workload is mapped to different types of components due to different types of components configured in the computing system.
2. The method according to claim 1, characterized in that, Using a tree structure to represent the hierarchical relationships of the computing system that the application runs includes: Configure a tree-structured hardware model, such that the root node of the tree represents the computing system on which the application runs, and multiple leaf nodes at one or more levels below the root node represent multiple components included in the computing system. Configure the component corresponding to the leaf node with a component identifier, as well as at least one of the following: throughput value, aggregation method, and subcomponent set.
3. The method according to claim 1, characterized in that, Configure the component corresponding to the leaf node with a component identifier, as well as at least one of the following: throughput value, aggregation method, and child component set: Configure throughput values for the components corresponding to the leaf nodes at the bottom of the component path. The component path is used to indicate the association path of the component and its child components in the tree structure. as well as Configure aggregation methods for components corresponding to leaf nodes that are not at the bottom level of the component path.
4. The method according to claim 3, characterized in that, The workload of a computing application includes any of the following: The workload is calculated using equations; The workload is computed via the compiler's intermediate representation; and The workload is dynamically calculated based on runtime performance metrics.
5. The method according to claim 2, characterized in that, The aggregation method includes: A maximum value aggregation mode, used to aggregate the execution time of components that exclusively occupy hardware resources; and The cumulative aggregation mode is used to perform aggregate calculations on the execution time of components that share hardware resources.
6. The method according to claim 5, characterized in that, Based on the mapping relationship and tree structure, the execution time of the computing component includes, via aggregation computation: In response to the aggregation method determining the component corresponding to the leaf node of the previous level in the tree structure as the maximum aggregation mode, the maximum execution time of the components in the current level is selected as the execution time of the component corresponding to the leaf node of the previous level; and In response to the aggregation method for determining the component corresponding to the leaf node of the previous level of the tree structure, which is the cumulative aggregation mode, the execution time of the components corresponding to all leaf nodes of the current level is accumulated so that the accumulated result can be used as the execution time of the component corresponding to the leaf node of the previous level.
7. The method according to claim 3, characterized in that, The application includes a loop structure, and the workload of the computational application includes: For each loop in the loop structure, from the innermost loop to the outermost loop, calculate the traffic required for each component to execute the application.
8. The method according to claim 4, characterized in that, Identifying performance bottlenecks in the computing system on which the application runs includes: Calculate the execution time of the components corresponding to the leaf nodes at the bottom of the tree structure, so that the execution time of each component can be calculated sequentially upwards along the component path from the bottom of the tree structure; and Based on the calculated execution time of the components, the performance bottleneck of the computing system on which the application runs and the total execution time of the computing system are determined.
9. The method according to claim 8, characterized in that, Identifying performance bottlenecks in the computing system on which the application runs includes: Compare the execution times of each component in the current layer to identify the component path containing the component with the longest execution time as the performance bottleneck.
10. A method for predicting the performance of a computing system on which an application runs, characterized in that, include: A tree structure is used to represent the hierarchical relationship between the first computing system and the second computing system that execute the application, respectively, for aggregate computing. The first computing system and the second computing system each include multiple components. Calculate the workload of the application; Determine the mapping relationship between the components included in the first computing system and the second computing system and the workload, respectively; as well as The performance bottlenecks and total execution times of the first and second computing systems are determined respectively in order to compare their performance. The performance bottlenecks and total execution times of the first and second computing systems are obtained by aggregation computation based on the mapping relationship and tree structure respectively. The workload of the computing application includes: generating a workload model based on multiple workload elements, each workload element indicating the number of operations or the amount of data transferred for a specific type of operation; Determining the mapping relationship between the components included in the computing system and the workload includes: in response to determining that the same workload is mapped to different types of components due to different types of components configured in the computing system.
11. A computing device, characterized in that, include: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 1-10.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a machine, performs the method according to any one of claims 1-10.
13. A computer program product, characterized in that, Includes a computer program, which, when executed by a machine, performs the method according to any one of claims 1-10.
Citation Information
Patent Citations
Efficiency bottleneck analysis method
CN114490295A
Method and system for evaluating computer hardware performance based on analogue simulation model
CN120353684A