Performance prediction method and device for chip execution operator, electronic equipment and medium
By obtaining the execution time of the reference chip and the utilization rate of the core components, and combining the hardware configuration specifications, the execution performance of the target chip is predicted, and the problem of accurately predicting execution performance on a general graphics processor in different configurations is solved, improving the accuracy of prediction and application support.
Patent Information
- Application Number
- CN202510648305.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-20
AI Technical Summary
In the absence of real execution or simulation, it is a technical challenge to accurately predict the execution performance of AI operators on a new generation or general-purpose graphics processor.
By obtaining the actual execution time of the reference chip to execute the current operator and the actual utilization rate of each core component, combined with the hardware configuration specifications, the actual workload of the core component is determined. Then, based on the hardware configuration specifications of the target chip, the predicted occupation time and the predicted execution time of each core component of the target chip that executes the same operator are predicted.
It realizes the execution performance of the accurate prediction operator on unknown chips without real execution or simulation, improves the accuracy of performance prediction, and provides strong support for resource management, hardware design, software optimization and system reliability.
Smart Images

Figure CN120162236A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence chip technology, and in particular to a method, device, electronic device and medium for predicting the performance of a chip execution operator. Background Art
[0002] With the rapid development of artificial intelligence (AI) and high-performance computing (HPC) applications, the architecture of the new generation of general-purpose graphics processing units (GPGPU) is constantly evolving, with significant differences in computing power, bandwidth, and microarchitecture design for each generation. Accurately predicting the execution performance of AI operators on the next generation or other configurations of GPGPU without actual execution or simulation has always been a technical challenge in the industry.
[0003] Therefore, how to accurately predict the execution performance of chip execution operators has become a technical problem that needs to be solved urgently in the industry. Summary of the invention
[0004] The present invention provides a method, device, electronic device and medium for predicting the performance of a chip execution operator, which are used to solve the technical problem of how to accurately predict the execution performance of a chip execution operator.
[0005] The present invention provides a method for predicting the performance of a chip execution operator, comprising: Obtaining the actual execution time of the reference chip executing the current operator and the actual utilization rate of each core component in the reference chip; Determine the actual workload of each core component in the reference chip based on the actual utilization rate of each core component in the reference chip, the actual execution time, and the hardware configuration specifications of each core component in the reference chip; Determining the predicted occupancy time of each core component when the target chip executes the current operator based on the actual workload of each core component in the reference chip and the hardware configuration specifications of each core component in the target chip; the target chip and the reference chip have the same type of core components; The predicted execution time of the target chip executing the current operator is determined based on the predicted occupancy time of each core component when the target chip executes the current operator and the predicted utilization rate of the target chip.
[0006] In some embodiments, the core components include a tensor core, a vector core, and a global memory; the global memory is used for access by the tensor core and the vector core.
[0007] In some embodiments, determining the actual workloads of the respective core components in the reference chip based on the actual utilization rates of the respective core components in the reference chip, the actual execution duration, and the hardware configuration specifications of the respective core components in the reference chip includes: Obtaining the actual clock frequency at which the reference chip executes the current operator, and the peak frequency of the reference chip; Based on the actual utilization rate of the tensor core, the actual clock frequency, the peak frequency, the actual execution duration, and the hardware configuration specifications of the tensor core, determining the actual workload of the tensor core.
[0008] In some embodiments, determining the actual workloads of the respective core components in the reference chip based on the actual utilization rates of the respective core components in the reference chip, the actual execution duration, and the hardware configuration specifications of the respective core components in the reference chip includes: Based on the actual utilization rate of the global memory, the actual execution duration, and the hardware configuration specifications of the global memory, determining the actual workload of the global memory.
[0009] In some embodiments, determining the actual workloads of the respective core components in the reference chip based on the actual utilization rates of the respective core components in the reference chip, the actual execution duration, and the hardware configuration specifications of the respective core components in the reference chip includes: Obtaining the number of instruction executions of each operation type of the vector core; Based on the number of instruction executions of each operation type, and the operands of each operation type of each warp in the reference chip for executing instructions, determining the actual workloads of the vector core for executing each operation type.
[0010] In some embodiments, determining the predicted occupancy durations of the respective core components when the target chip executes the current operator based on the actual workloads of the respective core components in the reference chip and the hardware configuration specifications of the respective core components in the target chip includes: Based on the actual workload of the tensor core, and the hardware configuration specifications of the tensor core in the target chip, determining the predicted occupancy duration of the tensor core when the target chip executes the current operator.
[0011] In some embodiments, determining the predicted occupancy durations of the respective core components when the target chip executes the current operator based on the actual workloads of the respective core components in the reference chip and the hardware configuration specifications of the respective core components in the target chip includes: Determine the predicted occupancy duration of the global memory when the target chip executes the current operator based on the actual workload of the global memory and the hardware configuration specifications of the global memory in the target chip.
[0012] In some embodiments, the determining the predicted occupancy duration of each core component when the target chip executes the current operator based on the actual workloads of the respective core components in the reference chip and the hardware configuration specifications of the respective core components in the target chip includes: Determine the predicted occupancy duration of the vector core in the target chip for each operation type based on the actual workloads of the vector core for each operation type and the hardware configuration specifications of the vector core in the target chip for each operation type; Determine the predicted occupancy duration of the vector core when the target chip executes the current operator based on the predicted occupancy durations of the vector core in the target chip for each operation type.
[0013] In some embodiments, the method further includes: Determine the predicted utilization rate of each core component based on the resource limitations of the current operator on each core component; Determine the predicted utilization rate of the target chip based on the predicted utilization rates of each core component.
[0014] In some embodiments, the method further includes: When the actual utilization rates of the respective core components in the reference chip are all less than a preset threshold, determine the actual execution duration as the predicted execution duration of the target chip for executing the current operator.
[0015] The present invention provides a performance prediction device for a chip to execute an operator, including: An acquisition module, configured to acquire the actual execution duration of the reference chip for executing the current operator and the actual utilization rates of the respective core components in the reference chip; A determination module, configured to determine the actual workloads of the respective core components in the reference chip based on the actual utilization rates of the respective core components in the reference chip, the actual execution duration, and the hardware configuration specifications of the respective core components in the reference chip; A calculation module, configured to determine the predicted occupancy duration of each core component when the target chip executes the current operator based on the actual workloads of the respective core components in the reference chip and the hardware configuration specifications of the respective core components in the target chip; the target chip and the reference chip have the same type of core components; A prediction module, configured to determine the predicted execution duration of the target chip for executing the current operator based on the predicted occupation durations of each core component when the target chip executes the current operator and the predicted utilization rate of the target chip.
[0016] The present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the performance prediction method for a chip to execute an operator is implemented.
[0017] The present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the performance prediction method for a chip to execute an operator is implemented.
[0018] The performance prediction method, device, electronic device, and medium for a chip to execute an operator provided by the present invention obtain the actual execution duration of a reference chip for executing the current operator and the actual utilization rates of each core component in the reference chip; determine the actual workloads of each core component in the reference chip based on the actual utilization rates, actual execution duration, and hardware configuration specifications of each core component in the reference chip; determine the predicted occupation durations of each core component when the target chip executes the current operator based on the actual workloads of each core component in the reference chip and the hardware configuration specifications of each core component in the target chip; determine the predicted execution duration of the target chip for executing the current operator based on the predicted occupation durations of each core component when the target chip executes the current operator and the predicted utilization rate of the target chip; since the target chip and the reference chip have the same type of core components, the execution duration of the reference chip for executing the current operator can be used to predict the execution duration of the target chip for executing the current operator, achieving accurate prediction of the execution performance of the operator on an unknown chip without actual execution or simulation; and during the prediction process, the hardware configuration specifications of the core components of the reference chip and the target chip are combined, and the prediction is carried out in units of the occupation duration of the core components, improving the accuracy of performance prediction and providing strong support for resource management, hardware design, software optimization, and system reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.
[0020] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0021] Figure 1 It is one of the schematic flowcharts of the method for predicting the performance of the chip executing operators provided by the present invention.
[0022] Figure 2 It is the schematic architecture diagram of the general graphics processing unit provided by the present invention.
[0023] Figure 3 It is the second schematic flowchart of the method for predicting the performance of the chip executing operators provided by the present invention.
[0024] Figure 4 It is the schematic structural diagram of the device for predicting the performance of the chip executing operators provided by the present invention.
[0025] Figure 5 It is the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0026] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0027] It should be noted that the terms "first", "second", etc. in the present invention are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units or modules does not necessarily have to be limited to those clearly listed steps or units or modules, but may include other steps or units or modules not clearly listed or inherent to these processes, methods, products, or devices.
[0028] Due to the continuous innovation of the architecture of general-purpose graphics processors, there are significant differences in the computing power, bandwidth, and micro-architecture design of each generation of general-purpose graphics processors. Accurately predicting the execution performance of various operators in an artificial intelligence model on the next-generation general-purpose graphics processor or general-purpose graphics processors with different hardware configurations without actual execution or simulation can not only help select a suitable general-purpose graphics processor for the artificial intelligence model but also provide a scientific basis for the hardware configuration and optimization of general-purpose graphics processors.
[0029] In related technologies, the Roofline model and architecture simulators, etc., are usually used to predict the performance of general-purpose graphics processors in executing operators, or historical performance data and mathematical models are used to predict the performance of general-purpose graphics processors in executing operators. These methods ignore the complexity at the architecture level of general-purpose graphics processors, cannot accurately simulate the architecture of general-purpose graphics processors, and rely on historical performance data, making them unable to adapt to different operators or general-purpose graphics processors with different architectures.
[0030] To address the deficiencies of related technologies, Figure 1 is one of the flow diagrams of the method for predicting the performance of a chip in executing an operator provided by the present invention. As Figure 1 shown, the method includes step 110, step 120, step 130, and step 140.
[0031] Step 110: Obtain the actual execution duration of a reference chip in executing the current operator, as well as the actual utilization rates of each core component in the reference chip.
[0032] Specifically, the execution entity of the method for predicting the performance of a chip in executing an operator provided by the embodiments of the present invention is a performance prediction device. This device can be implemented by software, such as a performance prediction program for a chip in executing an operator running on a computer; or it can be implemented by hardware, such as a computer or server that executes the method for predicting the performance of a chip in executing an operator.
[0033] The application scenario of the method provided by the embodiments of the present invention is to run a specific operator in a general-purpose graphics processor and predict the performance of the general-purpose graphics processor.
[0034] A reference chip refers to a general-purpose graphics processor used for comparison and evaluation. A general-purpose graphics processor with known performance, power consumption, architecture, and other characteristics can be selected as the reference chip. Core components are the main components in a chip that play a key role in the overall performance and function. They jointly determine the capabilities and application scenarios of the chip. These components are usually responsible for executing key computing tasks and data processing. For a general-purpose graphics processor, the types of its core components include tensor cores, vector cores, and global memory.
[0035] Figure 2It is a schematic diagram of the architecture of the general graphics processor provided by the present invention. As Figure 2 shown, the general graphics processor 200 may include multiple streaming processor clusters 210 (Streaming Processor Cluster, abbreviated as SPC) and a global memory 220 that can be accessed by all processors. The computing units in each streaming processor cluster 210 can generally be divided into tensor cores 211 (Tensor Core) and vector cores 212 (Vector Core).
[0036] The tensor core refers to a computing unit specifically used to execute tensor operations (such as matrix multiplication, convolution, etc.). The vector core refers to a computing unit used to execute vector operations (such as vector addition, multiplication, etc.), which is suitable for processing data parallel tasks. The global memory is a storage resource in the chip, accessed by the tensor core and vector core, for storing data and intermediate results.
[0037] An operator represents an entity that executes a specific operation or function. It can be a basic component that performs functions such as mathematical operations, data conversion, or logical judgment. In a neural network, an operator refers to the specific implementation of various layers or operations, which are the basic units that make up the neural network. The current operator can be the calculation of activation functions in a neural network, the normalization operation in the attention mechanism, and various custom operators, etc. The specific type of the operator is not specifically limited in the embodiments of the present invention.
[0038] The actual execution duration refers to the time elapsed from the start to the end of the execution of the current operator on the reference chip. The actual utilization rate refers to the busy degree of the core components during the execution of the computing task, which can be the ratio of the actual usage duration of each core component in the reference chip during the execution of the operator to the total available duration, usually expressed as a percentage.
[0039] It is possible to collect the process data of the reference chip executing the current operator, obtain the actual execution duration of the reference chip executing the current operator, and the actual utilization rate of each core component.
[0040] Step 120: Based on the actual utilization rate, actual execution duration of each core component in the reference chip, and the hardware configuration specifications of each core component in the reference chip, determine the actual workload of each core component in the reference chip.
[0041] Specifically, the hardware configuration specifications refer to the specific technical parameters and performance indicators of each core component in the chip, and these parameters determine the processing capacity, efficiency, and resource limitations of the core components.
[0042] The workload refers to the amount of work borne by each core component during the reference chip's execution of the current operator. For tensor cores and vector cores, the workload is the computational load, i.e., the amount of computational work; for global memory, the workload is the bandwidth load, i.e., the amount of bandwidth used.
[0043] By comprehensively considering the actual utilization rate, actual execution duration, and hardware configuration specifications of each core component in the reference chip, the actual workload borne by these core components during the execution of the current operator is determined.
[0044] For example, for a tensor core, its actual workload can be expressed as the actual utilization rate of the core during the execution of the operator multiplied by the actual execution duration, and then combined with its hardware configuration specifications (such as clock frequency, etc.) for measurement.
[0045] Step 130: Based on the actual workloads of each core component in the reference chip and the hardware configuration specifications of each core component in the target chip, determine the predicted occupancy durations of each core component when the target chip executes the current operator; the target chip and the reference chip have the same types of core components.
[0046] Specifically, the target chip refers to the chip that is currently being designed, evaluated, or optimized. It is the main object of research and development work. Since the target chip and the reference chip have the same types of core components, the reference chip can be used as a benchmark to compare the performance and characteristics of the target chip. The embodiments of the present invention do not specifically limit the instruction set architecture and microarchitecture of the chip, etc.
[0047] Since the types of core components are the same, the occupancy durations (busy time) of each core component when the target chip executes the current operator can be predicted based on the actual workloads of each core component in the reference chip and the hardware configuration specifications of each core component in the target chip, to obtain the predicted occupancy durations. The occupancy duration refers to the time occupied by a certain core component when executing a specific task. It reflects the length of time that the core component is in a busy state during the completion of the task.
[0048] Step 140: Based on the predicted occupancy durations of each core component when the target chip executes the current operator, and the predicted utilization rate of the target chip, determine the predicted execution duration of the target chip when executing the current operator.
[0049] Specifically, the predicted utilization rate refers to the predicted value of the utilization rate of each core component in the target chip when executing the current operator.
[0050] The predicted execution duration of the target chip when executing the current operator, that is, the predicted value of the execution duration, can be calculated based on the predicted occupancy durations of each core component when the target chip executes the current operator and the predicted utilization rate of the target chip.
[0051] There is a direct relationship between the predicted execution duration of the target chip for executing the current operator and the performance of the target chip for executing the current operator. The predicted execution duration is an important indicator to measure performance, which reflects the execution efficiency of the operator on the target chip. It can be understood that the larger the predicted execution duration, the lower the performance of the target chip for executing the current operator.
[0052] The performance prediction method for a chip to execute an operator provided by an embodiment of the present invention obtains the actual execution duration of a reference chip for executing the current operator and the actual utilization rates of various core components in the reference chip; determines the actual workloads of various core components in the reference chip based on the actual utilization rates, actual execution duration, and hardware configuration specifications of various core components in the reference chip; determines the predicted occupancy durations of various core components when the target chip executes the current operator based on the actual workloads of various core components in the reference chip and the hardware configuration specifications of various core components in the target chip; determines the predicted execution duration of the target chip for executing the current operator based on the predicted occupancy durations of various core components when the target chip executes the current operator and the predicted utilization rate of the target chip; since the target chip and the reference chip have the same type of core components, the execution duration of the reference chip for executing the current operator can be used to predict the execution duration of the target chip for executing the current operator, achieving accurate prediction of the execution performance of the operator on an unknown chip without actual execution or simulation; and during the prediction process, the hardware configuration specifications of the core components of the reference chip and the target chip are combined, and the prediction is carried out in units of the occupancy duration of the core components, improving the accuracy of performance prediction and providing strong support for resource management, hardware design, software optimization, and system reliability.
[0053] It should be noted that each embodiment of the present invention can be freely combined, the order can be swapped, or it can be executed independently, and does not need to rely on or depend on a fixed execution order.
[0054] In some embodiments, determining the actual workloads of various core components in the reference chip based on the actual utilization rates, actual execution duration, and hardware configuration specifications of various core components in the reference chip includes: Obtaining the actual clock frequency of the reference chip for executing the current operator and the peak frequency of the reference chip; Determining the actual workload of the tensor core based on the actual utilization rate, actual clock frequency, peak frequency, actual execution duration, and hardware configuration specifications of the tensor core.
[0055] Specifically, the clock frequency refers to the main clock frequency of the chip. It represents the number of oscillation cycles per second of the chip and is an important indicator to measure the running speed of the chip. The running speed of the chip is usually measured in the number of operations executed per second or the amount of data processed per second.
[0056] The clock frequency can reflect the operating speed of the tensor core. The actual clock frequency refers to the frequency at which the reference chip actually operates when executing the current operator. The actual clock frequency can reflect the speed of the tensor core during actual operation and directly affects its computing performance. The peak frequency refers to the highest operating frequency of the reference chip in theory and can reflect the maximum operating speed of the tensor core under ideal conditions.
[0057] The ratio of the actual clock frequency to the peak frequency can reflect the efficiency of the tensor core during actual operation. The closer the ratio is to 1, the closer the actual operating frequency of the tensor core is to its theoretical peak frequency, and the more fully the performance is utilized.
[0058] Collect the process data of the reference chip executing the current operator to obtain the actual clock frequency and actual execution duration of the reference chip executing the current operator, as well as the actual utilization rate of the tensor core. By querying the hardware specification (specification, spec) of the reference chip, the peak frequency of the reference chip can be obtained.
[0059] The peak frequency of the reference chip ChipA is Peak_clk; the actual clock frequency when executing the current operator is Real_clk; the actual execution duration is ChipA_duration; the hardware configuration specification of the tensor core is ChipA_Tcore_spec, which represents the theoretical peak computing power of the tensor core of the reference chip; the actual utilization rate of the tensor core is ChipA_Tcore_util. The actual workload of the tensor core ChipA_Tcore_GFLOP can be calculated as follows: ChipA_Tcore_GFLOP = ChipA_Tcore_util Real_clk / Peak_clk ChipA_duration ChipA_Tcore_spec.
[0060] The method for predicting the performance of a chip executing an operator provided by the embodiments of the present invention can accurately determine the actual workload of the tensor core according to the actual utilization rate, actual clock frequency, peak frequency, actual execution duration of the tensor core, and the hardware configuration specification of the tensor core, improving the accuracy of performance prediction.
[0061] In some embodiments, based on the actual utilization rate, actual execution duration of each core component in the reference chip, and the hardware configuration specification of each core component in the reference chip, determining the actual workload of each core component in the reference chip includes: Based on the actual utilization rate, actual execution duration, and hardware configuration specification of the global memory, determining the actual workload of the global memory.
[0062] Specifically, by querying the hardware specification (spec) of the reference chip, the hardware configuration specifications of the global memory can be obtained. The hardware configuration specifications of the global memory may include the capacity, bandwidth, cache line size, etc. of the global memory. In the embodiments of the present invention, the access amount of the memory bandwidth can be selected as the workload.
[0063] The actual utilization rate of the global memory is ChipA_Dram_util; the actual execution duration is ChipA_duration; the hardware configuration specifications of the global memory are ChipA_Dram_spec, indicating the bandwidth design parameters of the global memory of the reference chip. The actual workload ChipA_Dram_GB of the global memory can be calculated, representing the actual memory access amount of the current operator on the target chip: ChipA_Dram_GB = ChipA_Dram_util ChipA_Dram_spec ChipA_duration.
[0064] The performance prediction method for the chip to execute the operator provided by the embodiments of the present invention can accurately determine the actual workload of the global memory according to the actual utilization rate, actual execution duration, and hardware configuration specifications of the global memory, improving the accuracy of performance prediction.
[0065] In some embodiments, based on the actual utilization rate, actual execution duration, and hardware configuration specifications of each core component in the reference chip, determining the actual workload of each core component in the reference chip includes: Obtaining the number of instruction executions of each operation type of the vector core; Based on the number of instruction executions of each operation type and the operands of each operation type instruction executed by each warp in the reference chip, determining the actual workload of the vector core for executing each operation type.
[0066] Specifically, as a general-purpose computing unit, the vector core generally includes instructions such as vector floating-point calculation, vector integer calculation, vector special function calculation, vector logic calculation, scalar calculation, data access, conditional control, and synchronization processing. When calculating the actual workload of the vector core, various operation operations that the vector core can execute can be classified according to the operation type.
[0067] According to the instruction set of the vector core, it can be determined that the operation types of the instructions executed by the vector core can be divided into the following four categories: I. Floating-point operation type (float), including vector floating-point calculations (addition, subtraction, multiplication, division, square root, etc.); II. Integer operation type (int), including vector integer calculations (bit operations, addition, subtraction, multiplication, division, etc.); III. Special function operation type (sfu), including vector special function calculations (exponential, logarithmic, trigonometric functions, etc.); IV. General operation type (general). The instructions in the general operation type do not belong to floating-point operation instructions, integer operation instructions, and special function operation instructions, including vector logical operations (comparison, selection, masking, etc.), scalar calculations (various operations related to control flow), memory access operations (read and write operations of global memory, shared memory, texture memory, etc.), and thread control and synchronization operations (branch prediction, thread synchronization, etc.).
[0068] It is possible to count the number of executed instructions of each operation type of the vector core, and then determine the actual workload of the vector core for each operation type based on the number of executed instructions of each operation type and the operands of each operation type of each warp in the reference chip.
[0069] The number of executed instructions of the vector core for the floating-point operation type is ChipA_Vcore_float_instructions, and the operand of each floating-point operation type of each warp in the reference chip is ChipA_float_GOP_per_Warp_Inst. The actual workload Vcore_float_GOP of the vector core for each floating-point operation type can be calculated as follows: Vcore_float_GOP = ChipA_Vcore_float_instructions ChipA_float_GOP_per_Warp_Inst.
[0070] The number of executed instructions of the vector core for the integer operation type is ChipA_Vcore_int_instructions, and the operand of each integer operation type of each warp in the reference chip is ChipA_int_GOP_per_Warp_Inst. The actual workload Vcore_int_GOP of the vector core for each integer operation type can be calculated as follows: Vcore_int_GOP = ChipA_Vcore_int_instructions ChipA_int_GOP_per_Warp_Inst.
[0071] The number of instructions executed by the vector core for special function operation types is ChipA_Vcore_sfu_instructions. The operand for each instruction of each special function operation type executed by each warp in the reference chip is ChipA_sfu_GOP_per_Warp_Inst. The actual workload Vcore_sfu_GOP of the vector core for each special function operation type can be calculated as follows: Vcore_sfu_GOP = ChipA_Vcore_sfu_instructions ChipA_sfu_GOP_per_Warp_Inst.
[0072] The number of instructions executed by the vector core for general operation types is ChipA_Vcore_general_instructions. The operand for each instruction of each general operation type executed by each warp in the reference chip is ChipA_general_GOP_per_Warp_Inst. The actual workload Vcore_general_GOP of the vector core for each general operation type can be calculated as follows: Vcore_general_GOP = ChipA_Vcore_general_instructions ChipA_general_GOP_per_Warp_Inst.
[0073] By adding up the actual workloads of each operation type, the actual workload of the vector core can be obtained.
[0074] The method for predicting the performance of the operator executed by the chip provided by the embodiment of the present invention fully considers the complexity of the instruction types executed by the vector core, and adopts a method for determining the actual workload for different operation types, which can accurately determine the actual workload of the vector core and improve the accuracy of performance prediction.
[0075] In some embodiments, based on the actual workloads of each core component in the reference chip and the hardware configuration specifications of each core component in the target chip, determining the predicted occupation duration of each core component when the target chip executes the current operator includes: Based on the actual workload of the tensor core and the hardware configuration specifications of the tensor core in the target chip, determining the predicted occupation duration of the tensor core when the target chip executes the current operator.
[0076] Specifically, the actual workload of the tensor core is ChipA_Tcore_GFLOP; the hardware configuration specification of the tensor core in the target chip ChipB is ChipB_Tcore_spec_GFLOPS, which represents the theoretical peak computing power of the tensor core of the target chip; the predicted occupancy duration Tcore_busy_time of the tensor core when the target chip executes the current operator can be calculated as follows: Tcore_busy_time = ChipA_Tcore_GFLOP / ChipB_Tcore_spec_GFLOPS.
[0077] The performance prediction method for a chip to execute an operator provided by the embodiments of the present invention can accurately determine the predicted occupancy duration of the tensor core when the target chip executes the current operator, improving the accuracy of performance prediction.
[0078] In some embodiments, based on the actual workloads of the respective core components in the reference chip and the hardware configuration specifications of the respective core components in the target chip, determining the predicted occupancy durations of the respective core components when the target chip executes the current operator includes: Based on the actual workload of the global memory and the hardware configuration specification of the global memory in the target chip, determining the predicted occupancy duration of the global memory when the target chip executes the current operator.
[0079] Specifically, the actual workload of the global memory is ChipA_Dram_GB; the hardware configuration specification of the global memory in the target chip is ChipB_Dcore_spec_GB / s, which represents the theoretical peak memory bandwidth of the target chip; the predicted occupancy duration Dram_busy_time of the global memory when the target chip executes the current operator can be calculated as follows: Dram_busy_time = ChipA_Dram_GB / ChipB_Dcore_spec_GB / s.
[0080] The performance prediction method for a chip to execute an operator provided by the embodiments of the present invention can accurately determine the predicted occupancy duration of the global memory when the target chip executes the current operator, improving the accuracy of performance prediction.
[0081] In some embodiments, based on the actual workloads of the respective core components in the reference chip and the hardware configuration specifications of the respective core components in the target chip, determining the predicted occupancy durations of the respective core components when the target chip executes the current operator includes: Based on the actual workloads of the vector core for executing various operation types and the hardware configuration specifications of the vector core in the target chip for executing various operation types, determining the predicted occupancy durations of the vector core in the target chip for executing various operation types; Based on the predicted occupancy duration of each operation type executed by the vector core in the target chip, determine the predicted occupancy duration of the vector core when the target chip executes the current operator.
[0082] Specifically, the actual workload of the vector core executing floating-point operation types in the reference chip ChipA is ChipA_Vcore_float_GOP, and the hardware configuration specification of the vector core executing floating-point operation types in the target chip ChipB is ChipB_Vcore_float_spec_GFLOPS; the predicted occupancy duration of the vector core executing floating-point operation types in the target chip, Vcore_float_issue_time, can be calculated as follows: Vcore_float_issue_time = ChipA_Vcore_float_GOP / ChipB_Vcore_float_spec_GFLOPS.
[0083] The actual workload of the vector core executing integer operation types in the reference chip ChipA is ChipA_Vcore_int_GOP, and the hardware configuration specification of the vector core executing integer operation types in the target chip ChipB is ChipB_Vcore_int_spec_GFLOPS; the predicted occupancy duration of the vector core executing integer operation types in the target chip, Vcore_int_issue_time, can be calculated as follows: Vcore_int_issue_time = ChipA_Vcore_int_GOP / ChipB_Vcore_int_spec_GFLOPS.
[0084] The actual workload of the vector core executing special function operation types in the reference chip ChipA is ChipA_Vcore_sfu_GOP, and the hardware configuration specification of the vector core executing special function operation types in the target chip ChipB is ChipB_Vcore_sfu_spec_GFLOPS; the predicted occupancy duration of the vector core executing special function operation types in the target chip, Vcore_sfu_issue_time, can be calculated as follows: Vcore_sfu_issue_time = ChipA_Vcore_sfu_GOP / ChipB_Vcore_sfu_spec_GFLOPS.
[0085] The actual workload of the vector core in the reference chip ChipA for executing general operation types is ChipA_Vcore_general_GOP, and the hardware configuration specification of the vector core in the target chip ChipB for executing general operation types is ChipB_Vcore_general_spec_GFLOPS; the predicted occupancy duration Vcore_general_issue_time of the vector core in the target chip for executing general operation types can be calculated as follows: Vcore_general_issue_time = ChipA_Vcore_general_GOP / ChipB_Vcore_general_spec_GFLOPS.
[0086] The predicted occupancy durations of the vector core in the target chip for executing each operation type can be summed up to obtain the predicted occupancy duration Vcore_busy_time of the vector core when the target chip executes the current operator: Vcore_busy_time = Vcore_float_issue_time + Vcore_int_issue_time + Vcore_sfu_issue_time + Vcore_general_issue_time.
[0087] The performance prediction method for the chip to execute the operator provided by the embodiments of the present invention can accurately determine the predicted occupancy duration of the vector core when the target chip executes the current operator, improving the accuracy of performance prediction.
[0088] In some embodiments, based on the predicted occupancy durations of each core component when the target chip executes the current operator and the predicted utilization rate of the target chip, the predicted execution duration of the target chip for executing the current operator is determined.
[0089] Specifically, according to the predicted occupancy duration Tcore_busy_time of the tensor core, the predicted occupancy duration Vcore_busy_time of the vector core, the predicted occupancy duration Dram_busy_time of the global memory, and the predicted utilization rate Predict_utilization of the target chip, the predicted execution duration Predict_duration of the target chip for executing the current operator can be calculated as follows: Predict_duration = Max(Tcore_busy_time, Vcore_busy_time, Dram_busy_time) / Predict_utilization.
[0090] The above formula indicates that the predicted execution duration of the current operator on the target chip is equal to the longest predicted occupancy duration among all core components divided by the predicted utilization rate. This reflects the characteristic that the bottleneck resource for the target chip to execute the current operator determines the final performance.
[0091] In some embodiments, the method further includes: Determining the predicted utilization rate of each core component based on the resource limitations of the current operator on each core component; Determining the predicted utilization rate of the target chip based on the predicted utilization rates of each core component.
[0092] Specifically, in deep learning and high-performance computing, the current operator may be subject to resource limitations of each core component, which will affect the predicted utilization rate of each core component.
[0093] On the one hand, the performance of the current operator may be limited by the computing power or memory bandwidth of the tensor core. The upper limits of the tensor core and global memory utilization rates are mainly determined by the algorithm and pipelining implementation of the operator, and are less affected by the chip specifications and microarchitecture. Therefore, the predicted utilization rate of the target chip, Predict_utilization, can take the maximum value of the predicted utilization rate of the tensor core, ChipA_Tcore_util, and the predicted utilization rate of the global memory, ChipA_Dram_util. On the other hand, the performance of the current operator may be limited by the computing power of the vector core. The predicted utilization rate of the vector core, ChipA_Vcore_util, usually cannot reach 100%. However, the instruction granularity of the vector core is small, and its upper limit of utilization rate is less affected by the algorithm and pipelining implementation of the operator. It can be determined based on experience that 80% is the minimum value of the upper limit of the predicted utilization rate of the vector core.
[0094] The predicted utilization rate of the target chip, Predict_utilization, can be expressed as: 1. In the case where the computing power of the tensor core or the access to the memory bandwidth is limited: Predict_utilization = Max(ChipA_Tcore_util, ChipA_Dram_util); 2. In the case where the computing power of the vector core is considered limited: Predict_utilization = Max(ChipA_Vcore_util, 0.8).
[0095] The method for predicting the performance of a chip executing an operator provided by the embodiments of the present invention can adaptively adjust the predicted utilization rate according to the resource limitation situation of the current operator in the target chip, improving the accuracy of performance prediction.
[0096] In some embodiments, the method further includes: When the actual utilization rate of each core component in the reference chip is less than the preset threshold, determine the actual execution duration as the predicted execution duration of the target chip for executing the current operator.
[0097] Specifically, some operators are greatly affected by the chip startup overhead and scheduling delay, and are less restricted by the core computing power and memory bandwidth. This is manifested in that when these operators are executed in the reference chip, the actual utilization rate of each core component is relatively low. Therefore, when calculating the predicted execution duration of these operators, it can be considered that the execution durations of these operators in the reference chip and the target chip are equal.
[0098] The preset threshold can be set as needed. Compare the actual utilization rate of each core component in the reference chip with the preset threshold. If the actual utilization rate of each core component is less than the preset threshold, it can be considered that the predicted execution duration Predict_duration of the target chip for executing the current operator is equal to the actual execution duration ChipA_duration of the reference chip for executing the current operator.
[0099] For example, the preset threshold is 0.15, indicating that the set utilization rate is 15%. If the maximum value among the actual utilization rates in the tensor core, vector core, and global memory is less than the preset threshold, Max(ChipA_Tcore_util, ChipA_Vcore_util, ChipA_Dram_util) < 0.15, then it can be determined that Predict_duration = ChipA_duration.
[0100] The method for predicting the performance of a chip executing an operator provided by the embodiments of the present invention takes into account the special case that some operators are greatly affected by the chip startup overhead and scheduling delay, and improves the accuracy of performance prediction.
[0101] Figure 3 It is the second flowchart of the method for predicting the performance of a chip executing an operator provided by the present invention. As Figure 3 shown, the method includes: Step 310: Collect performance data of the reference chip.
[0102] The detailed performance metrics of the current operator executed on the reference chip can be collected by using a graphics processor performance analysis tool. The collected metrics include tensor core utilization rate, memory read / write bandwidth utilization rate, instruction count of the vector core, actual execution duration of the current operator, and actual clock frequency of the reference chip, etc.
[0103] For example, when a certain convolution operator is executed on the reference chip ChipA, the following can be collected: the actual execution duration is 2.5 milliseconds; the actual utilization rate of the tensor core is 85%; the actual utilization rate of the vector core is 30%; the actual utilization rate of the memory bandwidth is 65%; the actual clock frequency is 1.7 GHz (the peak clock frequency is 1.8 GHz). GHz stands for gigahertz.
[0104] Step 320: Calculate the actual workload of each core component according to the hardware configuration specifications of the reference chip.
[0105] The hardware configuration specifications of the reference chip are as follows: the theoretical computing power of the tensor core is 312 tera floating point operations per second (TFLOPS); the core computing power of the vector core is 39 TFLOPS; the memory bandwidth is 1.5 terabytes per second (TB / s).
[0106] The actual workload of each core component of the reference chip can be calculated. For example, the actual workload of the tensor core ChipA_Tcore_GFLOP is 0.85 (1.7 / 1.8) 2.5 312 = 626.2 giga floating point operations (GFLOP). The actual workload of the global memory ChipA_Dram_GB is 0.65 1.5 2.5 = 2.4 gigabytes (GB).
[0107] Step 330: Predict the performance of the target chip according to the hardware configuration specifications of the target chip.
[0108] The hardware configuration specifications of the target chip ChipB are as follows: the theoretical computing power of the tensor core is 1000 TFLOPS; the core computing power of the vector core is 32 TFLOPS; the memory bandwidth is 1.6 TB / s; the clock frequency is 1.0 GHz.
[0109] According to the actual workload of each core component of the reference chip and the hardware configuration specifications of each core component of the target chip, the predicted occupancy duration of each core component can be determined. For example, the predicted occupancy duration of the tensor core Tcore_busy_time = 626.2 / 1000 = 0.626 ms; the predicted occupancy duration of the global memory Dram_busy_time = 2.4 / 1.6 = 1.5 ms; the predicted occupancy duration of the vector core Vcore_busy_time = 0.92 ms.
[0110] The convolution operator is restricted by the resources of the tensor kernel, and its prediction utilization rate Predict_utilization = Max(0.85, 0.65) = 0.85. The predicted execution duration Predict_duration of the target chip = Max(0.626, 1.5, 0.92) / 0.85 = 1.76 ms.
[0111] Through actual testing, the actual execution time of this operator on the target chip is 1.8 ms, and the prediction error is only 3.3%, verifying the effectiveness of this method.
[0112] The performance prediction method for chip execution operators provided by the embodiments of the present invention is applicable to different chip products with homogeneous microarchitectures, and supports performance mapping prediction from known chips to unannounced chips; it provides important references for chip design, algorithm optimization, and system scheduling, can be used for early performance evaluation of new-generation chips, and accelerates the product R & D cycle; it does not require actual operation or construction of complex simulators, significantly reducing the prediction cost and time.
[0113] The device provided by the embodiments of the present invention will be described below, and the device described below can be correspondingly referred to the method described above.
[0114] Figure 4 It is a schematic structural diagram of a performance prediction device for chip execution operators provided by the present invention, as Figure 4 shown. The device includes: An acquisition module 410, configured to acquire the actual execution duration of a reference chip executing the current operator, and the actual utilization rates of each core component in the reference chip; A determination module 420, configured to determine the actual workloads of each core component in the reference chip based on the actual utilization rates of each core component in the reference chip, the actual execution duration, and the hardware configuration specifications of each core component in the reference chip; A calculation module 430, configured to determine the predicted occupancy durations of each core component when the target chip executes the current operator based on the actual workloads of each core component in the reference chip and the hardware configuration specifications of each core component in the target chip; the target chip and the reference chip have the same type of core components; A prediction module 440, configured to determine the predicted execution duration of the target chip executing the current operator based on the predicted occupancy durations of each core component when the target chip executes the current operator and the predicted utilization rate of the target chip.
[0115] The performance prediction device for a chip to execute an operator provided by an embodiment of the present invention obtains the actual execution duration of a reference chip to execute the current operator and the actual utilization rate of each core component in the reference chip; determines the actual workload of each core component in the reference chip based on the actual utilization rate of each core component in the reference chip, the actual execution duration, and the hardware configuration specifications of each core component in the reference chip; determines the predicted occupancy duration of each core component when the target chip executes the current operator based on the actual workload of each core component in the reference chip and the hardware configuration specifications of each core component in the target chip; determines the predicted execution duration of the target chip to execute the current operator based on the predicted occupancy duration of each core component when the target chip executes the current operator and the predicted utilization rate of the target chip; since the target chip and the reference chip have the same type of core components, the execution duration of the reference chip to execute the current operator can be used to predict the execution duration of the target chip to execute the current operator, achieving accurate prediction of the execution performance of the operator on an unknown chip without actual execution or simulation; and during the prediction process, the hardware configuration specifications of the core components of the reference chip and the target chip are combined, and the prediction is carried out in units of the occupancy duration of the core components, improving the accuracy of performance prediction and providing strong support for resource management, hardware design, software optimization, and system reliability.
[0116] Figure 5 is a schematic structural diagram of an electronic device provided by the present invention, as Figure 5 shown, the electronic device may include: a processor (Processor) 510, a communication interface (Communications Interface) 520, a memory (Memory) 530, and a communication bus (Communications Bus) 540. Among them, the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logical commands in the memory 530 to execute the methods described in the above embodiments, for example: obtain the actual execution duration of a reference chip to execute the current operator and the actual utilization rate of each core component in the reference chip; determine the actual workload of each core component in the reference chip based on the actual utilization rate of each core component in the reference chip, the actual execution duration, and the hardware configuration specifications of each core component in the reference chip; determine the predicted occupancy duration of each core component when the target chip executes the current operator based on the actual workload of each core component in the reference chip and the hardware configuration specifications of each core component in the target chip; the target chip and the reference chip have the same type of core components; determine the predicted execution duration of the target chip to execute the current operator based on the predicted occupancy duration of each core component when the target chip executes the current operator and the predicted utilization rate of the target chip.
[0117] In addition, when the logical commands in the above-mentioned memory can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several commands for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0118] The processor in the electronic device provided in the embodiments of the present invention can call the logical instructions in the memory to implement the above method. The specific implementation manner is the same as that of the foregoing method implementation manner, and the same beneficial effects can be achieved, which will not be elaborated herein.
[0119] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to execute the methods provided in the above-mentioned embodiments.
[0120] The specific implementation manner is the same as that of the foregoing method implementation manner, and the same beneficial effects can be achieved, which will not be elaborated herein.
[0121] The embodiments of the present invention provide a computer program product, including a computer program. When the computer program is executed by a processor, it implements the method as described above.
[0122] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0123] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting the performance of a chip execution operator, characterized in that: include: Obtaining the actual execution time of the reference chip executing the current operator and the actual utilization rate of each core component in the reference chip; Determine the actual workload of each core component in the reference chip based on the actual utilization rate of each core component in the reference chip, the actual execution time, and the hardware configuration specifications of each core component in the reference chip; Determine the predicted occupancy time of each core component when the target chip executes the current operator based on the actual workload of each core component in the reference chip and the hardware configuration specifications of each core component in the target chip; The target chip and the reference chip have the same type of core components; The predicted execution time of the target chip executing the current operator is determined based on the predicted occupancy time of each core component when the target chip executes the current operator and the predicted utilization rate of the target chip.
2. The method for predicting the performance of a chip execution operator according to claim 1, characterized in that: The core components include a tensor core, a vector core and a global memory; the global memory is used for access by the tensor core and the vector core.
3. The performance prediction method of a chip execution operator according to claim 2, characterized in that: The determining the actual workload of each core component in the reference chip based on the actual utilization rate of each core component in the reference chip, the actual execution time, and the hardware configuration specification of each core component in the reference chip includes: Obtaining an actual clock frequency of the reference chip executing the current operator and a peak frequency of the reference chip; An actual workload of the tensor core is determined based on the actual utilization of the tensor core, the actual clock frequency, the peak frequency, the actual execution duration, and a hardware configuration specification of the tensor core.
4. The method for predicting the performance of a chip execution operator according to claim 2, characterized in that: The determining the actual workload of each core component in the reference chip based on the actual utilization rate of each core component in the reference chip, the actual execution time, and the hardware configuration specification of each core component in the reference chip includes: An actual workload of the global memory is determined based on the actual utilization rate of the global memory, the actual execution duration, and a hardware configuration specification of the global memory.
5. The method for predicting the performance of a chip execution operator according to claim 2, characterized in that: The determining the actual workload of each core component in the reference chip based on the actual utilization rate of each core component in the reference chip, the actual execution time, and the hardware configuration specification of each core component in the reference chip includes: Obtaining the number of instruction executions of each operation type of the vector core; Based on the number of times instructions of each operation type are executed and the number of operations of each operation type executed by each thread warp in the reference chip, the actual workload of the vector core executing each operation type is determined.
6. The method for predicting the performance of a chip execution operator according to claim 3, characterized in that: The step of determining the predicted occupancy time of each core component when the target chip executes the current operator based on the actual workload of each core component in the reference chip and the hardware configuration specifications of each core component in the target chip comprises: Based on the actual workload of the tensor core and the hardware configuration specification of the tensor core in the target chip, a predicted occupancy time of the tensor core when the target chip executes the current operator is determined.
7. The method for predicting the performance of a chip execution operator according to claim 4, characterized in that: The step of determining the predicted occupancy time of each core component when the target chip executes the current operator based on the actual workload of each core component in the reference chip and the hardware configuration specifications of each core component in the target chip comprises: Based on the actual workload of the global memory and the hardware configuration specification of the global memory in the target chip, the predicted occupation time of the global memory when the target chip executes the current operator is determined.
8. The method for predicting the performance of a chip execution operator according to claim 5, characterized in that: The step of determining the predicted occupancy time of each core component when the target chip executes the current operator based on the actual workload of each core component in the reference chip and the hardware configuration specifications of each core component in the target chip comprises: Determine the predicted occupancy time of each operation type executed by the vector core in the target chip based on the actual workload of each operation type executed by the vector core and the hardware configuration specification of each operation type executed by the vector core in the target chip; Based on the predicted occupancy time of the vector core in the target chip executing each type of operation, the predicted occupancy time of the vector core when the target chip executes the current operator is determined.
9. The method for predicting the performance of a chip execution operator according to claim 1, characterized in that: The method further comprises: Determining predicted utilization rates of each core component based on resource constraints imposed on the current operator in each core component; Based on the predicted utilization of each core component, the predicted utilization of the target chip is determined.
10. The method for predicting the performance of a chip execution operator according to claim 1, characterized in that: The method further comprises: When the actual utilization rate of each core component in the reference chip is less than a preset threshold, the actual execution duration is determined as the predicted execution duration of the target chip executing the current operator.
11. A performance prediction device for a chip execution operator, characterized in that: include: An acquisition module is used to acquire the actual execution time of the reference chip executing the current operator and the actual utilization rate of each core component in the reference chip; A determination module, configured to determine an actual workload of each core component in the reference chip based on an actual utilization rate of each core component in the reference chip, the actual execution duration, and a hardware configuration specification of each core component in the reference chip; A calculation module, for determining the predicted occupancy time of each core component when the target chip executes the current operator based on the actual workload of each core component in the reference chip and the hardware configuration specifications of each core component in the target chip; the target chip and the reference chip have the same type of core components; The prediction module is used to determine the predicted execution time of the target chip executing the current operator based on the predicted occupancy time of each core component when the target chip executes the current operator and the predicted utilization rate of the target chip.
12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the performance prediction method of the chip execution operator described in any one of claims 1 to 10 is implemented.
13. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for predicting the performance of a chip execution operator according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Method, device and equipment for testing calculation performance of AI chip, and medium
CN113568821A
Chip performance determination method and device, electronic equipment and storage medium
CN115840685A
Operator performance determination method and device, computing equipment and storage medium
CN117667330A
Chip performance prediction method and device, equipment, medium and product
CN119047389A
Cited By
Time length prediction method and device, equipment, storage medium and program product
CN121958052A