Vector kernel calculation overhead quantification method and device, electronic equipment and storage medium

By analyzing the instruction execution information of the vector kernel execution operator, calculating the calculation load and overhead of each operation type, the accuracy of the vector kernel calculation overhead quantization is solved, and a comprehensive evaluation of the general graphics processor operator is realized, providing a basis for architecture design and optimization.

CN120179977AActive Publication Date: 2025-06-20SHANGHAI BIREN TECH CO LTD

Patent Information

Application Number
CN202510647064.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-20
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

How to accurately calculate the calculation overhead of vector cores when processing different operators, the current mainstream methods cannot effectively quantify the calculation overhead of vector cores.

Method used

By determining the number of instruction execution times of each operation type based on the instruction execution information of the current operator in the target chip, the calculation load and calculation overhead of each operation type are calculated, and the calculation overhead of the current operator in the target chip is finally determined.

Benefits of technology

It realizes the accurate quantification of the calculation overhead of the vector core when processing different operators, and can comprehensively and accurately evaluate the calculation overhead of the general graphics processor for the execution of various operators, providing a scientific basis for the general graphics processor architecture design and operator optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179977A_ABST
    Figure CN120179977A_ABST
Patent Text Reader

Abstract

The invention provides a vector kernel calculation overhead quantification method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: determining the instruction execution frequency of each operation type based on the instruction execution information of a current operator in a target chip; the target chip comprises a vector kernel; the vector kernel executes the current operator; determining the calculation load of each operation type based on the instruction execution times of each operation type; determining the calculation overhead of each operation type based on the calculation load of each operation type and the configuration calculation power of each operation type in the target chip; and determining the calculation overhead of the current operator in the target chip based on the calculation overhead of each operation type. According to the method and the device provided by the invention, the calculation overhead of the calculation vector core in processing different operators is accurately quantified, and a scientific basis is provided for architecture design and operator optimization of a general graphic processor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, electronic device and storage medium for quantifying vector core computing overhead. Background Art

[0002] With the rapid development of artificial intelligence technology, the scale and complexity of neural network models are growing exponentially, and the demand for computing power is increasing. General Purpose Graphics Processing Unit (GPGPU) has become the core hardware infrastructure supporting artificial intelligence (AI) training and reasoning with its highly parallel computing architecture and powerful floating-point computing capabilities. The current mainstream GPGPU architecture usually contains two core computing units: Tensor Core optimized for matrix operations and Vector Core for general vector computing.

[0003] Tensor Core focuses on accelerating specific operations such as matrix multiplication in deep learning, while Vector Core is responsible for handling more diverse parallel computing tasks. In actual application scenarios, although Tensor Core provides significant peak performance, a large number of complex algorithms still rely on Vector Core to complete, such as activation function calculations in convolutional neural networks, normalization operations in attention mechanisms, and various custom operators.

[0004] Therefore, how to accurately calculate the computational overhead of vector kernels when processing different operators has become a technical problem that needs to be solved urgently in the industry. Summary of the invention

[0005] The present invention provides a method, device, electronic device and storage medium for quantifying the computational overhead of a vector core, which are used to solve the technical problem of how to accurately calculate the computational overhead of a vector core when processing different operators.

[0006] The present invention provides a vector kernel computing overhead quantification method, comprising: Determining the number of instruction executions of each operation type based on instruction execution information of the current operator in the target chip; the target chip includes a vector core; the vector core executes the current operator; Determining the computational load of each operation type based on the number of instruction executions of each operation type; Determine the computational overhead of each computational type based on the computational load of each computational type and the configured computing power of each computational type in the target chip; Determine the computing overhead of the current operator in the target chip based on the computing overhead of each operation type.

[0007] In some embodiments, determining the computing load of each operation type based on the number of instruction executions of each operation type includes: Determine the computing load of each operation type based on the number of instruction executions of each operation type and the warp size for executing each operation type in the target chip.

[0008] In some embodiments, before determining the computing load of each operation type based on the number of instruction executions of each operation type and the warp size for executing each operation type in the target chip, the method further includes: In the case where the current operation type includes multiple operation subtypes, determine the operation subtype with equivalent computing power among the multiple operation subtypes; Based on the operands of each operation subtype and the configured computing power of each operation subtype in the target chip, determine the quantity conversion relationship between the instructions of each operation subtype and the instructions of the operation subtype with equivalent computing power when the computing power is equivalent; Based on the quantity conversion relationship and the number of instruction executions of each operation subtype, determine the number of instruction executions of the operation subtype with equivalent computing power; Determine the number of instruction executions of the operation subtype with equivalent computing power as the number of instruction executions of the current operation type.

[0009] In some embodiments, determining the computing overhead of each operation type based on the computing load of each operation type and the configured computing power of each operation type in the target chip includes: Determine the computing cycle of each operation type based on the ratio of the computing load of each operation type to the configured computing power of each operation type; Take the computing cycle of each operation type as the computing overhead of each operation type.

[0010] In some embodiments, determining the computing overhead of the current operator in the target chip based on the computing overhead of each operation type includes: In the case where the instructions of each operation type use the same instruction issue unit, accumulate the computing overhead of each operation type to obtain the computing overhead of the current operator in the target chip.

[0011] In some embodiments, determining the computing overhead of the current operator in the target chip based on the computing overhead of each operation type includes: When instructions of each operation type use different instruction issue units, the maximum calculation overhead of each operation type is determined as the calculation overhead of the current operator in the target chip.

[0012] In some embodiments, determining the number of instruction executions of each operation type based on the instruction execution information of the current operator in the target chip includes: Executing the current operator in the target chip or executing the current operator in an instruction simulator of the target chip to obtain the instruction execution information; Counting the number of executions of instructions of each operation type in the instruction execution information to determine the number of instruction executions of each operation type.

[0013] In some embodiments, the operation types include floating-point operation types, integer operation types, special function operation types, and general operation types; Instructions in the general operation type do not belong to floating-point operation instructions, integer operation instructions, and special function operation instructions.

[0014] In some embodiments, the method further includes: Adjusting the configured computing power of each operation type in the target chip to obtain the updated configured computing power of each operation type in the target chip; Based on the calculation load of each operation type and the updated configured computing power of each operation type in the target chip, determining the updated calculation overhead value of each operation type; Based on the updated calculation overhead value of each operation type, determining the updated calculation overhead value of the current operator in the target chip; Based on the calculation overhead of the current operator before the configured computing power adjustment and the updated calculation overhead value, determining the performance improvement value of the target chip for executing the current operator.

[0015] The present invention provides a vector core calculation overhead quantification device, including: An instruction statistics module, configured to determine the number of instruction executions of each operation type based on the instruction execution information of the current operator in the target chip; the target chip includes a vector core; the vector core executes the current operator; A load determination module, configured to determine the calculation load of each operation type based on the number of instruction executions of each operation type; An overhead calculation module, configured to determine the calculation overhead of each operation type based on the calculation load of each operation type and the configured computing power of each operation type in the target chip; An overhead summary module, configured to determine the calculation overhead of the current operator in the target chip based on the calculation overhead of each operation type.

[0016] The present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the vector core computation overhead quantification method as described above is implemented.

[0017] The present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the vector core computation overhead quantification method as described above is implemented.

[0018] The vector core computation overhead quantification method, device, electronic device, and storage medium provided by the present invention determine the instruction execution times of each operation type based on the instruction execution information of the current operator in the target chip; the target chip includes a vector core; the vector core executes the current operator; determine the computation load of each operation type based on the instruction execution times of each operation type; determine the computation overhead of each operation type based on the computation load of each operation type and the configured computing power of each operation type in the target chip; determine the computation overhead of the current operator in the target chip based on the computation overhead of each operation type; since during the process of the vector core of the target chip executing the current operator, the instruction execution times, computation load, and computation overhead of each operation type are quantitatively analyzed, the computation overhead of each operation type in the overall computing task can be determined, and finally the computation overhead of the current operator in the target chip can be obtained, realizing accurate quantification of the computation overhead of the vector core when processing different operators, which is beneficial to comprehensively and accurately evaluating the computation overhead of a general-purpose graphics processor when executing various operators, and providing a scientific basis for the architecture design and operator optimization of the general-purpose graphics processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.

[0020] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0021] Figure 1 It is a schematic diagram of the architecture of the general-purpose graphics processor provided by the present invention.

[0022] Figure 2 It is one of the flow diagrams of the vector core computation overhead quantification method provided by the present invention.

[0023] Figure 3 It is the second flow schematic diagram of the vector core calculation overhead quantification method provided by the present invention.

[0024] Figure 4 It is the structural schematic diagram of the vector core calculation overhead quantification device provided by the present invention.

[0025] Figure 5 It is the structural schematic diagram of the electronic device provided by the present invention. Specific Embodiments

[0026] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0027] It should be noted that the terms "first", "second", etc. in the present invention are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units or modules does not necessarily have to be limited to those steps or units or modules clearly listed, but may include other steps or units or modules not clearly listed or inherent to these processes, methods, products or devices.

[0028] ‌ Figure 1 It is the architecture schematic diagram of the general graphics processor provided by the present invention. As Figure 1 shown, the general graphics processor 100 may include a plurality of streaming processor clusters 110 (Streaming Processor Cluster, SPC) and a global memory 120 that can be accessed by all processors. The computing units in each streaming processor cluster 110 can generally be divided into tensor cores 111 (Tensor Core) and vector cores 112 (Vector Core). The tensor cores are used to perform tensor calculations. The vector cores are used to perform general vector calculations.

[0029] Various operators can be executed in a general-purpose graphics processing unit (GPGPU). An operator represents an entity that performs a specific operation or function. It can be a basic component that performs functions such as mathematical operations, data transformation, or logical judgment. In a neural network, an operator refers to the specific implementation of various layers or operations, which are the basic units that make up the neural network. Each operator performs a specific mathematical operation or data transformation, and they work together to achieve processes such as forward propagation and backward propagation of the neural network.

[0030] When evaluating the computational overhead of operators executed by a GPGPU, related technologies overly focus on the number of floating-point operations per second (FLOPS) alone. For example, a matrix multiplication operator of size (depth width height) is simplified to FLOPS (each multiply-accumulate operation is counted as 2 floating-point operations). This method can only evaluate the computational overhead of tensor cores and cannot effectively quantify and evaluate the computational overhead of vector cores, ignoring the complexity and diversity of vector cores as general computing units.

[0031] The main challenge in accurately evaluating the computational overhead of vector cores is that their instruction types are very rich. As general computing units, vector cores generally include instructions such as vector floating-point calculations, vector integer calculations, vector special function calculations, vector logical calculations, scalar calculations, data access, conditional control, and synchronization processing. Different algorithm tasks utilize vector core resources in different ways, making it difficult to establish a unified performance quantification model.

[0032] To solve the above technical problems, Figure 2 is one of the schematic flowcharts of the method for quantifying the computational overhead of vector cores provided by the present invention. As Figure 2 shown, the method includes steps 210, 220, 230, and 240.

[0033] Step 210: Based on the instruction execution information of the current operator in the target chip, determine the number of executed instructions for each operation type; the target chip includes a vector core; the vector core executes the current operator.

[0034] Specifically, the execution subject of the method for quantifying the computational overhead of vector cores provided by the embodiments of the present invention is a device for quantifying the computational overhead of vector cores. This device can be implemented by software, for example, implemented by a program for quantifying the computational overhead of vector cores running on a computer; it can also be implemented by hardware, such as a processor, chip, computer, or server that executes the method for quantifying the computational overhead of vector cores.

[0035] The application scenario of the method provided by the embodiment of the present invention is to run a specific operator in a general-purpose graphics processing unit and quantitatively calculate the computing overhead of the operator in the vector core of the general-purpose graphics processing unit.

[0036] The current operator can be the calculation of activation functions in neural networks, the normalization operation in the attention mechanism, and various custom operators, etc. The embodiment of the present invention does not make specific limitations on the specific type of the operator.

[0037] The target chip is a general-purpose graphics processing unit that needs to perform quantitative calculation of the computing overhead of the vector core. The target chip includes a vector core, and the current operator is executed by the vector core.

[0038] Instruction execution information refers to various data and details related to instruction execution when running the current operator on the target chip, which can include instruction type, execution times, clock cycles, and execution units, etc. According to the instruction execution information, the execution times of instructions of each operation type can be statistically obtained.

[0039] On the one hand, the current operator can be run on the target chip to obtain instruction execution information; on the other hand, the current operator can also be executed in the instruction simulator of the target chip to obtain instruction execution information. The execution times of instructions of each operation type in the instruction execution information are statistically counted to determine the execution times of instructions of each operation type.

[0040] The operation type refers to the types of operation operations that the vector core can execute. According to the instruction set of the vector core, the operation types of the instructions executed by the vector core can be divided into the following four categories: I. Floating-point operation type, including vector floating-point calculations (addition, subtraction, multiplication, division, square root, etc.); II. Integer operation type, including vector integer calculations (bit operations, addition, subtraction, multiplication, division, etc.); III. Special function operation type, including vector special function calculations (exponential, logarithmic, trigonometric functions, etc.); IV. General operation type, the instructions in the general operation type do not belong to floating-point operation instructions, integer operation instructions, and special function operation instructions, including vector logical operations (comparison, selection, masking, etc.), scalar calculations (various operations related to control flow), memory access operations (read and write operations of global memory, shared memory, texture memory, etc.), and thread control and synchronization operations (branch prediction, thread synchronization, etc.).

[0041] Step 220: Determine the computing load of each operation type based on the execution times of instructions of each operation type.

[0042] Specifically, the computational workload refers to the computational effort borne by the vector cores for each operation type when executing the current operator, reflecting the relative importance of that operation type in the overall computational task and the demand for computational resources.

[0043] Instructions of different operation types have different complexities. For example, floating-point multiply-accumulate operations are generally more complex than simple integer addition operations. The computational effort of a floating-point operation instruction is also greater than that of an integer addition instruction. That is to say, the computational efforts of instructions of different operation types are also different.

[0044] Therefore, by analyzing the execution process of the current operator, count the number of executed instructions for each operation type. Based on the number of executed instructions for each operation type and the computational effort of the instructions for each operation type, the computational workload of each operation type can be determined.

[0045] Step 230: Based on the computational workload of each operation type and the configured computing power of each operation type in the target chip, determine the computational overhead of each operation type.

[0046] Specifically, the computational overhead refers to the resources consumed when executing the current operator, including time, computational resources, etc. In the embodiments of the present invention, the computational overhead can be measured by time, that is, measured by the number of computational cycles (cycles) used by the vector cores to process each operation type. A computational cycle refers to the number of clock cycles required to complete a computational task or execute an instruction. It reflects the speed and efficiency of the target chip in executing computations. In chip performance evaluation, the computational cycle is an important indicator, which is closely related to factors such as the chip's clock frequency, instruction set architecture, and hardware design. The more computational cycles, the greater the computational overhead.

[0047] The configured computing power refers to the computing capabilities designed and provided by the target chip for different operation types, which is related to hardware resources (such as the number of computing units, clock frequency, etc.). It represents the maximum computational potential of the target chip for a specific operation type. By querying the hardware specification manual (specification, spec) of the target chip, obtain the configured computing power of each operation type in the target chip.

[0048] Based on the computational workload of each operation type and the configured computing power of each operation type in the target chip, determine the computational overhead of each operation type. For example, for a certain operation type, the ratio of the computational workload to the corresponding configured computing power can be used as the computational overhead of that operation type.

[0049] Step 240: Based on the computational overhead of each operation type, determine the computational overhead of the current operator in the target chip.

[0050] Specifically, by comprehensively calculating the computational costs of each operation type, the computational cost of the current operator in the target chip can be obtained.

[0051] The method for quantifying the computational cost of a vector core provided by the embodiments of the present invention determines the number of instruction executions of each operation type based on the instruction execution information of the current operator in the target chip; the target chip includes a vector core; the vector core executes the current operator; determines the computational load of each operation type based on the number of instruction executions of each operation type; determines the computational cost of each operation type based on the computational load of each operation type and the configured computing power of each operation type in the target chip; determines the computational cost of the current operator in the target chip based on the computational cost of each operation type; since during the process of the vector core of the target chip executing the current operator, the number of instruction executions, computational load, and computational cost of each operation type are quantitatively analyzed, the computational cost of each operation type in the overall computational task can be determined, and finally the computational cost of the current operator in the target chip can be obtained, realizing the accurate quantification of the computational cost of the vector core when processing different operators, which is beneficial to comprehensively and accurately evaluating the computational cost of a general-purpose graphics processor executing various operators, and provides a scientific basis for the architecture design and operator optimization of the general-purpose graphics processor.

[0052] It should be noted that each embodiment of the present invention can be freely combined, the order can be swapped, or each can be executed independently, and does not need to rely on or depend on a fixed execution order.

[0053] In some embodiments, determining the computational load of each operation type based on the number of instruction executions of each operation type includes: Determining the computational load of each operation type based on the number of instruction executions of each operation type and the warp size for executing each operation type in the target chip.

[0054] Specifically, in a target chip (general-purpose graphics processor), a warp is the basic execution unit. After a thread block is scheduled to a streaming multiprocessor cluster, the threads in the thread block will be further divided into multiple warps. The warp size refers to the number of threads included in a warp. For example, if the warp size is 32, it means that each warp consists of 32 consecutive threads. All threads within a warp execute in a lockstep manner, that is, they execute the same instruction at the same time, but can operate on different data.

[0055] Therefore, the computational load of each operation type can be determined according to the number of instruction executions of each operation type and the warp size for executing each operation type in the target chip. Specifically, the number of instruction executions of each operation type can be multiplied by the warp size to obtain the computational load of each operation type.

[0056] The vector kernel computation overhead quantification method provided by the embodiments of the present invention can accurately calculate the computation loads of various operation types through the number of executed instructions of each operation type and the warp size for executing each operation type.

[0057] In some embodiments, before determining the computation loads of various operation types based on the number of executed instructions of each operation type and the warp size for executing each operation type in the target chip, the method further includes: When the current operation type includes multiple operation subtypes, determining the operation subtype with equivalent computing power among the multiple operation subtypes; Based on the operands of each operation subtype and the configured computing power of each operation subtype in the target chip, determining the quantity conversion relationship between the instructions of each operation subtype and the instructions of the operation subtype with equivalent computing power when the computing power is equivalent; Based on the quantity conversion relationship and the number of executed instructions of each operation subtype, determining the number of executed instructions of the operation subtype with equivalent computing power; Determining the number of executed instructions of the operation subtype with equivalent computing power as the number of executed instructions of the current operation type.

[0058] Specifically, after classifying the instructions according to the operation type, there will be a situation where the instructions have a finer classification. For example, for the floating-point operation type, it can include many operation subtypes (such as addition, multiplication, multiply-accumulation, etc.), and the operand data types of the instructions of each operation subtype (such as 32-bit floating point, 16-bit floating point, etc.) are different, and further classification can be carried out.

[0059] The classified instructions can be processed by using the method of computing power normalization, so as to establish an accurate mapping relationship between the instructions and the hardware processing capabilities. The hardware processing capabilities are specifically the configured computing power in the hardware specification manual.

[0060] Taking the current operation type as the floating-point operation type as an example. The current operation type includes multiple operation subtypes, such as 32-bit floating point (32 bit Float Point, FP32), 16-bit brain floating point (16 bit Brain FloatPoint, BF16), 16-bit floating point (16 bit Float Point, FP16), etc.

[0061] When the current operation type includes multiple operation subtypes, determining the operation subtype with equivalent computing power among the multiple operation subtypes. The operation subtype with equivalent computing power refers to the operation subtype that can represent other operation subtypes in terms of computing power and performance, or can be converted and equivalent to other operation subtypes. For example, the FP32 operation type in the floating-point operation type can be used as the operation subtype with equivalent computing power.

[0062] Determine the quantity conversion relationship between the instructions of each operator type and the instructions of the operator type equivalent in computing power when they are equivalent in computing power, based on the operands of each operator type and the configured computing power of each operator type in the target chip.

[0063] For example, for floating-point operation types such as FP32, BF16, and FP16, the instructions of these operation types can be normalized to the FP32 operation type. Generally, the ratio of the number of operands processed by each warp instruction in a general-purpose graphics processing unit is the same as the ratio of the computing power indicators of the corresponding type in the hardware specification (SPEC). For example, in a certain chip, the number of operands processed by each FP32 / BF16 / FP16 floating-point multiply-accumulate instruction is 64 / 128 / 128, and the publicly disclosed FP32 / BF16 / FP16 floating-point performance indicators in the hardware specification of this chip are 66.9 / 133.8 / 133.8 teraflops (TFLOPS), and the ratio is consistent at 1:2:2. Therefore, one BF16 operation instruction and one FP16 operation instruction can be respectively equivalent to one FP32 operation instruction. It can also be considered that in this chip, the quantity conversion relationship between FP32, BF16, and FP16 operation instructions when they are equivalent in computing power is 1:1:1.

[0064] According to the quantity conversion relationship and the number of executions of the instructions of each operator type, determine the number of executions of the instructions of the operator type equivalent in computing power, and finally determine the number of executions of the instructions of the operator type equivalent in computing power as the number of executions of the instructions of the current operation type. For example, for one BF16 operation instruction, one FP16 operation instruction, and one FP32 operation instruction, the operator type equivalent in computing power is FP32, and the quantity conversion relationship is BF16:FP16:FP32 = 1:1:1. Then the above three instructions of different operator types can be converted into three FP32 operation instructions.

[0065] Similarly, during computing power evaluation, each instruction of the integer operation type that processes different data types can be normalized and counted as one 32-bit integer (INT32) accumulate instruction.

[0066] Special function instructions generally include reciprocal, reciprocal square root, sine function, cosine function, exponential, and logarithm, etc. These instructions use the special function unit (SFU) of the general-purpose graphics processing unit for calculation and have the same performance limit throughput. Therefore, each instruction of the special function operation type can be normalized to one general FP32 SFU instruction.

[0067] Instructions of the general operation type can be normalized to one non-computation instruction.

[0068] The vector core computing overhead quantification method provided by the embodiments of the present invention, in the case that the current operation type includes multiple operation sub-types, determines the computing power equivalent operation sub-type among the multiple operation sub-types, and normalizes the instructions of the multiple operation sub-types according to the computing power equivalent operation sub-type, so as to establish an accurate mapping relationship between the instructions and the hardware processing capabilities, which is beneficial to accurately calculating the computing loads of each operation type.

[0069] In some embodiments, based on the computing loads of each operation type and the configured computing power of each operation type in the target chip, determining the computing overhead of each operation type includes: Based on the ratio of the computing load of each operation type to the configured computing power of each operation type, determining the computing cycles of each operation type; Taking the computing cycles of each operation type as the computing overhead of each operation type.

[0070] Specifically, taking the floating-point operation type as an example for illustration.

[0071] When a certain chip executes a certain operator, the computing load of the floating-point operation type is 106.8 billion floating-point operations (GigaFloating Point Operations, GFLOP), and the configured computing power of the floating-point operation type of this chip is 33.8 kFLOP / cycle. Then the computing cycles of the floating-point operation type are: 106.8 GFLOP / 33.8 kFLOP / cycle = 3.16 M cycles.

[0072] Among them, M is mega; k is kilo; FLOP is the number of floating-point operations (Floating Point Operations).

[0073] Taking 3.16 M cycles as the computing overhead of the floating-point operation type when this chip executes this operator.

[0074] The vector core computing overhead quantification method provided by the embodiments of the present invention determines the computing overhead of each operation type according to the computing load of each operation type and the configured computing power of each operation type in the target chip, and realizes the quantification of the computing overhead of each operation type in the overall computing task.

[0075] In some embodiments, based on the computing overhead of each operation type, determining the computing overhead of the current operator in the target chip includes: In the case that the instructions of each operation type use the same instruction issue unit, adding up the computing overhead of each operation type to obtain the computing overhead of the current operator in the target chip.

[0076] Specifically, the Instruction Issue Unit is a key component responsible for reading instructions from the instruction cache and sending them to the execution unit.

[0077] If instructions of various operation types use the same instruction issue unit, it means that instructions of different operation types are scheduled and distributed through a single instruction issue unit. For example, instructions for floating-point operations, integer operations, and special function operations are all managed by the same instruction issue unit. These instructions of different operation types share the same instruction issue unit and they will compete for the issue resources. At a certain moment, if a certain type of instruction occupies the instruction issue unit, other types of instructions can only wait.

[0078] Therefore, it is necessary to accumulate the computational overheads of various operation types and determine the total computational overhead as the computational overhead of the current operator in the target chip.

[0079] In some embodiments, determining the computational overhead of the current operator in the target chip based on the computational overheads of various operation types includes: When instructions of various operation types use different instruction issue units, determining the maximum value of the computational overheads of various operation types as the computational overhead of the current operator in the target chip.

[0080] Specifically, if instructions of various operation types use different instruction issue units, it means that instructions of different operation types each use a dedicated instruction issue unit. For example, floating-point operation instructions are processed by a floating-point instruction issue unit, integer operation instructions are processed by an integer instruction issue unit, and special function operation instructions are processed by a special function instruction issue unit. Different types of instructions are independently processed in their respective instruction issue units, avoiding mutual competition and improving the overall issue efficiency.

[0081] Therefore, the maximum value of the computational overheads of various operation types can be determined as the computational overhead of the current operator in the target chip.

[0082] The method for quantifying the computational overhead of the vector kernel provided by the embodiments of the present invention determines the computational overhead of the current operator in the target chip according to whether the instruction issue units used by instructions of various operation types are the same, improving the accuracy of determining the computational overhead.

[0083] Figure 3 It is the second schematic flowchart of the method for quantifying the computational overhead of the vector kernel provided by the present invention. As Figure 3 shown, the method includes: Step 310, by executing the current operator on a certain chip (general-purpose graphics processor) or an instruction simulator, obtaining instruction execution information and counting the number of executions of instructions of various operation types.

[0084] For example, the current operator can be the softmax operator.

[0085] Step 320: Classify the instructions according to the operation type.

[0086] For example, the operation types can include floating-point operation type (float op inst), integer operation type (int op inst), special function operation type (sfu op inst), and general operation type (general op inst).

[0087] Step 330: Normalize the instructions of each operation type to determine the computational load (workload) of each operation type.

[0088] For example, the number of integer operation type instructions Int_instruction_count = 705 M; the number of floating-point operation type instructions float_instruction_count = 1669 M; the number of special function operation type instructions sfu_instruction_count = 201 M; the number of general operation type instructions general_instruction_count = 1784 M.

[0089] Since each instruction of a certain chip is in units of warps and two floating-point operation operations are calculated in one multiply-accumulate operation. Each FP32 floating-point multiply-accumulate instruction can be counted as 2 warp_size floating-point operation operations (float op, flop). Each integer instruction performs 2 warp_size integer operation operations (int op). And each special function operation type instruction can be recorded as warp_size special function operation operations (sfu op); each general operation type instruction is recorded as warp_size general operations (general op).

[0090] The warp size of a certain chip is 32. Correspondingly, the computational loads of the softmax operator in each operation type are respectively: Int_workload = 705 M 2 32 = 45.1 G int op; float_workload = 1669 M 2 32 = 106.8 G float op; sfu_workload = 201 M 32 = 6.4 G sfu operations; general_workload = 1784 M 32 = 57.1 G general operations; Among them, G stands for gigabit.

[0091] Step 340: Determine the configured computing power according to the hardware specification manual of a certain chip, and convert the floating-point operation instructions, integer operation instructions, special function operation instructions, and general operation instructions that can be completed by each clock cycle of the chip.

[0092] For example, through the hardware specification manual of the chip, the configured computing power can be determined as follows: The computing power configuration for floating-point operations is: 33.8k float op / cycle; The computing power configuration for integer operations is: 16.9k int op / cycle; The computing power configuration for special function operations is: 2.11k sfu op / cycle; The computing power configuration for general operations is: 16.9k general op / cycle; Among them, fop is the number of floating-point operation times (FLOP); int op is the number of integer operation times; sfu op is the number of special function operation times; general op is the number of general operation times.

[0093] Determine the calculation cycles of each operation type and use them as the calculation overhead.

[0094] For example, according to the above configured computing power, it can be calculated that: The calculation cycle of the integer operation type Int_cycles = 45.1 G int op / 16.9k int op / cycle = 2.67 M cycles; The calculation cycle of the floating-point operation type float_cycles = 106.8 G float op / 33.8k fop / cycle = 3.16 M cycles; The calculation cycle of the special function operation type sfu_cycles = 6.4 G sfu op / 2.11k sfu op / cycle = 3.03 M cycles; The calculation cycle of the general operation type, general_cycles = 57.1 G general op / 16.9k general op / cycle = 3.38 M cycles.

[0095] Step 350: Determine whether to share the instruction issue unit and determine the calculation overhead of the current operator in the chip.

[0096] Instructions of each operation type share one instruction issue unit. Sum up the calculation overheads of each operation type to obtain the total calculation overhead.

[0097] Vector_core_cycles = Int_cycles + float_cycles + sfu_cycles + general_cycles = 12.24 M cycles.

[0098] When actually running the softmax operator on this chip, it is detected that the operating frequency of the chip is 1.59 gigahertz (Ghz), and the actual time consumption is 8.36 milliseconds (ms). It can be analyzed that the actual calculation cycle of the softmax operator is 13.29 Mcycles. It can be analyzed that the utilization rate of the vector core of the softmax operator on this chip is 0.92, which is close to the full utilization state of the vector core computing power.

[0099] By calculating the percentage of the classification cycle to the total calculation cycle, such as (general_cycles + Int_cycles) / Vector_core_cycles = 49.4%, it can be known that the integer addition and general instructions of this operator occupy nearly half of the computing power of the vector core. The operator implementation and compiler optimization can be checked, and reducing this part of the instruction overhead can significantly improve the performance of this operator.

[0100] In some embodiments, the method further includes: Adjust the configured computing power of each operation type in the target chip to obtain the updated configured computing power of each operation type in the target chip; Based on the calculation load of each operation type and the updated configured computing power of each operation type in the target chip, determine the updated value of the calculation overhead of each operation type; Based on the updated value of the calculation overhead of each operation type, determine the updated value of the calculation overhead of the current operator in the target chip; Based on the calculation overhead of the current operator before the adjustment of the configured computing power and the updated value of the calculation overhead, determine the performance improvement value of the target chip for executing the current operator.

[0101] Specifically, the method provided by the embodiments of the present invention can also be used to improve the performance design of the target chip.

[0102] In the design of the target chip, the configured computing power of each operation type in the target chip can be adjusted to obtain the updated configured computing power of each operation type in the target chip. For example, in the above embodiment, the configured computing power of floating-point operations, integer operations, special function operations, and general operations remains unchanged, and the configured computing power of special function operations is doubled, that is, increased from 2.11k sfu op / cycle to 4.22k sfu op / cycle.

[0103] According to the computational load of each operation type and the updated configured computing power of each operation type in the target chip, determine the updated value of the computational overhead of each operation type. For example, the updated value of the computational overhead new sfu_cycles = 6.4 Gsfu op / 4.22k sfu op / cycle = 1.52 M cycles.

[0104] According to the updated value of the computational overhead of each operation type, determine the updated value of the computational overhead of the current operator in the target chip. For example, the updated value of the computational overhead new Vector_core_cycles = Int_cycles+ float_cycles+newsfu_cycles+ general_cycles = 10.73 M cycles.

[0105] According to the computational overhead of the current operator before the adjustment of the configured computing power and the updated value of the computational overhead, determine the performance improvement value of the target chip for executing the current operator. For example, the performance improvement value is 1 - new Vector_core_cycles / oldVector_core_cycles. It can be seen that there is a maximum performance improvement of 12.3%, and the proportion of the special function calculation time (sfu_cycles / Vector_core_cycles) will drop from 25.6% to 14.2%.

[0106] The method for quantifying the computational overhead of vector cores provided by the embodiments of the present invention can accurately calculate the performance improvement value of the target chip for executing the current operator, providing a scientific basis for the architecture design and operator optimization of general-purpose graphics processors.

[0107] In some embodiments, if the instructions of the integer operation type in the target chip are changed to be capable of parallel emission with the instructions of other operation types, the total computational overhead of the vector core is: Vector_core_cycles = max(Int_cycles, float_cycles + sfu_cycles + general_cycles) = 9.57 M cycles。

[0108] The performance improvement of the microarchitecture improvement of the chip integer computing unit's instruction parallel emission can be quantitatively known. It can be known from 1 - new Vector_core_cycles / old Vector_core_cycles that there is a performance improvement of up to 12.3% - 21.8%.

[0109] The device provided by the embodiments of the present invention will be described below. The device described below can be correspondingly referred to the method described above.

[0110] Figure 4 It is a schematic structural diagram of the vector core computing overhead quantification device provided by the present invention, as Figure 4 shown. The device includes: An instruction statistics module 410, configured to determine the number of executed instructions of each operation type based on the instruction execution information of the current operator in the target chip; the target chip includes a vector core; the vector core executes the current operator; A load determination module 420, configured to determine the computing load of each operation type based on the number of executed instructions of each operation type; An overhead calculation module 430, configured to determine the computing overhead of each operation type based on the computing load of each operation type and the configured computing power of each operation type in the target chip; An overhead summary module 440, configured to determine the computing overhead of the current operator in the target chip based on the computing overhead of each operation type.

[0111] The vector core computing overhead quantification device provided by the embodiment of the present invention determines the instruction execution times of each operation type based on the instruction execution information of the current operator in the target chip; the target chip includes a vector core; the vector core executes the current operator; determines the computing load of each operation type based on the instruction execution times of each operation type; determines the computing overhead of each operation type based on the computing load of each operation type and the configured computing power of each operation type in the target chip; determines the computing overhead of the current operator in the target chip based on the computing overhead of each operation type; since during the process of the vector core of the target chip executing the current operator, the instruction execution times, computing load, and computing overhead of each operation type are quantitatively analyzed, the computing overhead of each operation type in the overall computing task can be determined, and finally the computing overhead of the current operator in the target chip can be obtained, realizing the accurate quantification of the computing overhead of the vector core when processing different operators, which is beneficial to comprehensively and accurately evaluating the computing overhead of a general-purpose graphics processor when executing various operators, and provides a scientific basis for the general-purpose graphics processor architecture design and operator optimization.

[0112] Figure 5 is a schematic structural diagram of the electronic device provided by the present invention, as Figure 5 shown, the electronic device may include: a processor (Processor) 510, a communication interface (Communications Interface) 520, a memory (Memory) 530, and a communication bus (Communications Bus) 540. Among them, the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call the logical commands in the memory 530 to execute the methods described in the above embodiments, for example: Based on the instruction execution information of the current operator in the target chip, determine the instruction execution times of each operation type; the target chip includes a vector core; the vector core executes the current operator; determine the computing load of each operation type based on the instruction execution times of each operation type; determine the computing overhead of each operation type based on the computing load of each operation type and the configured computing power of each operation type in the target chip; determine the computing overhead of the current operator in the target chip based on the computing overhead of each operation type.

[0113] In addition, when the logical commands in the above-mentioned memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several commands for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0114] The processor in the electronic device provided in the embodiments of the present invention can call the logical instructions in the memory to implement the above method. The specific implementation manner is the same as that of the foregoing method embodiment, and the same beneficial effects can be achieved, which will not be elaborated herein.

[0115] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to execute the methods provided in the above-mentioned embodiments.

[0116] The specific implementation manner is the same as that of the foregoing method embodiment, and the same beneficial effects can be achieved, which will not be elaborated herein.

[0117] The embodiments of the present invention provide a computer program product, including a computer program, which when executed by a processor, implements the method as described above.

[0118] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0119] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for quantifying vector kernel computation overhead, characterized in that: include: Determining the number of instruction executions of each operation type based on instruction execution information of the current operator in a target chip; the target chip includes a vector core; The vector core executes the current operator; the operation type is the type of operation performed by the vector core; Determining the computational load of each operation type based on the number of instruction executions of each operation type; Determine the computational overhead of each computational type based on the computational load of each computational type and the configured computing power of each computational type in the target chip; Based on the computational overhead of each operation type, the computational overhead of the current operator in the target chip is determined.

2. The vector core computation overhead quantification method according to claim 1, characterized in that: The determining the computing load of each operation type based on the number of instruction executions of each operation type includes: The computation load of each operation type is determined based on the number of instruction executions of each operation type and the size of the thread warps executing each operation type in the target chip.

3. The vector core computation overhead quantification method according to claim 2, characterized in that: Before determining the computation load of each operation type based on the number of instruction executions of each operation type and the size of the thread warps executing each operation type in the target chip, the method further includes: In the case where the current operation type includes multiple operator types, determining an operator type with equivalent computing power among the multiple operator types; Based on the operands of each operator type and the configured computing power of each operator type in the target chip, determining the quantity conversion relationship between the instructions of each operator type and the instructions of the operator type with equivalent computing power when the computing power is equivalent; Based on the quantity conversion relationship and the number of instruction executions of each operator type, determine the number of instruction executions of the computing power equivalent operator type; The number of instruction executions of the computing power equivalent operator type is determined as the number of instruction executions of the current operation type.

4. The vector core computation overhead quantification method according to claim 1, characterized in that: The determining the computing overhead of each computing type based on the computing load of each computing type and the configured computing power of each computing type in the target chip includes: Determine the computing cycle of each computing type based on the ratio of the computing load of each computing type to the configured computing power of each computing type; The computation cycle of each operation type is regarded as the computation overhead of each operation type.

5. The vector core computation overhead quantification method according to claim 1, characterized in that: The determining, based on the computational overhead of each operation type, the computational overhead of the current operator in the target chip includes: When instructions of various operation types use the same instruction issuing unit, the computational overhead of each operation type is accumulated to obtain the computational overhead of the current operator in the target chip.

6. The vector core computation overhead quantification method according to claim 1, characterized in that: The determining, based on the computational overhead of each operation type, the computational overhead of the current operator in the target chip includes: In the case where instructions of different operation types use different instruction issuing units, the maximum value of the computational overhead of each operation type is determined as the computational overhead of the current operator in the target chip.

7. The vector core computation overhead quantification method according to claim 1, characterized in that: The step of determining the number of instruction executions of each operation type based on the instruction execution information of the current operator in the target chip includes: Executing the current operator in the target chip, or executing the current operator in an instruction simulator of the target chip, to obtain the instruction execution information; The execution times of instructions of each operation type in the instruction execution information are counted to determine the execution times of instructions of each operation type.

8. The vector core computation overhead quantification method according to claim 1, characterized in that: The operation types include floating point operation types, integer operation types, special function operation types and general operation types.

9. The vector core computation overhead quantification method according to any one of claims 1 to 8, characterized in that: The method further comprises: Adjusting the configured computing power of each computing type in the target chip to obtain the configured updated computing power of each computing type in the target chip; Determine a computing overhead update value for each computing type based on a computing load of each computing type and a configuration update computing power of each computing type in the target chip; Determining a computational cost update value of the current operator in the target chip based on computational cost update values ​​of each operation type; Based on the computational overhead of the current operator before the configuration computing power is adjusted and the computational overhead update value, a performance improvement value of the target chip executing the current operator is determined.

10. A vector core computation overhead quantization device, characterized in that: include: An instruction statistics module, used to determine the number of instruction executions of each operation type based on instruction execution information of the current operator in a target chip; the target chip includes a vector core; The vector core executes the current operator; the operation type is the type of operation performed by the vector core; A load determination module, for determining a computational load of each operation type based on the number of instruction executions of each operation type; An overhead calculation module, used to determine the computation overhead of each operation type based on the computation load of each operation type and the configured computing power of each operation type in the target chip; The cost summary module is used to determine the computation cost of the current operator in the target chip based on the computation cost of each operation type.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the vector core computing overhead quantification method described in any one of claims 1 to 9 is implemented.

12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the vector core computing overhead quantification method described in any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and storage medium

    CN112036561A

  • Task load calculation method and device, storage medium and terminal

    CN113626200A

  • Core particle-oriented neural network inference overhead estimation method and device, and electronic equipment

    CN115186821A

  • Operation acceleration processing method, operation accelerator using method and operation accelerator

    CN116685964A

  • Computational graph processing method based on federal learning, computer equipment and storage medium

    CN117114091A

Cited By

  • Multi-architecture compilation intermediate representation generation method and system driven by heterogeneous computing power fusion

    CN120849133A