Method and related product for estimating time required to execute a neural network model
By estimating the time required by the artificial intelligence processor to execute neural network models, the problem of inefficient splitting strategy selection in multi-core processing systems is solved, and the system operation performance is improved.
Patent Information
- Application Number
- CN202111680762.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-12-31
AI Technical Summary
In multi-core processing systems, traditional data parallel solutions are difficult to meet the requirements of small data and low latency for accelerators in inference scenarios, resulting in inefficient selection of splitting strategies.
By obtaining the hardware parameters of the neural network model and the artificial intelligence processor, processing and generating binary instructions, the time required to execute the neural network model is estimated.
This method can effectively estimate the running time of the neural network model, provide a reasonable performance basis, and improve the efficiency of splitting strategy selection and system operation performance.
Smart Images

Figure CN114356738B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of neural networks, and particularly to a method and related products for estimating the time required for an artificial intelligence processor to execute a neural network model. Background Art
[0002] In recent years, deep learning accelerators have been continuously proposed and, like general-purpose processors, are expanding from single-core to multi-core. Such an expanded multi-core structure can support data parallelism during the training phase to improve data throughput and accelerate the training speed. However, during the inference phase, compared to throughput, deep neural networks have higher requirements for end-to-end latency, which often determines the usability of the accelerator in a certain scenario. Traditional data parallelism schemes cannot meet the requirements of small data and low latency for the accelerator in the inference scenario.
[0003] In a multi-core processing system, the input data of a network can be split into different scales according to the number of available cores in the system and calculated on different cores; different data splitting strategies may have different performance. Therefore, to solve the problem of selecting a splitting strategy during the multi-core splitting operation on a multi-core processor, the selection of the splitting strategy needs to estimate the time required to execute the neural network model. Therefore, it is urgent to solve this time estimation problem. Summary of the Invention
[0004] To solve this technical problem, embodiments of this disclosure provide a method and related products for estimating the time required for an artificial intelligence processor to execute a neural network model.
[0005] In a first aspect, a method for estimating the time required for an artificial intelligence processor to execute a neural network model is provided. The estimation method includes the following steps:
[0006] Obtain the neural network model and the hardware parameters of the artificial intelligence processor;
[0007] Process the neural network model according to the neural network model and the hardware parameters of the artificial intelligence processor to obtain corresponding binary instructions;
[0008] Obtain the time required to execute the neural network model according to the hardware binary instructions corresponding to the neural network model.
[0009] In a second aspect, a device for estimating the time required for an artificial intelligence processor to execute a neural network model is provided. The device includes:
[0010] A parameter determination unit for obtaining the neural network model and the hardware parameters of the artificial intelligence processor;
[0011] A processing unit, configured to process the neural network model according to the neural network model and the hardware parameters of the artificial intelligence processor to obtain corresponding binary instructions;
[0012] An estimation unit, configured to obtain the time required to execute the neural network model according to the hardware binary instructions corresponding to the neural network model.
[0013] In a third aspect, a computing chip is provided, where the computing chip includes the device as described above.
[0014] In a fourth aspect, a computer-readable storage medium is provided, storing a computer program for electronic data exchange, where the computer program causes a computer to execute the method as described above.
[0015] The technical solution provided in this application performs time estimation, estimates the running time of the split neural network model, so as to provide a reasonable performance basis for the split strategy selection stage, and can greatly improve the efficiency of split strategy selection and the running performance of the system. Description of the Drawings
[0016] Figure 1 is a schematic diagram of an artificial intelligence processor architecture model;
[0017] Figure 2 is a flowchart of a method for estimating the time required for an artificial intelligence processor to execute a neural network model provided in this application
[0018] Figure 3 is a schematic diagram of an instruction queue;
[0019] Figure 4 is a schematic diagram of a device for estimating the time required for an artificial intelligence processor to execute a neural network model provided in this application. Detailed Embodiments
[0020] In order to enable those skilled in the art to better understand the disclosure solution, the technical solutions in the disclosure embodiments will be clearly and completely described below with reference to the drawings in the disclosure embodiments. Obviously, the described embodiments are only a part of the disclosure embodiments, rather than all of the embodiments. Based on the embodiments in the disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the disclosure.
[0021] Such as Figure 1As shown, it is a schematic diagram of an artificial intelligence processor architecture model. The working mode of this accelerator model does not require collaborative execution of neural network operations with the host. Instead, the entire neural network model is compiled on the host to generate all neural network accelerator instructions. During runtime, the input, weights, and constant data are prepared on the host, and the entire network is executed by the neural network accelerator, and the final output is returned to the host. Therefore, this model is a general neural network accelerator in terms of function, and in the neural network framework, there is no need to perform workload partitioning and allocation scheduling between the host and the device for the computational graph.
[0022] In Figure 1 it includes a scratchpad memory unit, a neural network arithmetic unit (NFU), and a control unit (CU). The control unit is responsible for reading the instruction sequence, decoding it, and sending it to each component for execution, as indicated by the arrows with tails in the figure. The CU controls the neural network arithmetic unit and the memory unit separately, which means that the two components can execute instructions simultaneously, overlapping each other's delays to improve the performance of the accelerator. The memory unit is functionally divided into a weight buffer unit (SB) and an input / output neuron buffer unit (NB). The neural network arithmetic unit includes a vector arithmetic unit (VFU) and a matrix arithmetic unit (MFU). Among them, the matrix arithmetic unit reads the data from the SB and NB respectively, and performs operations that require multiplication and accumulation, such as convolution operations, fully connected operations, and matrix multiplication operations, and stores the output results in the NB. The vector arithmetic unit reads the data from the NB and performs vectorized operations such as vector arithmetic operations, logical operations, comparison operations, activation operations, data transposition, and data type conversion, and stores the output results in the NB. The register (Reg) is connected to the control unit, arithmetic unit, and memory unit, supporting the export of data from the operation results, reading data from the memory unit, and immediate number assignment in instructions, and is used for functions such as address indexing in IO instructions, scale configuration in arithmetic instructions, and jump condition judgment.
[0023] Based on the characteristics of each component of the accelerator model, an ISA model can be proposed accordingly. The instruction set is single instruction multiple data streams (SIMD), including calculation (CP) class instructions responsible for execution by the above arithmetic units, input / output (IO) class instructions for data transfer, and some other configuration (CONFIG) class and control jump (CONTROL) class instructions. The configuration class instructions include custom configuration of activation operations and configuration operations of registers, etc. The control jump class instructions are responsible for synchronization instructions, conditional or unconditional jump instructions, etc. of each component. Among them, the IO and CP class instructions can be executed in parallel after being dispatched to the corresponding functional components respectively, and synchronization between different functional components and different types of instructions is achieved through the synchronization instructions in the control class instructions.
[0024] When the accelerator processes the neural network model, it is first necessary to load the necessary data from off-chip memory to on-chip storage through the data loading (Load) instruction of the IO class, and then execute the neural network operation through the corresponding CP instruction. Finally, the result is returned to the off-chip memory through the Store instruction.
[0025] As Figure 2 shown, it is a flowchart of a method for estimating the time required for an artificial intelligence processor to execute a neural network model provided by this application. The estimation method includes the following steps:
[0026] Step 201): Obtain the neural network model and the hardware parameters of the artificial intelligence processor.
[0027] In this embodiment, the hardware parameters of the artificial intelligence processor include, but are not limited to, the DDR bandwidth, LLC bandwidth, and main frequency information of the current device.
[0028] Step 202): Process the neural network model according to the neural network model and the hardware parameters of the artificial intelligence processor to obtain corresponding binary instructions.
[0029] In this embodiment, the step of processing the neural network model according to the network information of the neural network model and the hardware parameters of the artificial intelligence processor includes:
[0030] Obtain the corresponding computational graph according to the neural network model;
[0031] In the computational graph, traverse the nodes of the computational graph in the order of topological sorting, and generate the execution trees of the operators in turn;
[0032] Compile and optimize the execution tree according to the hardware parameters of the artificial intelligence processor;
[0033] Parse the compiled and optimized execution tree to obtain the instruction blocks of the compiled and optimized operators;
[0034] Rearrange the instruction blocks to obtain corresponding hardware binary instructions.
[0035] For further details, first, the computation graph consists of neural network layer nodes and data nodes, marks the data type of each data node, and records the data dependencies between layers and between layers and data. After establishing the computation graph, target-independent optimization can be selected for implementation. Secondly, in the computation graph, traverse the neural network layer nodes in the order of topological sorting, and generate the execution trees of operators in sequence. Then, perform a series of compilation optimizations related to the target hardware part (partially target-independent optimization) based on the execution trees. Finally, parse the execution tree, rearrange the instruction blocks of the operators according to the idea of software pipelining, insert synchronization instructions, and splice the instruction sequences of each layer together to generate the final instruction sequence.
[0036] Step 203): Obtain the time required to execute the neural network model according to the hardware binary instructions corresponding to the neural network model.
[0037] In this embodiment, the step of obtaining the time required to execute the neural network model according to the hardware binary instructions corresponding to the neural network model includes:
[0038] Classify the hardware binary instructions to obtain an I / O instruction set and a computation instruction set;
[0039] Send the I / O instructions in the I / O instruction set to the I / O instruction queue, send the computation instructions in the computation instruction set to the computation instruction queue, and obtain the waiting time when there is an instruction blockage in the instruction queue;
[0040] Obtain the execution time of each I / O instruction and the execution time of each computation instruction. The sum of the execution times of all I / O instructions is the I / O execution time, the sum of the execution times of all computation instructions is the computation execution time, and take the maximum value from the I / O execution time and the computation execution time to obtain the execution time;
[0041] Obtain the time required to execute the neural network model according to the waiting time and the execution time.
[0042] During the operation of each operator, the time overhead mainly includes two main parts: I / O and computation; for the two main parts, there are two instruction queues for I / O and computation, and instructions will be sent to these two queues according to the instruction category, as Figure 3 shown. The instruction-level time estimation module will simulate these two instruction queues with two integer queues, and each queue will be set with a maximum instruction length L. If the next instruction is still of the same category after the instruction queue is full, there will be an instruction blockage. This will generate additional waiting time.
[0043] The entire instruction execution simulation process is iteratively carried out in two steps. The first is the instruction sending step, which fills the instruction queue as much as possible. The second is the instruction execution process, which empties the two queues and increases the total execution time. By continuously iterating through these two steps until all instruction queues are emptied, the obtained total execution time is the final execution time.
[0044] Next is the estimation of the specific execution time of each instruction. The execution time calculation formula for each instruction is generally provided by the hardware department and is slightly corrected through actual measurement of instructions. The following introduces the calculation methods of some main instructions and some special instruction correction parts.
[0045] 1) conv
[0046] The convolution calculation time is very fixed and is obtained by multiplying 6 fields of the instruction. xo * yo * kx * ky * fi * fo, which represents the MACs (multiplication and accumulation) of the data. By representing fi and fo (the number of input and output features) as the number of rows, the MACs can be converted into cycles. In addition, an input sparsity will affect the convolution execution time, and there is also a relationship between the input sparsity and the instruction execution time. Here, this sparsity will be passed in as a configurable parameter together with the hardware information by the upper layer. The final calculation formula is:
[0047] conv_cycles = xo * yo * kx * ky * fi * fo * sparserate
[0048] 2) mlp
[0049] Similarly, it has an input sparse coefficient, and the calculation formula is:
[0050] mlp_cycles = fi * fo * batch * sparserate
[0051] 3) Basic I / O instructions
[0052] Basic I / O instructions are divided into two categories. One is the reading of weights, and the other is the reading and copying of other data. These two basic I / O behaviors are the same, but the bandwidths are different. Finally, the corresponding I / O times will be calculated according to the two bandwidths passed down by the upper layer.
[0053] 4) Stride I / O instructions
[0054] When the Stride is too large, it will cause the phenomenon of extended time. Compared with normal linear reading and linear copying, the I / O offset size with such an offset will be positively correlated with the running time. In this embodiment, therefore, this type of instruction will multiply the running time by a correction coefficient ratio_late according to the value of Stride / data_size.
[0055] As Figure 4 shown, it is a schematic diagram of a device for estimating the time required for an artificial intelligence processor to execute a neural network model proposed in an embodiment of the present application. The device includes:
[0056] A parameter determination unit 401, configured to obtain the neural network model and the hardware parameters of the artificial intelligence processor;
[0057] A processing unit 402, configured to process the neural network model according to the neural network model and the hardware parameters of the artificial intelligence processor to obtain corresponding binary instructions;
[0058] An estimation unit 403, configured to obtain the time required to execute the neural network model according to the hardware binary instructions corresponding to the neural network model.
[0059] An embodiment of the present application discloses a computing chip, and the computing chip includes: As Figure 4 shown in the device.
[0060] An embodiment of the present application also discloses a computer-readable storage medium, storing a computer program for electronic data exchange, wherein the computer program causes a computer to execute the method as Figure 2 shown.
[0061] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.
[0062] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0063] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0064] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
Claims
1. A method for estimating the time required for an artificial intelligence processor to execute a neural network model, characterized in that, The prediction method includes the following steps: Obtain the hardware parameters of the neural network model and the artificial intelligence processor; Process the neural network model according to the neural network model and the hardware parameters of the artificial intelligence processor to obtain corresponding binary instructions; Obtain the time required to execute the neural network model according to the hardware binary instructions corresponding to the neural network model; The step of processing the neural network model according to the network information of the neural network model and the hardware parameters of the artificial intelligence processor includes: Obtain a corresponding computational graph according to the neural network model; In the computational graph, traverse the nodes of the computational graph in the order of topological sorting, and generate the execution trees of the operators in sequence; According to the hardware parameters of the artificial intelligence processor, compile and optimize the execution tree; Parse the execution tree after compilation and optimization to obtain the instruction blocks of the operators after compilation and optimization; Rearrange the instruction blocks to obtain corresponding hardware binary instructions.
2. The method according to claim 1, wherein The step of obtaining the time required to execute the neural network model according to the hardware binary instructions corresponding to the neural network model includes: Classify the hardware binary instructions to obtain an I / O instruction set and a computational instruction set; Send the I / O instructions in the I / O instruction set to the I / O instruction queue, send the computational instructions in the computational instruction set to the computational instruction queue, and obtain the waiting time when there is an instruction blockage in the instruction queue; Obtain the execution time of each I / O instruction and the execution time of each computational instruction. The sum of the execution times of all I / O instructions is the I / O execution time, and the sum of the execution times of all computational instructions is the computational execution time. Take the maximum value from the I / O execution time and the computational execution time to obtain the execution time; Obtain the time required to execute the neural network model according to the waiting time and the execution time.
3. The method according to claim 1, wherein The hardware parameters include: DDR bandwidth, LLC bandwidth, layer frequency.
4. The method according to claim 2, characterized in that, The determination method of the execution time of the I / O instruction includes: Determine the input data size and weight data size of the I / O instruction. The I / O instruction time is equal to the result of dividing the sum of the input data size and the weight data size by the memory bandwidth.
5. The method according to claim 2, characterized in that, The determination method of the execution time of the computational instruction includes: Determine the input data size and weight data size of the computational instruction. The execution time of the computational instruction is equal to the result of dividing the computational amount of each computational instruction by the layer frequency and the data type time parameter.
6. The method according to claim 5, wherein If the input data and weight data of the computational instruction are non-sparse data, the computational amount of the computational instruction is equal to the product of the input data size and the weight data size; If the input data and weight data of the computational instruction are sparse data, the computational amount of the computational instruction is equal to the product of the input data size, the weight data size, and the sparse coefficient.
7. An apparatus for estimating the time required for an artificial intelligence processor to execute a neural network model, characterized in that, The device includes: A parameter determination unit for obtaining the hardware parameters of the neural network model and the artificial intelligence processor; A processing unit, configured to process the neural network model according to the neural network model and the hardware parameters of the artificial intelligence processor to obtain corresponding binary instructions; An estimation unit, configured to obtain the time required to execute the neural network model according to the hardware binary instructions corresponding to the neural network model; Processing the neural network model according to the network information of the neural network model and the hardware parameters of the artificial intelligence processor includes: Obtaining a corresponding computational graph according to the neural network model; Traversing the nodes of the computational graph in the order of topological sorting in the computational graph, and sequentially generating execution trees of operators; Compiling and optimizing the execution tree according to the hardware parameters of the artificial intelligence processor; Parsing the compiled and optimized execution tree to obtain instruction blocks of the compiled and optimized operators; Rearranging the instruction blocks to obtain corresponding hardware binary instructions.
8. A computing chip, characterized in that, The computing chip includes: the device according to claim 7.
9. A computer-readable storage medium, characterized in that, Storing a computer program for electronic data exchange, wherein the computer program causes a computer to execute the method according to any one of claims 1-6.
Citation Information
Patent Citations
Neural network processing method and device, computer device and storage medium
CN110674936A