Methods for interfacing with hardware accelerators
Through the combination of pipeline schemes and single routines, the pipelined operation subset in the operation set is determined and executed, which solves the problem of inefficient execution of computing tasks in the prior art, and realizes the efficient utilization of hardware accelerators and significant improvement in computing efficiency.
Patent Information
- Application Number
- CN202080051285.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-07-15
- Filing Date
- 2020-06-30
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-06-30
AI Technical Summary
In the prior art, when performing a computing task composed of an operation set, it is difficult to efficiently process multiple operations, resulting in ineffective computing efficiency.
Determine a subset of pipelined operations in the operation set through a pipeline scheme, create a single routine to enable the hardware accelerator to perform these operations, trigger and perform composite operations with a single API call to achieve the best utilization of the hardware accelerator.
The execution efficiency of computing tasks is improved, especially when processing matrix vector multiplication of large-density matrices and training of deep neural networks, which significantly reduces the calculation steps and time complexity.
Smart Images

Figure CN114127689B_ABST
Abstract
Description
Background Art
[0001] The present invention relates to the field of digital computer systems and, more particularly, to systems for performing computing tasks consisting of a set of operations.
[0002] Hardware acceleration enables the use of specially manufactured computer hardware to perform some functions more efficiently than is possible in software running on a general-purpose CPU. For example, an operation can be calculated in special-purpose hardware designed to calculate the operation faster than on a general-purpose computer processor. However, there is a need to improve the calculation of many of these operations. Summary of the invention
[0003] Various embodiments provide a method, computer system, and computer program product for performing a computing task consisting of a set of operations.
[0004] In one aspect, an embodiment of the present invention relates to a computer-implemented method for executing a computing task consisting of at least a set of operations. The method includes: determining a subset of pipelineable operations in the set of operations according to a pipeline scheme; creating a single routine for enabling the determined subset of operations to be executed by a hardware accelerator, the routine having input data indicating the computing task and configuration parameter values as arguments, wherein a call to the routine causes the subset of operations on the hardware accelerator to be scheduled according to the configuration parameter values; upon receiving the input data of the computing task that called the routine, thereby causing the hardware accelerator to execute the computing task according to the schedule.
[0005] On the other hand, an embodiment of the present invention relates to a computer system configured to: determine a subset of pipelineable operations of at least one operation set of a computing task according to a pipeline scheme; create a single routine for enabling the determined operation subset to be executed by a hardware accelerator, the routine having input data indicating the computing task and configuration parameter values as independent variables, wherein a call to the routine causes the operation subset on the hardware accelerator to be scheduled according to the configuration parameter values; and upon receiving the input data of the computing task that calls the routine, thereby causing the hardware accelerator to execute the computing task according to the schedule.
[0006] In another aspect, an embodiment of the present invention relates to a computer program product, comprising a computer-readable storage medium having computer-readable program code. The computer-readable program code is configured to: determine a subset of pipelineable operations of at least one set of operations of a computing task according to a pipeline scheme; create a single routine for enabling the determined set of operations to be performed by a hardware accelerator, the routine having input data indicating the computing task and configuration parameter values as arguments, wherein a call to the routine causes the subset of operations on the hardware accelerator to be scheduled according to the configuration parameter values; upon receiving the input data of the computing task that called the routine, thereby causing the hardware accelerator to perform the computing task according to the schedule. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Embodiments of the present invention are explained in more detail below, by way of example only, with reference to the accompanying drawings, in which:
[0008] Figure 1 Depicted is an example structure of a hardware accelerator.
[0009] Figure 2A is a flow chart of a method for using a hardware accelerator to perform a computing task consisting of a set of operations according to an example of the present subject matter.
[0010] Figure 2B A pipelined scheme for matrix-matrix multiplication is described.
[0011] Figure 3A An example hardware accelerator for training deep neural networks is shown.
[0012] Figure 3B Depicted example code.
[0013] Figure 3C A diagram depicting the flow of the task for training a deep neural network.
[0014] Figure 4 is a diagram showing the process of training a deep neural network.
[0015] Figure 5 An example structure of a crossbar array for performing training of a deep neural network is shown. DETAILED DESCRIPTION
[0016] The description of various embodiments of the present invention will be presented for the purpose of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, practical applications, or technical improvements existing in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
[0017] The present subject matter can accelerate the computation performed by the hardware accelerator by using as many units of the hardware accelerator as possible in parallel. In contrast to the serial execution of operations, the present subject matter can take advantage of pipelining because it gives the hardware accelerator information not only about a small portion of the task to be performed, but also about the entire task.
[0018] In the case where the computing task is the training of a deep neural network (DNN), the present subject matter not only gives information about a small part of the network to the hardware accelerator, but also gives information about the entire network required for pipeline operations. The present subject matter can make it possible to group them into one or more compound operations, rather than sending commands for individual network operations (e.g., matrix multiplication, convolution, activation, etc. ...) to the hardware accelerator one by one. The hardware accelerator can then perform these compound operations and execute them according to predefined and optimized pipelines. For example, due to the non-von Neumann characteristics of computational memory, computing resources located on different crossbar arrays can be reused in the form of pipelines. The acceleration obtained by compound operations and pipeline operations may be particularly beneficial for linear algebra applications.
[0019] The present subject matter can provide a software interface for interfacing with a hardware accelerator. The software interface can include functions that enable access to the hardware functions of the hardware accelerator. The single routine can be a function of these functions of the software interface. When the calling program calls the single routine, a command can be issued to the hardware accelerator to perform a computing task. The command to the hardware accelerator can be passed as a composite operation representing a sequence of basic operations supported by the hardware accelerator. The composite operation can be, for example, forward propagation and / or backward propagation of training. Once the hardware accelerator sends the data back to the software interface, the software interface can provide the data to the original calling program. A pipeline scheme (or execution pipeline) can be defined for at least part of the composite operation, for example, a pipeline scheme can be defined for each composite operation. This can allow optimal use of the computing power of the hardware accelerator.
[0020] According to one embodiment, the computing task includes any one of the following: training a deep neural network DNN, matrix-vector multiplication, and matrix multiplication.
[0021] This embodiment may be particularly advantageous for matrix-vector multiplications with large density matrices. For example, due to physical limitations, the crossbar array of a hardware accelerator may only reach a certain size of the matrix to be processed. To this end, the multiplication of large matrices can be split up. This embodiment may enable the user to pass a complete matrix-vector multiplication as a composite operation. The matrix can be broken into appropriate pieces and distributed across different crossbar arrays of the hardware accelerator. The individual matrix-vector multiplications can then be performed in parallel. For example, a matrix that does not fit on a single crossbar may be used The matrix M will be multiplied by the vector This embodiment may enable the multiplication to be performed using the following instructions:
[0022] Make a single API call to this routine,
[0023] Decompose M into A, B, C and D through the computational memory software stack,
[0024] Compute A*X*B*Y*C*X*D*Y* and
[0025] Add the calculation results of the calculation memory software stack.
[0026] This is in contrast to another multiplication technique with the following instructions:
[0027] For example, the user splits M into A, B, C and D.
[0028] Make 4 API calls to calculate A*x, B*y, C*x, and D*y, respectively, and
[0029] It is up to the user to accumulate the matrix accordingly.
[0030] According to one embodiment, the at least set of operations includes a first set of operations for forward propagation of training, and / or a second set of operations for backward propagation of training, and / or a third set of operations for both forward and backward propagation of training. The method includes: generating corresponding composite operations for each set of operations in the first, second and third sets of operations, wherein calling the routine includes executing a single application programming interface (API) call for each composite operation of at least a portion of the generated composite operations. Composite operations can be generated or defined so that a single API call is sufficient to trigger and execute all operations that generate composite operations. Composite operations can be generated so that they are configured to receive a single input and provide the result of performing a computing task (or the result of a set of operations) as output. This can enable a single routine to use the value indicating the input data and the configuration parameter value of the computing task as independent variables. By a single call to the routine, an output indicating the expected result can be obtained.
[0031] According to one embodiment, the configuration parameters include parameters describing the structure of the deep neural network and parameters required for configuring the training of the deep neural network.
[0032] According to one embodiment, the method further comprises providing an application programming interface API to the hardware accelerator, and creating a routine using the API.The hardware accelerator may be, for example, an artificial intelligence based hardware accelerator.
[0033] According to one embodiment, the method also includes providing a computational graph describing a computational task, the computational task involving a deep neural network, determining at least one operation set by parsing the computational graph to identify at least one operation set using nodes of the computational graph, generating a user graph such that each operation set in the at least one operation set is represented by a node of the user graph, wherein the calling routine includes identifying each node of the user graph representing a corresponding operation set, and for each identified node executing a single API call for the operation set represented by the identified node.
[0034] For some applications, the program / sequence of operations is represented as a computational graph (data flow graph), where the nodes represent units of computation. This embodiment can make it possible to convert such a computational graph into a flow that fully uses the computational memory hardware (for example, by generating a new representation using compound operations). To this end, a graph parser can be used to group the pipelined operations in the graph into compound operations. The graph parser can receive a computational graph as input and can output a transformation graph with an appropriate sequence of operations merged into a compound operation. Using this graph parser, programs written in an established deep learning framework can be used directly with a computational memory deep learning accelerator.
[0035] According to one embodiment, the method further comprises receiving an output from the hardware accelerator indicating a result of the computing task.
[0036] According to one embodiment, a pipelining scheme is provided such that each of the operation subsets includes mutually independent operations that can be executed in parallel.
[0037] According to one embodiment, the hardware accelerator operates according to a pipeline scheme using a memristor crossbar array. A subset of pipelineable operations is determined so that each subset of the operations of the subset can be performed in parallel on different crossbar arrays of the memristor crossbar array. The analog memory crossbar array provides an inexpensive vector-matrix computation engine with O(1) computational complexity, providing promising significant acceleration for neural network and linear algebra applications.
[0038] According to one embodiment, a hardware accelerator operates according to a pipeline scheme using a memristor crossbar array, and a computing task includes training a deep neural network, wherein each layer of the deep neural network is associated with two crossbar arrays of the hardware accelerator, and the two crossbar arrays include the same values, wherein causing the hardware accelerator to perform the computing task includes: for each layer of the deep neural network, using one of the two crossbar arrays for forward propagation, and using the other crossbar array only for backward propagation.
[0039] Figure 1 An example architecture of a hardware accelerator is depicted.The hardware accelerator 100 may be, for example, an analog and / or digital based accelerator.
[0040] The hardware accelerator 100 may be configured to perform computational tasks such as training a neural network, running inference using a trained neural network, image processing, summing integers, and the like.
[0041] Like most tasks, computational tasks can be decomposed into sets of operations. For example, in the case of summing numbers, the task can be decomposed into prefix sum operations, which enable the sum of integers to be obtained in an optimal way. In the case of machine learning, most computational tasks are a combination of one or more vector-matrix-multiplications and activation functions. For example, deep neural networks involve vector-matrix multiplications, where the vector x of neuron activations is i will be combined with the weight matrix w ij Multiply them together to generate a new vector y for the neuron excitation of the next layer j This decomposes the computation task into a multiplication-accumulation operation (∑w ij x i ), followed by a nonlinear suppression function.
[0042] Therefore, depending on the computing task, different architectures of the hardware accelerator 100 can be designed to implement the operation of the task. In other words, a person skilled in the art with a given computing task can provide an architecture of a hardware accelerator that enables at least part of the computing task. In the following, the hardware accelerator 100 is described with reference to artificial intelligence applications, but it is not limited thereto.
[0043] The hardware accelerator 100 includes an integrated circuit 101. The integrated circuit 101 is configured to perform operations on analog and / or digital signals. The integrated circuit 101 includes a plurality of physically implemented functional units 103A-N, and the functional units 103A-N are provided so that the conventional instruction fetch and decode steps of the instruction cycle are not required to perform the computational task. For example, the functional units 103A-n may form a hierarchy of chips, including a memristor array, an ADC at the periphery of the array, an embedded DRAM (eDRAM) for buffering intermediate items, and a digitized array output, such as for implementing the multiply-accumulate operations involved in the forward reasoning of the DNN.
[0044] The functionality of the hardware accelerator 100 depends on the functional units 103A-N selected for the hardware accelerator 100. For example, parameters such as the size of the memristor crossbar array, the number of crossbars, the number of ADCs, etc., may be used in order to define an algorithm according to which the hardware accelerator 100 may perform a computational task. For example, the algorithm may utilize parallel computing and pipelining schemes to reduce the number of steps of the computational task and thus may reduce the time complexity compared to another algorithm that performs the computations sequentially.
[0045] Thus, depending on the algorithm used to operate the hardware accelerator 100, the functional units 103A-N may be configured to receive and provide data between each other according to the algorithm. To this end, the hardware accelerator 100 may include a component 105 that controls and sequences events in time. The component 105 may include one or more finite state machines. The finite state machine may be driven by loading a control vector into the hardware accelerator 100, such as a mapping of the functional units 103A-N, and the pipeline scheme may be determined offline and loaded into a control register that drives the finite state machine.
[0046] Figure 2A is a flow chart of a method for performing a computing task consisting of a set of operations using a hardware accelerator (eg, 100 ) according to an example of the present subject matter.
[0047] For the purpose of simplicity, reference is made to the computational task described as a matrix-matrix multiplication. Figure 2A In the case of matrix-matrix multiplication, the multiplication can be decomposed into a sequence of matrix-vector multiplications, where the group of operations is matrix-vector multiplication.
[0048] In order to optimally or maximally use the hardware accelerator 100, a pipeline scheme may be used. The pipeline scheme may define a pipeline divided into a plurality of stages, wherein each stage completes a portion of a computing task in parallel, and the stages are related to each other to form a pipeline. The pipeline scheme may be determined based on the structure and function of the functional unit and the computing task, for example, the determination of the pipeline scheme may take into account knowledge about the hardware capabilities of the hardware accelerator, such as the number of memristor crossbar operations that may be calculated in parallel.
[0049] Following the matrix-matrix multiplication example, the computational task may be to perform a matrix multiplication M 1 xM 2 …xM 5 For example, each matrix in the matrix can be a 4×4 matrix. In order to perform this chain of matrix multiplications in an optimal way, the following method or process can be used: The matrix M 1 ×M 2 …×M 4 Each matrix of can be stored in a corresponding cross array, and finally the matrix M 5 can be decomposed into column vectors and the vectors can be fed into a crossbar array, such as Figure 2B Based on this process, a pipeline scheme can be defined to optimally perform the following Figure 2B The multiplication M shown in Table 220 1 ×M 2 …×M 5 , where 5 levels (or time steps) 222.1-5 are defined, and one or more matrix-vector multiplications may be performed in each level. Figure 2B As shown, in the first stage 222.1, the storage matrix M is used n For example, the cross array is fed with the vector x 1 The 4 elements of x can be executed 2 1 =M n x 1 The first stage 222.1 can be X 2 1 As output (the result of the multiplication) is provided to the second stage 222.2. In the second stage 222.2, the matrix M is stored. n The crossbar array can be executed x 2 2 =M n x 2 The second matrix-vector multiplication is performed because the crossbar array becomes idle after completing the first stage. In parallel with the second multiplication, a third multiplication can be performed, namely x 3 1 =M n-1 x2 1 Since the third multiplication requires the result of the first multiplication, the first multiplication is only performed in the second stage 222.2 after the first multiplication is performed. In the last two stages 222.4-5, all crossbar arrays run the corresponding multiplications in parallel, thereby achieving full utilization of the hardware accelerator.
[0050] Therefore, based on the pipeline scheme, a subset of pipelineable operations may be determined from the set of operations according to the pipeline scheme in step 201. The subset of pipelineable operations may, for example, include operations that may be performed in parallel, for example, in a given stage of the pipeline. The determined subset of operations may allow full or optimal utilization of the hardware accelerator 100. Figure 2B In an embodiment, the first subset of operations may include x 2 1 =M n x 1 operation, the second subset of operations may include x 2 2 =M n 2 and x 3 1 =M n-1 x 2 1 Two operations, the third subset of operations may include x 2 3 =M n x 3 、x 3 2 =M n-1 x 2 2 and x 4 1=M n-2 x 3 1 Three operations, and so on.
[0051] It has been defined, for example Figure 2B The pipeline of operations to be performed shown in the figure, the present method can be advantageous because it can only require a single routine to enable the execution of the entire computing task. In step 203, a single routine can be created so that the independent variables of the routine can indicate to the hardware accelerator the data that enables the execution of the pipeline, for example, no further input from the routine is required. For example, the independent variables can include values indicating input data and configuration parameter values of the computing task. In one example, an API can be provided to interface with the hardware accelerator 100, wherein the single routine can be a function of the API. In this case, the call of the single routine can be referred to as an API call. In another example, a single routine can be defined using a function of the API.
[0052] The call of the routine causes a subset of operations on the hardware accelerator 100 to be scheduled according to the configuration parameter values. For example, the configuration parameter values may be loaded into the hardware accelerator 100 as a control vector to drive a finite state machine that properly manipulates inputs and outputs after each cycle / phase.
[0053] For example, the call to the routine may be performed as follows: 1) a single API call is made that references all five matrices; 2) the software stack inserts M 1 M 2 M 3 and M 4 3) The row vectors of X are pipelined through the crossbar. This is in contrast to the approach of making at least 5 API calls to compute a single matrix-matrix multiplication.
[0054] Steps 201 and 203 may be performed offline, for example, before using the hardware accelerator 100 for calculation.
[0055] Upon receiving the input data of the computing task, a routine may be called in step 205 so that the hardware accelerator 100 may perform the computing task in step 207 according to the schedule. The result of the computing task may be received from the hardware accelerator 100. According to the above example, the hardware accelerator may include storing the matrices M 1 To M 4 In this case, the arguments to the routine may include the matrix M 5 The vector x 1 to x4, as input data and as indicator matrix M 1 、M 2 , M3 and M 4 For example, instead of executing the following four calls, mm 1 =Matmul(M 4 , M 5 );mm 2 =Matmul(M3, mm 1 ); mm3 = Matmul(M 2 , mm 2 ); and OUTPUT = Matmul(M 1 , mm3), a single call (such as an API call) can be executed as follows OUTPUT = composition (config, M 5 ), where the configuration parameter can be defined as config = MatrixMatrixMultiplicationChain(M 1 , M 2 , M3, M 4 ).
[0056] Figure 3A An example hardware accelerator 300 is shown for training a DNN having an input layer 301, one hidden layer 303, and an output layer 305. In this case, the set of operations may include operations for forward propagation of training and / or operations for backward propagation of training.
[0057] The three layers have 784, 250, and 10 neuromorphic neuron devices, respectively. The output layer has 10 neuromorphic neuron devices representing 10 possible numbers 0 to 9, and the input layer has 784 neuromorphic neuron devices representing the number of pixels of the input MNIST image. Each of the neuron devices can be configured to use an activation function for generating an activation function based on the current state of the neuron device (e.g., x i The hardware accelerator 300 may also include two crossbar arrays or memristor crossbar arrays (not shown) for respectively calculating the weight elements W JI and W KJ The multiplication with the activation vector x, for example, has elements W JI The matrix-vector multiplication of the matrix W of the input layer and the activation vector x of the input layer can be achieved by a first memristive crossbar array, by representing each matrix element with the conductance of the corresponding memristive element of the first memristive crossbar array, wherein the multiplication of the matrix W and the vector x can be performed by inputting a voltage representing the vector value x to the first memristive crossbar array, and the resulting current indicates the product of W and x. The resistive storage element (or device) of the crossbar array can be, for example, one of a phase change memory (PCM), a metal oxide resistive RAM, a conductive bridge RAM, and a magnetic RAM. In Figure 3A In this example, the functional unit may include at least two crossbar arrays and a neuromorphic neuron device.
[0058] Knowing the computational task of training a 3-layer DNN and having access to the operation of the functional units of the hardware accelerator 300, a pipeline scheme can be defined with a given number of stages (see Figure 3C ), where in each stage, one or more operations may be performed in parallel by the functional units of the hardware accelerator 300.
[0059] Instead of Figure 3B As shown in the code 310, there is an API call for each layer operation (e.g., matrix multiplication, convolution, activation, pooling, etc.), which can be used as Figure 3B The single API call 313 shown in code 312 of FIG. The input to the API call 313 may be an MNIST image and configuration parameters 314 describing the DNN, as indicated by code 312. By executing code 312, multiple operations may be chained and performed together.
[0060] Figure 3C Depicted is a diagram illustrating an example of a method for Figure 3A A first diagram 330 of an execution scheme or algorithm for training a DNN of FIG. 10 and a diagram illustrating an example of the present subject matter for Figure 3A A second diagram 350 of an execution scheme for training a DNN of FIG. 350 and a diagram illustrating another example according to the present subject matter Figure 3A A third diagram 360 of an execution scheme for training a DNN.
[0061] For example, training of a DNN may require input of multiple image sets, and for each image set, forward propagation can be performed without changing the synaptic weights, so that the prediction error of the DNN to be backpropagated can be estimated by combining the errors obtained for that image set (rather than just one image).
[0062] The first diagram 330 is a computational diagram indicating the flow of a computational task. For example, the weight 331 of the first set of inputs and the input vector 332 may be multiplied in response to the first API call of the matmul function 333. The result of the first API call is used to execute the second API call of the sigmoid function 334. The result of the second API call is used to execute the third API call of the Matmul function 335, including multiplying the weight 336 and the vector result of the second API call. The result of the third API call is used to execute the fourth API call of the sigmoid function 337. The vector obtained from the fourth API call and the label 338 of the input 332 may be used to calculate the loss function 339. The difference between the vector obtained from the fourth API call and the label 338 may be used to calculate the prediction error δ performed by the DNN. The calculated prediction error δ may be back-propagated. And the increment ΔW of all weights may be used to update the weights 331 and 336 after back-propagation, as shown in Figure 340. These API calls may be repeated for each additional input 332 until the computational task is performed, for example, the computational task may require 100 input images for forward propagation. After the last API call for the first set of inputs is completed, the second set of inputs enters the first graph. Therefore, when processing the first set (or the second set of inputs), the computational tasks performed following the flow of the first graph 300 may not benefit from the fact that the weights 336 and 331 do not change for each set of inputs, e.g., each of the crossbars storing the weights 336 and 331 is not used for parallel computation.
[0063] In order to take advantage of parallel computing, the process described by the second diagram 350 can be used. The second diagram 350 is a computational graph indicating the process of a computational task according to an example of the present subject. In order to implement the process of the second diagram 350, two pipeline schemes can be defined, one for the forward propagation of training and the other for the backward propagation of training. In this case, the group of inputs 332 is provided as an input to a composite operation 353 in combination with weights 331 and 336, which can be called by a single routine to perform forward propagation. The composite operation 353 can process the input according to the pipeline scheme, for example, if the input set includes two images, during the first stage, the first crossbar array only processes the first image, and during the second stage / cycle of the pipeline, the second crossbar stores weights 336, and in parallel, the second image is processed using the first crossbar array storing weights 331. As described above, the prediction error is estimated using a loss function 339. The prediction error can be back-propagated using matrix-vector multiplication. This is indicated by another composite operation 355. The compound operation 355 can process the input of the backward propagation of the prediction error according to the pipeline scheme in a similar manner as the forward propagation. And, as shown in diagram 380, the weights 331 and 336 can be updated using the increment ΔW of all weights.
[0064] Thus, during training of the DNN, the second graph 350 enables forward and backward propagation to be performed in different composite operations. This separation between forward and backward propagation can be advantageous because the second graph 350 design can be used only for inference (without the need to perform backward propagation). In addition, the process of the second graph 350 can work directly with techniques that require information about the entire batch (e.g., batch normalization) and that occur in the stage between the forward and backward propagation processes. This is in Figure 4 , where batch normalization can still be kept separate or independent from the pipeline schemes used for forward and backward propagation. This can also enable greater freedom in choosing the loss function since the loss function is not covered by the two pipeline schemes. Briefly, Figure 4 Two schemes are described, the first scheme having the operations of convolution 402, rectified linear unit 404, convolution 406, rectified linear unit 408, batch normalization 410, convolution 412, rectified linear unit 414, convolution 416, and rectified linear unit 418, and the second scheme having compound operation 420, batch normalization 422, and compound operation 424.
[0065] Back to Figure 3C, in order to further utilize parallel computing, the process described by the third figure can be used. The third figure 360 is a calculation diagram of the process of the instruction computing task according to the example of this theme. In order to realize the process of the third diagram 360, a pipeline scheme is defined for forward and backward propagation and loss function calculation. In this case, input set 332 is provided as the input of composite operation 363 in combination with weights 331 and 336, and this composite operation can be called by a single routine to perform forward propagation and backward propagation according to the pipeline scheme that attempts to parallelize the operation as much as possible. Those operations to be parallelized involve matrix-vector multiplication using a cross matrix and activation function and loss function calculation using neurons. For example, when the second crossbar array is used for back-propagation error signal, the first crossbar array can be used for calculating the matrix-vector multiplication of forward propagation. In this example, additional memory may be needed to save the activation and error signal of forward and backward propagation calculation.
[0066] Thus, during DNN training, the third graph 360 enables forward and backward propagation to be performed in the same compound operation. This may be advantageous because it may require less memory consumption. For example, once ΔW is calculated, the pre-stored layer activations may be discarded and the memory may be reused for another sample in the batch. Another advantage may be that the execution flow of the third graph may require less overhead. For example, at the beginning and end of a compound operation, there may always be an overhead period where not all arrays are used. By reducing the number of compound operations, this overhead may be reduced.
[0067] Another advantage of the process of the third diagram 360 may be that the process can be combined with Figure 5 Combining the array replication techniques shown, for example, two crossbar arrays of a DNN can be replicated (i.e., multiple crossbar arrays containing the same weights) so that one crossbar array is used only for the forward pass and the other is used only for the backward pass, as shown in Figure 5 As shown, Figure 5 Layer 1 (item 502) and layer 2 (item 504) of relate to the input layer 301 and hidden layer 303 of the DNN, respectively. Arrays Array1 and Array 2 are crossbar arrays that perform matrix-vector multiplications that occur between the input layer and the hidden layer, and between the hidden layer and the output layer, respectively. This can allow multiple operations to be performed simultaneously on the same layer. Specifically, Figure 5 Data 514 is shown being input through layer 1 array 2 (item 510) and then via forward propagation 518 through layer 2 array 2 (item 512); then from layer 2 array 2 to layer 2 array 1 (item 508); and then from layer 2 array 1 (item 508) via back propagation 516 through layer 1 array 1 (item 506).
[0068] Various aspects of embodiments of the present invention are described herein with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to embodiments of the present invention. It will be understood that each frame of the flowchart and / or block diagram and the combination of frames in the flowchart and / or block diagram can be implemented by computer-readable program instructions.
[0069] Embodiments of the present invention may be systems, methods and / or computer program products. A computer program product may include a computer-readable storage medium (or multiple media) having computer-readable program instructions thereon, the computer-readable program instructions being used to cause a processor to perform various aspects of embodiments of the present invention.
[0070] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punch card or a raised structure in a groove on which instructions are recorded, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be interpreted as a temporary signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (e.g., a light pulse by an optical fiber cable), or an electrical signal transmitted by a wire.
[0071] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in a computer-readable storage medium within the corresponding computing / processing device.
[0072] The computer-readable program instructions for performing the operation of an embodiment of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" programming language or similar programming languages. The computer-readable program instructions can be executed completely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or completely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider through the Internet). In some embodiments, in order to perform various aspects of an embodiment of the present invention, the electronic circuit including, for example, a programmable logic circuit, a field programmable gate array (FPGA) or a programmable logic array (PLA) can execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit.
[0073] Various aspects of embodiments of the present invention are described herein with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to embodiments of the present invention. It will be understood that each frame of the flowchart and / or block diagram and the combination of frames in the flowchart and / or block diagram can be implemented by computer-readable program instructions.
[0074] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device create a device for implementing the functions / actions specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which can guide the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable storage medium having the instructions stored therein includes an article of manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes of the flowchart and / or block diagram.
[0075] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more boxes of the flowchart and / or block diagram.
[0076] Flowchart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention.In this regard, each frame in flow chart or block diagram can represent the module, segment or part of instruction, and it comprises one or more executable instructions for realizing the logical function of appointment.In some alternative embodiments, the function mentioned in the frame may not occur in the order mentioned in the figure.For example, two frames shown continuously can actually be performed substantially simultaneously, or these frames can sometimes be performed in reverse order, depending on the function involved.It will also be noted that the combination of the frame in each frame of block diagram and / or flow chart illustration and block diagram and / or flow chart illustration can be realized by the dedicated hardware-based system of performing specified function or action or performing the combination of special hardware and computer instruction.
Claims
1. A computer-implemented method for performing a computing task consisting of at least one set of operations, the method comprising: include: Determining a subset of pipelineable operations in the operation set according to a pipeline scheme; creating a single routine for enabling the determined subset of operations to be performed by a hardware accelerator, the routine having as arguments values indicative of input data for the computing task and configuration parameter values, wherein a call to the routine enables scheduling of the subset of operations on the hardware accelerator according to the configuration parameter values; When input data of the computing task is received, the routine is called so that the hardware accelerator executes the computing task according to the schedule.
2. The method of claim 1, wherein the computing task comprises any one of: training a deep neural network, using the trained neural network to perform inference, matrix-vector multiplication, and matrix-matrix multiplication.
3. The method according to claim 2, wherein the at least one operation set comprises a first operation set for forward propagation of the training, a second operation set for backward propagation of the training, and a third operation set for both the forward propagation and the backward propagation of the training; the method further include: A respective composite operation is generated for each of the first, second, and third sets of operations, wherein calling the routine includes executing a single application programming interface (API) call for each composite operation of at least a portion of the generated composite operations.
4. The method according to claim 2, wherein the configuration parameters include parameters describing the structure of the deep neural network and parameters required for configuring the training of the deep neural network.
5. The method of claim 1, further comprising providing an application programming interface (API) to the hardware accelerator and creating the routine using the API, wherein the call to the routine is a single API call.
6. The method according to claim 1, further comprising: include: A computation graph describing the computation task is provided, the computation task involving a deep neural network, the at least one operation set is determined by parsing the computation graph to identify the at least one operation set using nodes of the computation graph, a user graph is generated such that each operation set in the at least one operation set is represented by a node of the user graph, wherein calling the routine includes identifying each node of the user graph representing a corresponding operation set, and performing, for each identified node, a single API call for the operation set represented by the identified node.
7. The method of claim 1, further comprising receiving an output from the hardware accelerator indicating a result of the computing task.
8. The method of claim 1, wherein the pipelining scheme is provided such that each of the operation subsets includes mutually independent operations that can be executed in parallel.
9. The method of claim 1 , wherein the hardware accelerator operates according to the pipeline scheme using a memristor crossbar array, wherein the subset of pipelineable operations is determined such that each subset of operations of the subset can be executed in parallel on different crossbar arrays of the memristor crossbar array.
10. The method of claim 1, wherein the hardware accelerator operates according to the pipeline scheme using a memristor crossbar array, the computational task comprises training a deep neural network, wherein each layer of the deep neural network is associated with two crossbar arrays of the hardware accelerator, the two crossbar arrays comprising the same values, wherein causing the hardware accelerator to perform the computational task include: For each layer of the deep neural network, one of the two crossbar arrays is used for forward propagation, and the other crossbar array is used only for backward propagation.
11. A computer program product comprising a computer readable storage medium having computer readable program code embodied therewith, the computer readable program code being configured to: Determining a subset of pipelineable operations of at least one operation set of the computing task according to the pipeline scheme; creating a single routine for enabling the determined subset of operations to be performed by a hardware accelerator, the routine having as arguments values indicative of input data for the computing task and configuration parameter values, wherein a call to the routine causes the subset of operations on the hardware accelerator to be scheduled according to the configuration parameter values; Upon receiving input data of the computing task that calls the routine, the hardware accelerator is thereby caused to execute the computing task according to the schedule.
12. The computer program product of claim 11, wherein the computing task comprises one of: training a deep neural network, matrix-vector multiplication, and matrix-matrix multiplication.
13. A computer program product according to claim 12, at least one operation set includes a first operation set for forward propagation of the training, a second operation set for backward propagation of the training, and a third operation set for both the forward propagation and the backward propagation of the training, and the computer-readable program code is also configured to: generate a corresponding synthetic operation for each operation set in the first operation set, the second operation set, and the third set of operation sets, wherein calling the routine includes executing a single application programming interface (API) call for each synthetic operation of at least a part of the generated synthetic operations.
14. The computer program product of claim 12, the configuration parameters comprising parameters describing a structure of the deep neural network and parameters required for configuring the training of the deep neural network.
15. The computer program product of claim 11, further configured to create the routine using an API to the hardware accelerator.
16. The computer program product of claim 11 , further configured to provide a computational graph describing the computational task, the computational task involving a deep neural network, determining at least one pipelineable operation set by parsing the computational graph for identifying the at least one operation set using nodes of the computational graph, generating a user graph such that each operation set in the at least one operation set is represented by a node of the user graph, wherein calling the routine comprises identifying each node of the user graph representing a corresponding operation set, and executing, for each identified node, a single API call for the operation set represented by the identified node.
17. The computer program product of claim 11, further configured to receive an output from the hardware accelerator indicating a result of the computing task.
18. The computer program product of claim 11, the pipelining scheme being provided such that each of the subsets includes mutually independent operations that can be executed in parallel.
19. A computer system configured to: Determining a subset of pipelineable operations of at least one operation set of the computing task according to the pipeline scheme; creating a single routine for enabling the determined subset of operations to be performed by a hardware accelerator, the routine having as arguments values indicative of input data for the computing task and configuration parameter values, wherein a call to the routine causes the subset of operations on the hardware accelerator to be scheduled according to the configuration parameter values; Upon receiving input data of the computing task that calls the routine, the hardware accelerator is thereby caused to execute the computing task according to the schedule.