Accelerator for Dense and Sparse Matrix Computations

By distinguishing dense and sparse matrix operations in computing devices and using special computing devices to process them separately, the problem of inefficient hybrid matrix calculation in the prior art is solved, and more efficient computing performance and energy utilization are achieved.

CN115023685BActive Publication Date: 2025-07-04MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180011738.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-29
Filing Date
2021-01-20
Publication Date
2025-07-04
Estimated Expiration
2041-01-20

AI Technical Summary

Technical Problem

Existing dedicated computing hardware is inefficient when dealing with hybrid dense and sparse matrix calculations, especially when dealing with large numbers of zero-value elements, resulting in unnecessary power consumption and computational delays.

Method used

By receiving the matrix calculation signal in the computer processing device, and using the sparse data inspection device to distinguish dense and sparse operands, the operations are forwarded to a dedicated dense computing device or sparse computing device for processing, the sparse computing device optimizes the calculation by ignoring the zero-value operands.

Benefits of technology

It realizes reducing energy consumption and delays when processing mixed matrix calculations, improving calculation efficiency, and avoiding unnecessary calculations especially when sparse data does not contribute to the final result.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115023685B_ABST
    Figure CN115023685B_ABST
Patent Text Reader

Abstract

A method for increasing the computer hardware efficiency of matrix calculations. The method includes receiving, at a computer processing machine, a digital signal encoding one or more operations of a matrix calculation, each operation including one or more operands. The method further includes: in response to the sparse data inspection device of the computer processing machine determining that the operation of the matrix calculation includes all dense operands, forwarding the operation to the dense calculation device of the computer processing machine, the dense calculation device being configured to perform the operation of the matrix calculation based on the dense operands. The method further includes: in response to the sparse data inspection device determining that the operation of the matrix calculation includes one or more sparse operands, forwarding the operation to the sparse calculation device, the sparse calculation device being configured to perform the operation of the matrix calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Matrix calculations can be accelerated using specialized computing hardware. However, such specialized computing hardware is typically inefficient when performing matrix calculations with a relatively large proportion of zero-valued elements. Summary of the Invention

[0002] The present Summary of the Invention is provided to introduce a selection of concepts in a simplified form that will be further described in the Detailed Description below. The present Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Additionally, the claimed subject matter is not limited to embodiments that solve any or all disadvantages noted in any part of this disclosure.

[0003] A method for increasing the computer hardware efficiency of matrix calculations includes: receiving, at a computer processing device, a digital signal encoding one or more operations of a matrix calculation, each operation including one or more operands. In response to a determination by a sparse data inspection device of the computer processing machine that the operation of the matrix calculation includes all dense operands, forwarding the operation to a dense calculation device of the computer processing machine, the dense calculation device being configured to perform the operation of the matrix calculation based on the dense operands. In response to a determination by the sparse data inspection device that the operation of the matrix calculation includes one or more sparse operands, forwarding the operation to a sparse calculation device, the sparse calculation device being configured to perform the operation of the matrix calculation. Brief Description of the Drawings

[0004] Figures 1A - 1B An exemplary accelerator pipeline for matrix calculations is shown.

[0005] Figure 2 A method for increasing the computer hardware efficiency of matrix calculations is shown.

[0006] Figure 3 An exemplary microarchitecture of an accelerator for performing matrix calculations on sparse and / or dense matrix data is shown.

[0007] Figure 4 An exemplary computing system is shown. Detailed Description

[0008] Computer computations often include matrix computations, such as matrix multiplication. For example, neural networks (e.g., convolutional neural networks and / or deep neural networks) can be implemented using one or more matrix multiplications. Thus, an accelerator can perform one or more matrix multiplications, providing numerous benefits (e.g., lower latency, higher bandwidth, and / or lower power utilization) for neural network implementations. Non-limiting examples of computer programs and / or algorithms that can be substantially implemented using matrix computations include graphics processing programs (e.g., rendering software) and artificial intelligence models, such as: deep neural networks (DNN), convolutional neural networks (CNN; e.g., deep convolutional neural networks (DCNN)), recurrent neural networks (RNN; e.g., long short-term memory (LSTM)), and many other artificial intelligence models. These artificial intelligence models can be implemented using multiple layers of matrix computations that start with an input vector and process the input vector at each layer of the matrix computation to compute an arbitrary function, such as a function learned using a machine learning training algorithm. Neural networks that utilize such matrix computations are capable of achieving state-of-the-art results in many applications, such as computer vision, machine translation, speech assistance, etc. As an example of a neural network model that can be substantially implemented using matrix computations, a CNN can include multiple layers that can be mathematically decomposed into large-scale matrix multiplications (convolutions), followed by element-wise non-linear transformations, such as sigmoid or rectified linear unit functions. As another example, a DNN can perform a large number of matrix computations.

[0009] However, matrix computations can involve large computational runtime costs, memory, and power consumption. For example, a neural network for processing image data (e.g., photo data, video data, etc.) can perform hundreds or thousands of arithmetic operations to process just a single image. As another example, existing neural networks can have more than 100 layers, where processing each layer includes performing hundreds or even thousands of arithmetic operations.

[0010] Specialized computing hardware can potentially be used to reduce the runtime costs, memory utilization, processing time, and power consumption of matrix computations. Relative to sequential computations, such as matrix computations using conventional general-purpose computing hardware, specialized computing hardware can significantly reduce these costs. For example, an accelerator can achieve benefits such as lower latency, higher bandwidth, and / or lower power utilization compared to a general-purpose processor.

[0011] In some examples, such specialized computing hardware can perform computations organized in a pipeline. For example, as Figure 1AAs shown, accelerator pipeline 100 is a non-limiting example of a pipeline for data flow and logic in matrix computations. For example, accelerator pipeline 100 includes an instruction issue stage 102 that receives instructions (e.g., issued from different logic devices and / or streamed from a storage device). Multiplication instructions are sent to multiplier engine 104. Similarly, memory instructions (e.g., load / store instructions) are sent to memory engine 106. Instructions are completed in a defined order in the pipeline (e.g., in program order, according to control and / or data dependencies, and / or in any other suitable order that introduces appropriate semantics for matrix operations), and the results of the instructions are finalized in a commit stage 108 (e.g., by writing the computed results to registers and / or causing one or more storage devices to be updated to perform the load / store operations processed by the memory engine). In each cycle, a number of instructions are issued (e.g., read from an instruction memory and assigned to a pipeline that matches their type). The number of instructions per cycle is referred to as the issue bandwidth. In the next cycle, each instruction is processed in the pipeline by reading the operand values of the instruction from a register file, executing the instruction using its designated engine, and writing the result to a register and / or storing data via memory engine 106. After reaching the commit stage 108, the instruction is marked as executed and no longer exists in the pipeline.

[0012] As a non - limiting example of the general design of an accelerator, the accelerator can compute matrix computations asynchronously, concurrently, and / or substantially in parallel to reduce the runtime cost. For example, the logic hardware can execute the steps of the matrix computation in a scheduled manner to achieve an efficient data flow throughout the computation. For example, an efficient parallel implementation can ensure that input data is accessed from memory in an optimal order when needed. The efficient parallel implementation can solve sub - problems of the matrix computation to exploit spatial and / or temporal locality, reducing memory access and computation latency. Another non - limiting example of dedicated computing hardware for matrix computations is a systolic array (SA) for matrix multiplication. A systolic array for matrix multiplication can be implemented as a plurality of tightly - coupled 2D multiply - accumulate (MAC) computing nodes that are highly synchronized to process data in synchronization with a schedule (e.g., a clock) when the data arrives at the systolic array. Matrix multiplication can be decomposed into local operations for computing multiple parts of the matrix product. Matrix multiplication can be decomposed in any suitable way, e.g., by computing block parts of the matrix, consecutive diagonals of the matrix, etc. Thus, in an SA for matrix multiplication, the 2D MAC computing nodes can be configured to: perform matrix multiplication decomposed into local operations using only nearest - neighbor communication between the MAC computing nodes. The local operations can be computed using only the input data and / or intermediate computation values required for the local operations. Thus, an SA for matrix multiplication can reduce (e.g., minimize) memory access latency and / or power consumption. For example, connections between non - local computing nodes can represent a large potential power cost. By specifically using local connections, an SA for matrix multiplication can significantly reduce power consumption relative to sequential matrix computations. As a non - limiting example, an SA for matrix multiplication can perform a series of multiplication operations representing specific element - by - element products in the matrix multiplication, and can perform a series of accumulation (e.g., addition) operations to add the element - by - element products and save the resulting sum in an appropriate memory destination corresponding to the output matrix produced by the matrix multiplication. The multiplication and addition can be performed asynchronously (e.g., when the input data is available, the multiplications can be performed simultaneously and / or in parallel, and when the ongoing series of multiplication operations produces products, the additions can be performed concurrently and / or in parallel).

[0013] In addition to matrix multiplication, an accelerator (e.g., a systolic array-based accelerator) can be configured to perform various other matrix calculations. Non-limiting examples of matrix calculations that can be implemented at least in part via the accelerator include principal component analysis (PCA), Fourier transforms, and / or matrix addition and subtraction. Additionally, matrix calculations as described herein can also refer to generalized tensor calculations (e.g., vector calculations and higher tensor calculations), which can similarly be implemented using dedicated computing hardware for performing the planned calculations. For example, non-limiting examples of vector / matrix / tensor calculations that can be effectively implemented using the techniques of the present disclosure include pooling operations (e.g., max pooling), Hadamard product operations, etc.

[0014] Returning to the example of matrix multiplication, a typical systolic array is designed for dense matrix operations, such as dense multiplication. A "dense matrix" is used herein to refer to a matrix in which all elements are explicitly defined, including elements with zero values. A dense matrix can have a memory consumption for storage that depends on the size of the matrix dimensions (e.g., 2x2, 3x3, etc.), as each element of the matrix is explicitly defined and stored. Storing and computing using a dense matrix can be particularly effective when the matrix includes a relatively small number of zero elements.

[0015] However, inputs for matrix calculations typically include many zero-valued elements. For example, data with more than 50% zero-valued elements can be referred to as sparse data. When the data in a calculation (e.g., in AI and / or DNN calculations) is sparse, providing the accelerator with instructions to operate on the sparse data and generate additional sparse data that does not contribute to the overall result (e.g., sparse data that does not affect the overall learning or prediction of an AI system) may unnecessarily waste accelerator power.

[0016] For example, due to non-linear (e.g., rectified linear unit) activation and quantization, the input for each layer of a neural network can include many zero-valued elements. In some examples, a matrix can be stored as a "sparse matrix", which is used herein to refer to a matrix in which only non-zero elements are explicitly defined. As a non-limiting example, a sparse matrix can be stored in the form of two vectors, namely an order vector indicating which elements of the sparse matrix are populated (e.g., a bit vector that indicates a "1" value for non-zero entries and a "0" value for zero entries in row-column lexicographical order) and a data vector that includes all non-zero elements (e.g., listed in row-column lexicographical order). When the number of non-zero entries is relatively small, storing and computing using a sparse matrix can be particularly efficient because only non-zero entries are explicitly defined. Thus, only non-zero elements need to be stored, and in some cases, computations can be simplified or optimized based on the implicit encoding of zero-valued elements (e.g., skipping a portion of the computation that corresponds to computing the product of one or more values, including one of the implicitly encoded zero values). In some examples, sparse matrix data can be "unpacked" to populate a dense matrix, e.g., by explicitly storing all non-zero and zero elements indicated by the sparse matrix.

[0017] It is believed that much (although not necessarily all) of the data loaded from memory in many matrix computations is sparse. An accelerator that processes sparse data in the same way as dense data can incur a high power cost when performing computations on sparse data that do not contribute to the overall final result of the computation (e.g., unnecessarily multiplying zero-valued entries, producing zero-valued products that do not contribute any numerical weight to the result of the neural network computation). However, an accelerator can be configured to process dense and sparse matrices separately and may not be able to handle a mixture of dense and sparse data. In some examples, an accelerator can be configured to unpack a sparse matrix and process the data from the sparse matrix as dense matrix data. However, such unpacking and processing can cause excessive energy consumption due to unpacking, storing, and performing computations on all zero values in the sparse matrix. Thus, previous accelerators may not be optimal when dealing with a mixture of dense and sparse matrix data. Similarly, specialized processing of sparse data can be inefficient in processing dense matrix data compared to accelerated dense matrix data computations.

[0018] Figure 1BAn accelerator pipeline 100’ is shown, which is configured to efficiently process a mixture of dense and sparse matrix data, such as to perform matrix multiplication operations (or any other suitable matrix operations). Similar to pipeline 100, the accelerator pipeline 100’ is configured to process instructions received in an instruction issue stage 102’ and ultimately write out results in a commit stage 108’. However, the accelerator pipeline 100’ is configured to distinguish instructions that process one or more sparse data operands from instructions that only process dense data. The accelerator pipeline 100’ uses a sparse data inspection device 103’ for this distinction, which is configured to: process the issued instructions and their operands when the issued instructions (e.g., arithmetic and / or memory load / store operations) and their operands are received in the issue stage 102’. The sparse data inspection device 103’ is configured to evaluate whether one or more operands of an instruction are sparse. As a non-limiting example, if an operand has a zero value, the sparse data inspection device 103’ may evaluate the operand as sparse. Then, based on such an evaluation performed by the sparse data inspection device 103’, a multiplier engine 104’ processes instructions that operate on dense data in a dedicated dense data sub-pipeline. The multiplier engine 104 may be configured to process instructions that operate on sparse data in a separate dedicated sparse data sub-pipeline.

[0019] In some examples, a separate dedicated sparse data sub-pipeline can be configured to perform special operations configured for sparse data (e.g., operations configured to process sparse data with reduced energy consumption, latency, and / or memory usage). As a non-limiting example, in the SA for matrix multiplication based on MAC nodes, the sparse data sub-pipeline can strategically ignore instructions regarding sparse data. It should be understood that in matrix multiplication, the final result matrix is the sum of element-wise products from two input matrices. Thus, if any element of the input matrix is sparse (e.g., a zero-valued element or a very small element), all the element-wise products involving that element will be sparse (e.g., zero or near-zero values) and will not substantially contribute to the final sum in the result matrix. Thus, for matrix multiplication operations, the sparse data sub-pipeline can simply ignore instructions operating on sparse data. However, the multiplier engine 104’ can compute the accurate final result matrix because all instructions operating on dense data (which actually contribute to the final result) are processed in the dense data sub-pipeline. Other accelerated computations (e.g., other matrix computations) can have similar sparse properties, enabling the sparse data sub-pipeline to substantially ignore incoming instructions (e.g., by using no-op instructions to replace at least a portion of the incoming instructions). For example, neural network computations can be implemented as a sum of products similar to matrix multiplication (e.g., tensor multiplication), which can enable instructions operating on sparse data to be ignored because they may not contribute to the final result of the computation.

[0020] Relative to the sequential processing of dense and / or sparse matrices and relative to dedicated accelerators for dense or sparse matrices (e.g., relative to the simpler accelerator pipeline 100 for dense matrices), the accelerator pipeline 100’ can achieve reduced energy consumption. As indicated by the double arrows for multiplication instructions and memory instructions from the instruction issue stage 102’ to the sparse data check device 103’ and the memory engine 106’, the instruction issue bandwidth of the accelerator pipeline 100’ is increased (e.g., relative to previous accelerators), enabling the accelerator pipeline to process a large number of issued instructions. For example, the accelerator can simultaneously process dense instructions and sparse instructions in a dedicated dense sub-pipeline and a dedicated sparse sub-pipeline, respectively. In some examples, the accelerator pipeline can be configured to feed more instructions to the accelerator than the accelerator can process (e.g., the accelerator pipeline can “oversupply” the accelerator), allowing the accelerator to continue processing instructions without sparse instructions and buffering the remaining instructions into an effective on-chip memory for future processing (e.g., taking advantage of sparsity).

[0021] Figure 2An exemplary method for increasing the hardware efficiency of matrix computations in an accelerator pipeline (e.g., accelerator pipeline 100 or 100') is shown.

[0022] Figure 3 A non - limiting example of a micro - architecture 300 for an accelerator pipeline (e.g., for accelerator pipeline 100 or 100') is shown. The micro - architecture 300 is configured to increase the hardware efficiency of matrix computations including multiple matrix operations. As a non - limiting example, the micro - architecture 300 can be configured to implement method 200.

[0023] At 202, method 200 includes receiving, at a computer processing device, a digital signal encoding one or more operations of a matrix computation, each operation including one or more operands. For example, the digital signal can encode matrix multiplication and / or neural network computations, as described above. Non - limiting examples of digital signals for encoding matrix operations include computer - readable instructions, machine code, assembly code, bytecode, source code, etc.

[0024] At 204, method 200 includes determining, by a sparse - data checking device of a computer processing machine, whether the operations of the matrix computation include all dense operands. The operands of the matrix computation can be either sparse or dense. For example, sparse operands can include zero - valued operands and / or operands having a value less than a threshold. Thus, the sparse - data checking device can be configured to perform any suitable operations (e.g., arithmetic operations) to determine whether each operand of the operation is sparse or dense.

[0025] At 206, in response to determining that the operations of the matrix computation include all dense operands, method 200 includes forwarding the operations to a dense - computing device of the computer processing machine, the dense - computing device being configured to perform the operations of the matrix computation based on the dense operands. At 208, in response to determining that the operations of the matrix computation include one or more sparse operands, method 200 includes forwarding the operations to a sparse - computing device, the sparse - computing device being configured to perform the operations of the matrix computation. Thus, operations involving any sparse operands can be efficiently processed by the sparse - computing device. Additionally, all operations issued to the dense - computing device involve all - dense operations, so the dense - computing device does not unnecessarily use any computing resources to perform arithmetic operations associated with sparse results. For example, the dense - computing device can specifically perform arithmetic operations on non - zero data, thereby eliminating unnecessary latency and / or energy consumption associated with explicitly performing operations such as multiplying by zero (which always produces a zero product regardless of other operands) and / or adding zero (equivalent to no operation).

[0026] The microarchitecture 300 is configured to: process instructions received in the instruction issue stage 302 (e.g., digital signals encoding matrix operations), and ultimately write out the results in the commit stage 308. The microarchitecture 300 is configured to receive incoming instructions at the instruction issue stage 302. As indicated by the double arrows and like the accelerator pipeline 200, the bandwidth can accommodate the simultaneous processing of incoming multiplication instructions and / or memory instructions. The microarchitecture 300 can enable efficient accelerated matrix calculations, regardless of what program uses the microarchitecture 300 for acceleration (e.g., the microarchitecture 300 provides energy savings, latency, and / or throughput advantages independent of specific matrix processing, AI, ML, and / or neural network programs and / or data).

[0027] The method 200 and / or the microarchitecture 300 can be used for various matrix calculations. For example, the method 200 and / or the microarchitecture 300 can be used to improve the hardware efficiency of matrix multiplication (e.g., using a dense computing device configured to perform multiply-accumulate operations). As another non-limiting example, the method 200 and / or the microarchitecture 300 can be used to improve the hardware efficiency of neural network calculations.

[0028] The sparse data checking device 303 is configured to distinguish instructions regarding sparse data from instructions regarding dense data. For example, the sparse data checking device 303 can identify sparse data based on instructions involving one or more zero-valued operands, based on a predefined tag associated with the data and / or the instruction, or in any other suitable manner. As an example, a multiplication instruction involving one or more zero-valued operands (and optionally operands close to zero) can be considered sparse because multiplying 0 by any value results in a product of 0 (or close to zero, in the case of multiplying by an operand close to zero).

[0029] In some examples, the sparse data checking device 303 is configured to: determine whether an operand is sparse or dense based on determining that zero-valued operands are sparse and non-zero-valued operands are dense. In other examples, the sparse data checking device 303 is configured to: determine whether an operand is sparse or dense based on determining that an operand is dense if its value exceeds a predefined threshold and sparse if its value does not exceed the predefined threshold. For example, the sparse data checking device 303 can be configured to identify sparse data based on instructions involving one or more operands having values less than a sparse threshold.

[0030] As an example, the sparse threshold can be selected based on how an instruction is used in an overall computation. For example, the sparse threshold for matrix multiplication can be the minimum factor size, where it is expected that the smaller factor will produce product values that are unlikely to significantly affect the final result of the computation. For example, the sparse value can be a small floating point number such as 0.001, 0.00001, 0.0000001, etc. In some examples, the sparse value can be an adjustable value (e.g., a hardware hyperparameter of the sparse data checking device 303). For example, the sparse value can be redefined before and / or during the computation in order to reduce the energy, latency, and / or memory cost associated with the computation. As an example, the threshold sparse value can be optimized in order to reduce the energy cost by treating most of the instruction operands in the instruction operands as sparse, while being constrained based on the accuracy criteria of the computation. As a non-limiting example, for matrix multiplication computations, the threshold sparse value can be selected based on the expected distribution of the input data matrix in order to reduce the cost, while ensuring that there is at most a threshold deviation in the accuracy of the matrix product for matrices extracted from the expected distribution. As another non-limiting example, for neural network computations, the threshold sparse value can be selected based on treating as many computations as possible as sparse in order to reduce the cost, while maintaining at least a threshold average prediction accuracy. By adjusting an appropriate threshold sparse value, the sparse data checking device 303 can enable reduced costs (e.g., power, memory, and / or latency) for implementing various computations involving sparse operands. The above examples are non-limiting, and the threshold sparse value can be selected based on any appropriate criteria. As an example, in a matrix computation of accumulating maximum values, large negative values can be considered sparse because such large negative values are unlikely or less likely to be the maximum value. As another example, in a matrix computation of accumulating the product of multiple factors, values close to 1 can be considered sparse because multiplying by a value close to 1 does not significantly change the final product result. Instructions operating on sparse data are processed by the sparse computing device 304B together with the look-ahead sparse instruction queue (LASIQ) 305B and the sparse register identifier queue (SRIQ) 309B. Instructions operating on dense data are processed by the dense computing device 304A together with the look-ahead dense instruction queue (LADIQ) 305A.

[0031] Although the checking of data sparsity is described herein with respect to an operand being sparse when the operand is a zero value or near zero value, in other embodiments, the sparse data checking device 303 may be configured to: sort computer operations (e.g., steps such as matrix calculations) based on other suitable criteria regarding the operands of such operations. Thus, the sorted computations may be appropriately forwarded to a dense computing device configured to perform computer operations and may be appropriately forwarded to a sparse computing device configured to perform the same computer operations at a lower computational cost. For example, the sparse computing device may utilize the nature of the sorted data passed to the sparse computing device (similar to utilizing sparsity to more efficiently perform matrix calculations described herein).

[0032] In some examples, any instruction having all dense operands is enqueued into the LADIQ 305A to forward the instruction to the dense computing device 304A. Thus, the dense computing device 304A may be configured to execute instructions from the LADIQ 305A in order. In some examples, more instructions than the number of instructions that the dense computing device 304A is configured to process in subsequent cycles may be fed to the LADIQ 305A. For example, based on the instruction bandwidth, throughput, and / or latency of the computing resources associated with the dense computing device 304A, the dense computing device 304A may only be able to perform a limited number of arithmetic operations in one cycle. Thus, when the LADIQ 305A remains full and instructions are forwarded from the LADIQ 305A, the dense computing device 304A may process the maximum number of instructions in multiple subsequent cycles. For example, using the LADIQ 305A to provide instructions to the execution engine may enable the dense computing device 304A to complete multiple parts of a matrix calculation without any waiting that may be caused by receiving a digital signal specifying operand data when checking the sparsity of operand data in the sparse data checking device 303 and / or when processing sparse data by the sparse computing device 304A.

[0033] Similarly, in some examples, instructions are forwarded to the sparse computing device 304B by enqueuing any instruction having one or more sparse operands into the LASIQ 305B. In some examples, the sparse computing device 304B is configured to automatically store the sparse result values in the SRIQ 309B in program order. For example, the sparse computing device 304B can identify the location associated with the result value of each sparse instruction in order to automatically write zero values into the register file for each sparse instruction. In some examples, more instructions than the number of instructions that the sparse computing device is configured to process in subsequent cycles can be fed into the LASIQ 305B. For example, a large number of instructions having sparse results can be provided to the LASIQ 305B, and these instructions can utilize no-operation substitution and / or be used to automatically derive the location for storing the sparse result values. Accordingly, the computations associated with the sparse results can be skipped (e.g., using no-operation substitution) and / or delayed (e.g., enqueued for later writing of the sparse result values). Since a large number of instructions are processed by the sparse computing device 303, the LADIQ 305A and / or the LASIQ 305B can remain at and / or near full capacity, thereby allowing for efficient, concurrent execution of the sparse and / or dense portions of matrix computations.

[0034] The sparse data inspection device 303 is generally configured to process all available issued instructions to classify such instructions into a sparse instruction set and / or a dense instruction set for processing. Accordingly, the instructions can be immediately processed by the sparse computing device 304B or the dense computing device 304A. Alternatively or additionally, when the instructions are processed, such instructions can be enqueued into the LASIQ 305B and / or the LADIQ 305A for subsequent processing by the sparse computing device 304B and / or the dense computing device 304A, respectively. As will be described below, the LASIQ 305B, the LADIQ 305A, and the SRIQ 309B allow the sparse data inspection device 303, the sparse computing device 304B, and the dense computing device 304A to concurrently process multiple incoming instructions while ensuring that the results of the processed instructions appear in the correct order (e.g., in program order). In some examples, the sparse computing device 304B can be a shorter-latency pipeline relative to the dense computing device 304A pipeline (e.g., the sparse computing device 304B can be implemented using specialized processing techniques applicable to sparse matrices to avoid redundant / unnecessary computations on zero elements).

[0035] Compared with the dense computing device 304A, the sparse computing device 304B is a smaller / shorter pipeline (e.g., having fewer computing steps, lower power consumption, and / or lower latency). In some examples, the sparse computing device 304B can substantially ignore incoming instructions that have been detected as sparse (e.g., because such instructions on sparse data may not actually affect the final result of the overall computation).

[0036] For example, the sparse computing device 304B may not actually multiply multiple values. For example, the sparse computing device 304B may utilize no-operation instructions instead of multiplying sparse data. Even when no multiplication is performed, the sparse computing device 304B can determine where to store non-zero elements for later processing without performing any actual multiplication and while ignoring zero elements, so that the non-zero elements can be processed later. In other examples, the sparse computing device 304B may perform sparse-optimized multiplication to effectively multiply multiple values while ignoring zero elements. As an example, the sparse computing device 304B may be configured to ignore a particular operand of all incoming instructions, but for each incoming instruction, calculate a memory address at which a sparse value (e.g., a constant 0 value, which represents the expected result from an instruction regarding one or more sparse operands) is automatically written. Thus, the sparse computing device 304B may be able to effectively determine all such memory addresses at which the sparse value (e.g., the constant 0 value) is written without actually performing any arithmetic on the incoming operands. Thus, the sparse computing device 304B calculates all sparse values generated by instructions regarding sparse data while reducing energy, time, and / or memory costs because no specific results need to be processed. In some examples, it may not even be necessary to explicitly track the results of sparse instructions (e.g., for a sum of products, products equal to 0 may be completely omitted from the final sum). In any case, the sparse computing device 304B is configured to: when it is necessary to track any such results for an overall calculation, determine the order and / or location at which the results of instructions regarding sparse data are written. In some examples, the sparse computing device 304B is configured to: calculate intermediate results of a matrix calculation with one or more sparse operands by automatically assuming sparse result values. For example, in matrix multiplication, when a zero-valued element appears in the input matrix, based on having the zero-valued element as a factor, one or more intermediate results of the matrix calculation are necessarily zero-valued. Thus, the sparse computing device 304B may be configured to automatically calculate intermediate results associated with sparse operands by writing the zero-valued results to the associated locations in the product matrix representing the calculation result. For example, when a zero-valued operand appears, an entire row and / or entire column may be automatically assumed to have a sparse result. Thus, the sparse computing device 304B can efficiently and automatically write all such sparse results without performing specific calculations on the operands. In some examples, the sparse computing device 304B may be configured to utilize no-operation instructions instead of executable instructions for an operation. For example, an operation that performs a particular multiplication and / or addition step in a matrix multiplication operation may be replaced with a no-operation instruction because the addition of result values is mathematically equivalent to a no-operation.

[0037] For example, as Figure 3As shown, the sparse computing device 304B is configured to: within a sparse register identifier queue (SRIQ309B), track a sparse instruction identifier for each instruction regarding sparse data and an output register for that instruction. The SRIQ 309B is in turn configured to: forward a given instruction to the commit stage after all instructions issued before the given instruction have been committed. In addition, the SRIQ 309B is configured to forward the value of the corresponding register to the register file. In some embodiments, the SRIQ 309B guarantees that instructions retire (i.e., complete) in program order, thus ensuring the correct result of matrix calculations. In other words, the sparse computing device can be configured to forward instruction identifiers to the SRIQ 309B, and thus the SRIQ can be configured to write corresponding results according to the program order associated with the instruction identifiers. Therefore, the computational work of storing sparse values resulting from operations on sparse data can be skipped, delayed, and / or efficiently implemented by the SRIQ 309B.

[0038] The sparse data check device 303 is configured to forward instructions to the sparse computing device 304B and / or the dense computing device 304A when they are available for processing the instructions. For example, if a series of instructions issued during a cycle (e.g., all instructions in the cycle or some of the instructions in the cycle) are instructions regarding dense data, the first instruction (e.g., the "youngest" or most recently issued instruction) can be fed into the dense computing device 304A for processing. Additionally, any other instructions regarding dense data can be buffered into the LADIQ 305A. For example, the LADIQ 305A can be an on-chip queue of any appropriate size. Similarly, if a series of instructions issued during a cycle have sparse data values, the first instruction can be fed into the sparse computing device 304B, and other instructions regarding sparse data can be buffered into the LASIQ 305B. Like the LADIQ 305A, the LASIQ 305B can be an on-chip queue of any appropriate size. In subsequent cycles, the next instruction in the LADIQ 305A can be fed into the dense computing device 304A, and / or the next instruction in the LASIQ 305B can be fed into the sparse computing device 304A. Instructions are processed from the instruction issue stage 302 as long as there is space in both the LASIQ 305B and the LADIQ 305A. At the instruction issue stage 302, the accelerator microarchitecture 300 has not yet determined whether the issued instructions operate on dense data or sparse data. Thus, the accelerator microarchitecture uses the LASIQ 305B and the LADIQ 305A to queue incoming operations to ensure that any incoming operation can be processed and / or queued (e.g., even if a part of the pipeline such as the sparse computing device 304B or the dense computing device 304A is too busy to immediately process additional instructions). Therefore, if either the LASIQ 305B or the LADIQ 305A is completely full, the instruction issue stage 302 is configured to stop issuing new instructions until there is space in both queues.

[0039] As described above, the issue bandwidth for multiplication instructions can be increased, e.g., such that instructions regarding sparse data and dense data can be enqueued and / or processed concurrently. Additionally, the issue bandwidth for memory instructions can be increased similarly to facilitate loading data from memory for multiplication instructions. As a non-limiting example, the memory instruction bandwidth can be increased by increasing the bandwidth of the load / store device 306A of the memory engine 306. Alternatively or additionally, as another non-limiting example, the memory instruction bandwidth can be increased by adding one or more additional load / store devices 306B. In some examples, when more than two load / store devices are used, a load queue 307 can be used to enqueue memory load results to ensure that load instructions retire (e.g., complete) in program order. For example, although not typically used in accelerator hardware, a load queue can be used in a superscalar compute pipeline to ensure program-order loading of data from memory.

[0040] The methods and processes described herein can be bound to the computing system of one or more computing devices. In particular, such methods and processes can be implemented as executable computer applications, network-accessible computing services, application programming interfaces (APIs), libraries, or combinations of the foregoing and / or other computing resources. As another non-limiting example, the accelerator pipeline 100’ and / or the microarchitecture 300 can exist as (a) logical subsystem(s) (e.g., a processor and / or a coprocessor device) in a computing system for performing any suitable computations (e.g., for matrix multiplication).

[0041] Figure 4 A simplified representation of a computing system 400 is schematically illustrated, which is configured to provide any to all of the computing functions described herein. The computing system 400 can take the form of one or more personal computers, network-accessible server computers, tablet computers, home entertainment computers, gaming devices, mobile computing devices, mobile communication devices (e.g., smart phones), virtual / augmented / mixed reality computing devices, wearable computing devices, Internet of Things (IoT) devices, embedded computing devices, and / or other computing devices. As a non-limiting example, the computing system 400 can implement the accelerator pipeline 100 or the accelerator pipeline 100’. As another non-limiting example, the computing system 400 can include the microarchitecture 300. As another non-limiting example, the computing system 400 can implement the method 200.

[0042] The computing system 400 includes a logical subsystem 402 and a storage subsystem 404. The computing system 400 can optionally include an input / output subsystem 406, a communication subsystem 408, and / or Figure 4Other subsystems not shown. For example, the logic subsystem 402 may include an accelerator pipeline 100' and / or a microarchitecture 300.

[0043] The logic subsystem 402 includes one or more physical devices configured to execute instructions. For example, the logic subsystem may be configured to execute instructions as part of one or more applications, services, or other logical constructs. The logic subsystem may include one or more hardware processors configured to execute software instructions. Additionally or alternatively, the logic subsystem may include one or more hardware or firmware devices configured to execute hardware or firmware instructions. The processors of the logic subsystem may be single-core or multi-core, and the instructions executed thereon may be configured for sequential, parallel, and / or distributed processing. Optionally, the various components of the logic subsystem may be distributed among more than two separate devices, which may be remotely located and / or configured for coordinated processing. Aspects of the logic subsystem may be virtualized and executed by remotely accessible, networked computing devices configured in a cloud computing configuration.

[0044] The storage subsystem 404 includes one or more physical devices configured to temporarily and / or permanently hold computer information, such as data and instructions executable by the logic subsystem. When the storage subsystem includes more than two devices, the devices may be collocated and / or remotely located. The storage subsystem 404 may include volatile, non-volatile, dynamic, static, read / write, read-only, random access, sequential access, location-addressable, file-addressable, and / or content-addressable devices. The storage subsystem 404 may include removable and / or built-in devices. When the logic subsystem executes instructions, the state of the storage subsystem 404 may be transformed, for example, to hold different data.

[0045] Aspects of the logic subsystem 402 and the memory subsystem 404 may be integrated together into one or more hardware logic components. For example, such hardware logic components may include program and application specific integrated circuits (PASIC / ASIC), program and application specific standard products (PSSP / ASSP), system-on-chips (SOC), and complex programmable logic devices (CPLD). As a non-limiting example, the accelerator microarchitecture 300, the instruction issue stage 302, the sparse data inspection device 303, the LASIQ 305B, the LADIQ 305A, the sparse computing device 304B, the dense computing device 304A, the SRIQ 309B, the memory engine 306, the load store device 306A, the load store device 306B, the load queue 307, and / or the commit stage 308 may all be implemented via any suitable combination of hardware logic components. In some examples, the hardware accelerator may have limited on-chip memory capacity and low access latency (e.g., "fast memory"). If the workload of the accelerator exceeds the size of the on-chip memory, a remote slower memory may be used to buffer the irrelevant data. However, this may significantly reduce the speed of processing the workload. However, the techniques disclosed herein may enable efficient on-chip processing of matrix computations (e.g., by reducing latency, power consumption, and / or the on-chip storage requirements for maintaining zero values of sparse data), thereby mitigating the possible consequences of limited on-chip memory. However, it should be understood that the techniques described herein are beneficial even when a large amount of on-chip memory is available.

[0046] The logic subsystem and the storage subsystem may cooperate to instantiate one or more logical machines and / or engines. As used herein, the terms "machine" and "engine" are used interchangeably to refer to a combination of hardware, firmware, software, instructions, and / or any other components that cooperate to provide computer functionality. In other words, a "machine" and / or "engine" is never an abstract concept and always has a tangible form. A machine and / or engine may be instantiated by a single computing device or may include two or more sub-components instantiated by more than two different computing devices. In some embodiments, a machine and / or engine includes local components (e.g., a software application executed by a computer processor) that cooperate with remote components (e.g., cloud computing services provided by a network of server computers). The software and / or other instructions that give a particular machine and / or engine its functionality may optionally be saved as one or more unexecuted modules on one or more appropriate storage devices. As a non-limiting example, according to the present disclosure, accelerator microarchitecture 300, instruction issue stage 302, sparse data checking device 303, LASIQ 305B, LADIQ 305A, sparse computing device 304B, dense computing device 304A, SRIQ 309B, memory engine 306, load store device 306A, load store device 306B, load queue 307, and / or commit stage 308 may be implemented as a machine / engine.

[0047] The machine / engine can be implemented using any suitable combination of existing technologies and / or future machine learning (ML), artificial intelligence (AI), and / or natural language processing (NLP) technologies. Alternatively or additionally, the machine / engine described in this disclosure can be used as a component in the implementation of ML, AI, and / or NLP technologies. As a non-limiting example, matrix computations (e.g., matrix multiplication) can be used in the implementation of neural networks. Thus, the accelerator microarchitecture 300 can be used to efficiently perform such matrix computations. Non-limiting examples of ML, AI, and / or NLP technologies include support vector machines, multi-layer neural networks, convolutional neural networks (e.g., including spatial convolutional networks for processing images and / or videos, temporal convolutional neural networks for processing audio signals and / or natural language sentences, and / or any other suitable convolutional neural network configured to convolve and pool features across one or more time and / or space dimensions), recurrent neural networks (e.g., long short-term memory networks), associative memories (e.g., lookup tables, hash tables, Bloom filters, neural Turing machines, and / or neural random access memories), word embedding models (e.g., GloVe or Word2Vec), unsupervised spatial and / or clustering methods (e.g., nearest neighbor algorithms, topological data analysis, and / or k-means clustering), graphical models (e.g., (hidden) Markov models, Markov random fields, (hidden) conditional random fields, and / or AI knowledge bases), and / or natural language processing technologies (e.g., tokenization, stemming, constituency and / or dependency parsing, and / or intent recognition, segmentation models, and / or super-segmentation models (e.g., hidden dynamic models)).

[0048] In some examples, the methods and processes described herein can be implemented using one or more differentiable functions, where the gradient of the differentiable function can be computed and / or estimated with respect to the input and / or output of the differentiable function (e.g., with respect to training data, and / or with respect to an objective function). Such methods and processes can be determined at least in part by a set of trainable parameters. Thus, the trainable parameters of a particular method or process can be adjusted by any suitable training procedure to continuously improve the functionality of the method or process.

[0049] Non-limiting examples of training procedures for adjusting trainable parameters include supervised training (e.g., using gradient descent or any other suitable optimization method), zero-shot, few-shot, unsupervised learning methods (e.g., classification based on classes derived from unsupervised clustering methods), reinforcement learning (e.g., deep Q-learning based on feedback), and / or generative adversarial neural network training methods, belief propagation, RANSAC (Random Sample Consensus), contextual bandit methods, maximum likelihood methods, and / or expectation maximization. In some examples, multiple methods, processes, and / or components of the systems described herein can be simultaneously trained with respect to an objective function that measures the performance of the collective functionality of multiple components (e.g., with respect to reinforcement feedback and / or with respect to labeled training data). Simultaneously training multiple methods, processes, and / or components can improve such collective functionality. In some examples, one or more methods, processes, and / or components can be trained independently of other components (e.g., offline training on historical data).

[0050] This disclosure is presented by way of example and with reference to the associated drawings. Components, process steps, and other elements that are substantially the same in one or more of the figures may be identically identified and described with a minimum of repetition. However, it should be noted that the identically identified elements may also differ to some extent. It should also be noted that some of the figures may be schematic and not drawn to scale. The various drawing scales, aspect ratios, and numbers of components shown in the figures may be deliberately distorted to make certain features or relationships easier to see.

[0051] In one example, a method for increasing the computer hardware efficiency of matrix calculations includes: receiving, at a computer processing device, a digital signal encoding one or more operations of a matrix calculation, each operation including one or more operands; in response to a determination by a sparse data inspection device of the computer processing machine that the operations of the matrix calculation include all dense operands, forwarding the operations to a dense computing device of the computer processing machine, the dense computing device being configured to perform the operations of the matrix calculation based on the dense operands; and in response to a determination by the sparse data inspection device that the operations of the matrix calculation include one or more sparse operands, forwarding the operations to a sparse computing device, the sparse computing device being configured to perform the operations of the matrix calculation. In this example or any other example, the sparse data inspection device is configured to: determine whether an operand is sparse or dense based on determining that a zero-valued operand is sparse and a non-zero-valued operand is dense. In this example or any other example, the sparse data inspection device is configured to: determine whether an operand is sparse or dense based on determining that an operand is dense if the value of the operand exceeds a predefined threshold and the operand is sparse if the value of the operand does not exceed the predefined threshold. In this example or any other example, the predefined threshold is a hardware hyperparameter of the sparse data inspection device. In this example or any other example, the matrix calculation is matrix multiplication. In this example or any other example, the dense computing device is configured to perform multiply-accumulate operations. In this example or any other example, the matrix calculation is a neural network calculation. In this example or any other example, the sparse computing device is configured to automatically save the sparse result value to the location resulting from the operations of the matrix calculation. In this example or any other example, the sparse computing device is configured to replace the executable instructions of the operation with no-operation instructions. In this example or any other example, forwarding an operation with all dense operands to the dense computing device includes: enqueuing the operation with all dense operands into a look-ahead dense instruction queue, wherein the dense computing device is configured to execute the operations from the look-ahead dense instruction queue in order. In this example or any other example, the method further includes feeding the look-ahead dense instruction queue with more operations than the number of operations that the dense computing device is configured to process in subsequent cycles. In this example or any other example, forwarding an operation with one or more sparse operands to the sparse computing device includes: enqueuing the operation with one or more sparse operands into a look-ahead sparse instruction queue. In this example or any other example, the sparse computing device is configured to: automatically store the sparse result value from the operation in the look-ahead sparse instruction queue in the program order of the operation in the look-ahead sparse instruction queue. In this example or any other example, the method further includes feeding the look-ahead sparse instruction queue with more operations than the number of operations that the sparse computing device is configured to process in subsequent cycles.

[0052] In one example, a computer system for performing matrix calculations includes: a sparse computing device configured to compute the result of an operation having one or more sparse operands; a dense computing device configured to compute the result of an operation having all dense operands; an instruction issue stage configured to receive a digital signal encoding one or more operations of a matrix calculation; and a sparse data inspection device configured to: forward an operation having one or more sparse operands to the sparse computing device and forward an operation having all dense operands to the dense computing device. In this example or any other example, the sparse data inspection device is configured to: determine an operand as dense if the value of the operand is greater than a predefined threshold, and is configured to: determine an operand as sparse if the value of the operand is less than the predefined threshold. In one example, a computer system for performing matrix calculations includes: a sparse computing unit configured to perform matrix calculations on sparse data; a dense computing unit configured to perform matrix calculations on dense data; an instruction issue stage configured to receive a plurality of instructions, including an instruction for operating on sparse data and an instruction for operating on dense data; and a sparse data inspection device configured to distinguish between an instruction for operating on sparse data and an instruction for operating on dense data, wherein the sparse data inspection device is further configured to: forward an instruction determined to be for operating on sparse data to the sparse computing device and forward an instruction determined to be for operating on dense data to the dense computing device. In this example or any other example, the sparse data inspection device is configured to: detect a plurality of instructions for operating on sparse data and queue the plurality of instructions for operating on sparse data into a speculative sparse instruction queue. In this example or any other example, the sparse data inspection device is configured to: detect a plurality of instructions for operating on dense data and queue the plurality of instructions for operating on dense data into a speculative dense instruction queue. In this example or any other example, the sparse computing device is configured to forward an instruction identifier to a sparse register identifier queue, and the sparse register identifier queue is configured to write corresponding results according to the program order associated with the instruction identifier.

[0053] It should be understood that the configurations and / or methods described herein are exemplary in nature and these specific embodiments or examples should not be considered in a limiting sense as many variations are possible. The specific routines or methods described herein may represent one or more of any number of processing strategies. As such, the various acts shown and / or described may be performed in the sequence shown and / or described, in other sequences, in parallel, or omitted. Likewise, the order of the above processes may be changed.

[0054] The subject matter of this disclosure includes all novel and non-obvious combinations and sub-combinations of various processes, systems, and configurations, as well as other features, functions, acts, and / or properties disclosed herein, and any and all equivalents thereof.

Claims

1. A method for increasing the computer hardware efficiency of matrix calculations, comprising: At a computer processing device, receiving a digital signal encoding one or more operations of the matrix calculation, each operation including one or more operands; In response to the sparse data checking device of the computer processing device determining that the operations of the matrix calculation include all dense operands, forwarding the operations to the dense calculation device of the computer processing device, the dense calculation device being configured to perform the operations of the matrix calculation based on the dense operands; And In response to the sparse data checking device determining that the operations of the matrix calculation include one or more sparse operands, forwarding the operations to the sparse calculation device of the computer processing device, the sparse calculation device being configured to perform the operations of the matrix calculation.

2. The method according to claim 1, wherein the sparse data checking device is configured to: determine whether an operand is sparse or dense based on determining zero-valued operands as sparse and non-zero-valued operands as dense.

3. The method according to claim 1, wherein the sparse data checking device is configured to: determine whether an operand is sparse or dense based on determining the operand as dense if the value of the operand exceeds a predefined threshold and determining the operand as sparse if the value of the operand does not exceed the predefined threshold.

4. The method according to claim 3, wherein the predefined threshold is a hardware hyperparameter of the sparse data checking device.

5. The method according to claim 1, wherein the matrix calculation is matrix multiplication.

6. The method according to claim 1, wherein the dense calculation device is configured to perform multiplication and accumulation operations.

7. The method according to claim 1, wherein the matrix calculation is neural network calculation.

8. The method according to claim 1, wherein the sparse calculation device is configured to automatically save the sparse result value to the position resulting from the operations of the matrix calculation.

9. The method according to claim 1, wherein the sparse calculation device is configured to replace the executable instructions of the operation with no-operation instructions.

10. The method according to claim 1, wherein forwarding the operation having all dense operands to the dense computing device comprises: Queue the operations with all dense operands into a preceding dense instruction queue, wherein the dense calculation device is configured to execute the operations from the preceding dense instruction queue in sequence.

11. The method according to claim 10 further comprises: Feed the preceding dense instruction queue with more operations than the number of operations that the dense calculation device is configured to process in subsequent cycles.

12. The method according to claim 1, wherein forwarding the operation having one or more sparse operands to the sparse computing device comprises: Queue the operations with one or more sparse operands into a preceding sparse instruction queue.

13. The method according to claim 12, wherein the sparse calculation device is configured to: automatically store the sparse result values from the operations in the preceding sparse instruction queue in the program order of the operations in the preceding sparse instruction queue.

14. The method according to claim 12 further comprises: Feed the preceding sparse instruction queue with more operations than the number of operations that the sparse calculation device is configured to process in subsequent cycles.

15. A computer system for performing matrix calculations, comprising: A sparse computing device configured to compute the result of an operation having one or more sparse operands; A dense computing device configured to compute the result of an operation having all dense operands; An instruction issue stage configured to receive a digital signal encoding one or more operations of a matrix computation; And A sparse data inspection device configured to forward an operation having one or more sparse operands to the sparse computing device and an operation having all dense operands to the dense computing device.

16. The computer system according to claim 15, wherein the sparse data inspection device is configured to determine that an operand is dense if the value of the operand is greater than a predefined threshold, and the sparse data inspection device is configured to determine that the operand is sparse if the value of the operand is less than the predefined threshold.

17. A computer system for performing matrix computations, comprising: A sparse computing unit configured to perform matrix computations on sparse data; A dense computing unit configured to perform the matrix computations on dense data; An instruction issue stage configured to receive a plurality of instructions, the plurality of instructions including instructions for operating on sparse data and instructions for operating on dense data; And A sparse data inspection device configured to distinguish between instructions for operating on sparse data and instructions for operating on dense data, wherein the sparse data inspection device is further configured to: forward an instruction determined to be for operating on sparse data to the sparse computing unit and an instruction determined to be for operating on dense data to the dense computing unit.

18. The computer system according to claim 17, wherein the sparse data inspection device is configured to: detect a plurality of instructions for operating on sparse data and queue the plurality of instructions for operating on sparse data in a speculative sparse instruction queue.

19. The computer system according to claim 17, wherein the sparse data inspection device is configured to: detect a plurality of instructions for operating on dense data and queue the plurality of instructions for operating on dense data in a speculative dense instruction queue.

20. The computer system according to claim 17, wherein the sparse computing unit is configured to forward an instruction identifier to a sparse register identifier queue, and the sparse register identifier queue is configured to write corresponding results according to a program order associated with the instruction identifier.

Citation Information

Patent Citations

  • Efficient inner product operations

    US20190347256A1