Method for storing weight matrix, inference system and computer readable storage medium
By segmenting the sparse weight matrix into zero-value and non-zero-value sub-blocks, and assigning non-zero-value weights to cluster storage and calculations, the problem of reduced performance and high energy consumption of sparse weight matrix in artificial neural networks is solved, and more efficient inference performance and energy efficiency are achieved.
Patent Information
- Application Number
- CN202010048497.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-10
- Filing Date
- 2020-01-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-01-16
AI Technical Summary
When the sparse weight matrix is stored and calculated in artificial neural networks, there are problems of reduced performance and high energy consumption, especially because it contains a large number of zero-value coefficients, resulting in low efficiency of multiplication and addition operations.
The sparse weight matrix is divided into at least one first sub-block and a second sub-block, wherein the first sub-block contains only zero-value weights, the second sub-block contains non-zero-value weights, and the non-zero-value weights in the second sub-block are assigned to the cluster of the circuit for storage and calculation.
By segmenting and mapping non-zero value weights into the cluster, the inference performance of artificial neural networks is improved, and trivial calculations of multiplying by zero or adding zero are avoided in matrix-vector multiplication operations, which improves energy efficiency.
Smart Images

Figure CN111445004B_ABST
Abstract
Description
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS]
[0002] This application claims priority to and the benefits of U.S. Provisional Application No. 62 / 793,731 filed on January 17, 2019 and U.S. Non-Provisional Patent Application Serial No. 16 / 409,487 filed on May 10, 2019, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present disclosure generally relates to methods of storing weights of neuron synapses for a sparse artificial neural network. Background Art
[0004] Artificial neural networks (ANNs) are used for a variety of tasks, including image recognition, natural language processing, and various pattern matching and classification tasks. Generally speaking, an artificial neural network includes an input layer, an output layer, and one or more hidden layers, each of which includes a series of neurons. The output of each layer of neurons is connected to all neuron inputs of the next layer. Each connection between neurons has a "weight" associated with the connection. The activation of each neuron is calculated by performing a weighted sum of the inputs of the neuron and passing the linear combination of the weighted inputs to a threshold activation function using a transfer function. Therefore, the artificial neural network performs a matrix-vector multiplication (MVM) of the input vector multiplied by the weight matrix, then sums (e.g., a linear combination of the input signals), and then thresholds the sum by a comparator.
[0005] Once an artificial neural network is trained to perform a specific task, the artificial neural network can accurately predict the output when provided with input, a process known as artificial neural network inference. The weights of the trained artificial neural network can be stored locally to the neuron-neuron interconnections to quickly and energy efficiently perform multiplication and addition operations of the artificial neural network. For example, some prior art systems utilize analog memory elements to store neuron weights, where the conductance of the analog memory elements represents the weights. The higher the conductance, the higher the weight, and therefore the greater the impact of the neuron input utilizing this conductance. However, sparse weight matrices including a large number of zero-valued coefficients may reduce the performance of the inference process and may be energy inefficient due to performing trivial calculations such as multiplying or adding zeros. Summary of the invention
[0006] The present disclosure relates to various embodiments of a method for storing a sparse weight matrix for a trained artificial neural network in a circuit including a series of clusters. In one embodiment, the method includes partitioning the sparse weight matrix into at least one first sub-block and at least one second sub-block. The at least one first sub-block includes only zero-valued weights and the at least one second sub-block includes non-zero-valued weights. The method also includes assigning the non-zero-valued weights in the at least one second sub-block to at least one cluster in the series of clusters of the circuit. The circuit is configured to perform a matrix-vector multiplication (MVM) between the non-zero-valued weights of the at least one second sub-block and an input vector.
[0007] The method may include identifying a cluster in the series of clusters that is not assigned at least one non-zero-valued weight during said assigning of the non-zero-valued weights.
[0008] The method may include completely powering off (power gating) the clusters that are not assigned at least one non-zero valued weight.
[0009] The circuit may include an array of memristors.
[0010] Each of the memristors may be a resistive random access memory (RRAM), a conductive-bridging random access memory (CBRAM), a phase-change memory (PCM), a ferroelectric field effect transistor (FerroFET), or a spin-transfer torque random access memory (STT RAM).
[0011] Assigning the non-zero valued weight may include setting the resistance of each of the memristors utilizing a series of selectors connected in series with the memristors.
[0012] The sparse weight matrix may have a size of 512×512, the at least one first sub-block may have a size of 256×256, and the at least one second sub-block may have a size of 128×128, 64×64, or 32×32.
[0013] Partitioning the sparse weight matrix may include recursively comparing a size of the at least one second sub-block to a size of a smallest cluster in the series of clusters.
[0014] If the size of the at least one second sub-block is equal to the size of the minimum cluster, the method may also include: calculating a first energy cost for processing the non-zero valued weight using an unblocked element cluster comprising an unblocked element buffer and at least one digital arithmetic logic unit; calculating a second energy cost for processing the non-zero valued weight using the minimum cluster; determining a lower energy cost between the first energy cost and the second energy cost; and assigning the non-zero valued weight to the unblocked element cluster or the minimum cluster based on the lower energy cost.
[0015] If the size of the at least one second sub-block is larger than the size of the minimum cluster, the method may also include: further dividing the at least one second sub-block into a series of sub-blocks, the series of sub-blocks having a size matching the size of a first series of clusters in the series of clusters; calculating a first total energy cost for processing the non-zero value weights of each of the series of sub-blocks using the first series of clusters; calculating a second total energy cost for processing the non-zero value weights of the second sub-block using a single cluster having the same size as the second sub-block; determining the lower total energy cost of the first total energy cost and the second total energy cost; and assigning the non-zero value weights of the series of sub-blocks to the first series of clusters or assigning the non-zero value weights of the at least one second sub-block to the single cluster based on the lower total energy cost.
[0016] The present disclosure also relates to various embodiments of a system for performing inference using an artificial neural network with a sparse weight matrix. In one embodiment, the system includes a network-on-chip, the network-on-chip including a series of clusters, each cluster in the series of clusters including an array of memristor crossbars. In one embodiment, the system also includes a processor and a non-transitory computer-readable storage medium, in which instructions are stored, which, when executed by the processor, cause the processor to: partition the sparse weight matrix into at least one first sub-block and at least one second sub-block, the at least one first sub-block including only zero-valued weights and the at least one second sub-block including non-zero-valued weights; and assign the non-zero-valued weights in the at least one second sub-block to at least one cluster in the series of clusters of the circuit. The circuit is configured to perform a matrix-vector multiplication (MVM) between the non-zero-valued weights of the at least one second sub-block and an input vector.
[0017] The instructions, when executed by the processor, may also cause the processor to identify a cluster in the series of clusters that is not assigned at least one non-zero-valued weight.
[0018] The instructions, when executed by the processor, may also cause the processor to completely power off the clusters that are not assigned at least one non-zero valued weight.
[0019] Each memristor of the array of memristor crossbar switches may be a resistive random access memory (RRAM), a conductive bridging random access memory (CBRAM), a phase change memory (PCM), a ferroelectric field effect transistor (FerroFET), or a spin transfer torque random access memory (STT RAM).
[0020] The on-chip network may further include a series of selectors connected in series with the memristor crossbar, and the instructions, when executed by the processor, may further cause the processor to assign the non-zero valued weight by setting the resistance of the memristor crossbar using the selectors.
[0021] The sparse weight matrix may have a size of 512×512, the at least one first sub-block may have a size of 256×256, and the at least one second sub-block may have a size of 128×128, 64×64, or 32×32.
[0022] The instructions, when executed by the processor, may also cause the processor to recursively compare the size of the at least one second sub-block to a size of a smallest cluster in the series of clusters.
[0023] If the size of the at least one second sub-block is equal to the size of the minimum cluster, the instructions may also cause the processor to: calculate a first energy cost for processing the non-zero valued weight using an unblocked element cluster comprising an unblocked element buffer and at least one digital arithmetic logic unit; calculate a second energy cost for processing the non-zero valued weight using the minimum cluster; determine the lower energy cost of the first energy cost and the second energy cost; and assign the non-zero valued weight to the unblocked element cluster or the minimum cluster based on the lower energy cost.
[0024] If the size of the at least one second sub-block is greater than the size of the smallest cluster, the instructions, when executed by the processor, may also cause the processor to: further divide the at least one second sub-block into a series of sub-blocks, the series of sub-blocks having a size matching the size of a first series of clusters in the series of clusters; calculate a first total energy cost for processing the non-zero valued weights of each of the series of sub-blocks using the first series of clusters; calculate a second total energy cost for processing the non-zero valued weights of the second sub-block using a single cluster having the same size as the second sub-block; determine the lower total energy cost of the first total energy cost and the second total energy cost; and based on the lower total energy cost, assign the non-zero valued weights of the series of sub-blocks to the first series of clusters or assign the non-zero valued weights of the at least one second sub-block to the single cluster.
[0025] The present disclosure also relates to various embodiments of a non-transitory computer-readable storage medium. In one embodiment, the non-transitory computer-readable storage medium stores software instructions, which, when executed by a processor, cause the processor to: partition a sparse weight matrix of an artificial neural network into at least one first sub-block and at least one second sub-block, wherein the at least one first sub-block includes only zero-valued weights and the at least one second sub-block includes non-zero-valued weights; and assign the non-zero-valued weights in the at least one second sub-block to at least one cluster of an on-chip network including an array of memristor crossbar switches.
[0026] This summary is provided to introduce optional features and concepts of embodiments of the present disclosure, which are further described in the detailed description below. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. One or more of the described features may be combined with one or more other described features to provide a viable device. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] These and other features and advantages of the disclosed embodiments will become more apparent by referring to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same reference numerals are used throughout the figures to refer to the same features and components. These figures are not necessarily drawn to scale.
[0028] Figure 1 is a flowchart illustrating the tasks of a method for performing artificial neural network inference using a memristor accelerator according to one embodiment of the present disclosure.
[0029] Figure 2 FIG. 1 is a diagram showing a trained artificial neural network according to an embodiment of the present disclosure. The trained artificial neural network may be used in Figure 1 The method shown is used to perform artificial neural network inference.
[0030] Figure 3 Show according to Figure 1 The sparse weight matrices partitioned for the tasks shown in .
[0031] Figure 4 A series of clusters of a hardware accelerator according to one embodiment of the present disclosure are shown, each of which includes an array of memristor crossbar switches and can be used in the task of storing non-zero valued weights of a second sub-block (eg, sub-region).
[0032] Figure 5 A network-on-chip (NoC) according to one embodiment of the present disclosure is shown, wherein the NoC includes a series of computing nodes, a series of routing nodes connected to the computing nodes, and a series of data lines and address lines connected to the computing nodes and the routing nodes.
[0033] Figure 6 Draw Figure 5 One of the compute nodes of the NoC shown in is connected to one of the routing nodes.
[0034] [Explanation of Symbols]
[0035] 100: Methods;
[0036] 110, 120, 140, 150: tasks;
[0037] 130: Task / Matrix-Vector Multiplication (MVM) Task;
[0038] 200: Artificial neural networks;
[0039] 201: input layer;
[0040] 202: input layer neurons / neurons;
[0041] 203: hidden layer;
[0042] 204: Hidden layer neurons / neurons;
[0043] 205: output layer;
[0044] 206: output layer neurons / neurons;
[0045] 207, 208: connections / synapses;
[0046] 300: sparse weight matrix;
[0047] 301: first sub-block;
[0048] 302: second sub-block;
[0049] 303: Sub-area / third sub-area;
[0050] 304: Sub-area / fourth sub-area;
[0051] 305: Sub-area / fifth sub-area;
[0052] 400: cluster / minimum cluster / destination cluster;
[0053] 401: memristor crossbar switch / memristor / memristor crossbar switch array;
[0054] 402: Field effect transistor (FET);
[0055] 403: Sample and hold (S / H) array;
[0056] 404: Analog-to-digital converter (ADC);
[0057] 500: Network on Chip (NoC);
[0058] 501: computing node;
[0059] 502: routing node / NoC routing node;
[0060] 503: data line and address line;
[0061] 504, 505: address and data bus;
[0062] 506: bus interface unit. DETAILED DESCRIPTION
[0063] The present disclosure relates to various systems and methods for storing weight coefficients of a sparse weight matrix for a trained artificial neural network in a circuit and performing an artificial neural network inference process using the circuit. The circuit includes a series of clusters, and each cluster includes an array of memristor crossbar switches (e.g., resistive random access memory (RRAM), conductive bridging random access memory (CBRAM), phase change memory (PCM), ferroelectric field effect transistor (FerroFET) and spin transfer torque random access memory (STT RAM)). The memristor can be analog or digital (e.g., single bit or multiple bits). Each cluster is configured to perform analog or digital matrix-vector multiplication (MVM) between the weights of the sparse weight matrix and the input vector when data flows from one layer of the artificial neural network to another layer during inference. The systems and methods of the present disclosure include partitioning the sparse weight matrix into at least two sub-blocks, wherein at least one sub-block contains only zero weight coefficients, and then only mapping the sub-blocks containing non-zero weight coefficients to the array of memristor crossbar switches. In this way, the systems and methods of the present disclosure are configured to improve the performance of artificial neural network inference processes and are configured to be energy efficient by avoiding performing trivial calculations such as multiplying or adding zeros in MVM operations.
[0064] Hereinafter, exemplary embodiments will be described in more detail with reference to the accompanying drawings, and in all the drawings, the same reference numerals refer to the same elements. However, the present invention may be implemented in various forms and should not be considered to be limited to the embodiments shown herein. Specifically, these embodiments are provided as examples to make the present disclosure thorough and complete, and to fully convey the various aspects and features of the present invention to those skilled in the art. Therefore, processes, elements and technologies that are not necessary for those of ordinary skill in the art to fully understand the various aspects and features of the present invention may no longer be described. Unless otherwise noted, the same reference numerals throughout all the drawings and this written description represent the same elements, and therefore, they may no longer be described in detail.
[0065] In the accompanying drawings, the relative sizes of elements, layers and regions may be exaggerated and / or simplified for clarity. For ease of explanation, spatially relative terms such as "beneath", "below", "lower", "under", "above", "upper" and the like may be used herein to describe the relationship between one element or feature shown in the figure and another (other) element or feature. It should be understood that the spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientations shown in the figure. For example, if the device in the figure is turned over, the elements described as being "below", "below" or "below" other elements or features will now be oriented to be "above" the other elements or features. Therefore, the exemplary terms "below" and "below" may encompass both the above and below orientations. The device may have other orientations (e.g., rotated 90 degrees or in other orientations) and the spatially relative descriptors used herein should be interpreted accordingly.
[0066] It should be understood that although the terms "first", "second", "third", etc. may be used herein to describe various elements, components, regions, layers and / or sections, these elements, components, regions, layers and / or sections should not be limited to these terms. These terms are used to distinguish one element, component, region, layer or section from another element, component, region, layer or section. Therefore, the first element, component, region, layer or section described below may be referred to as a second element, component, region, layer or section without departing from the spirit and scope of the present invention.
[0067] It should be understood that when an element or layer is referred to as being "on," "connected to," or "coupled to" another element or layer, the element or layer may be directly on, directly connected to, or directly coupled to the other element or layer, or one or more intervening elements or layers may be present. In addition, it should be understood that when an element or layer is referred to as being "between" two elements or layers, the element or layer may be the only element or layer between the two elements or layers, or one or more intervening elements or layers may also be present.
[0068] The terms used herein are for the purpose of illustrating specific embodiments and are not intended to limit the present invention. Unless the context clearly indicates otherwise, the singular form "a and an" used herein is intended to also include plural forms. It should also be understood that when the terms "comprises, comprising" and "includes, including" are used in this specification, it is to indicate the existence of the stated features, integers, steps, operations, elements and / or components, but does not exclude the existence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. The term "and / or" used herein includes any and all combinations of one or more items in the relevant listed items. For example, expressions such as "at least one of..." modify the elements of the entire series when they are located before a series of elements rather than modifying the individual elements in the series of elements.
[0069] As used herein, the terms "substantially," "about," and the like are used as terms of approximation, not as terms of degree, and are intended to take into account the inherent deviations of measurements or calculations that one of ordinary skill in the art would know. In addition, the use of "may" when describing embodiments of the present invention refers to "one or more embodiments of the present invention." As used herein, the terms "use," "using," and "used" may be considered synonymous with the terms "utilize," "utilizing," and "utilized," respectively. In addition, the term "exemplary" is intended to refer to an example or illustration.
[0070] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms (such as those defined in commonly used dictionaries) should be interpreted as having a meaning consistent with their context in the relevant technology and / or this specification, and should not be interpreted as having an idealized or overly formal meaning unless explicitly defined herein.
[0071] Figure 11 is a flowchart illustrating the tasks of a method 100 for performing artificial neural network inference using a memristor accelerator according to one embodiment of the present disclosure. In the illustrated embodiment, the method 100 includes a task 110 of partitioning a sparse weight matrix of a trained artificial neural network into at least one first sub-block containing only zero-valued weight coefficients and at least one second sub-block containing non-zero-valued weight coefficients. The task 110 of partitioning the sparse weight matrix may be performed by any suitable one or more processes, such as, for example, diagonalization of the sparse weight matrix. Diagonalization of the sparse weight matrix includes determining eigenvalues of the sparse weight matrix. As used herein, the term "sparse weight matrix" refers to a weight matrix in which most or approximately most of the weight coefficients are zero or substantially zero.
[0072] Figure 2 FIG. 2 shows a trained artificial neural network 200 according to an embodiment of the present disclosure. The artificial neural network 200 can be used to Figure 1 In the method 100 for performing artificial neural network inference shown. In the illustrated embodiment, the artificial neural network 200 includes an input layer 201 having a series of input layer neurons 202, at least one hidden layer 203 including a series of hidden layer neurons 204, and an output layer 205 having a series of output layer neurons 206. The output of each of the input layer neurons 202 is connected to each of the hidden layer neurons 204 through a series of connections or synapses 207, and the output of each of the hidden layer neurons 204 is connected to each of the output layer neurons 206 through a series of connections or synapses 208. Each connection 207, 208 between neurons 202, 204, 206 has a "weight" associated therewith, and the weights of the connections 207, 208 from one layer of the artificial neural network 200 to the next layer are stored in a weight matrix. In one or more embodiments, the artificial neural network 200 may have any other suitable number of layers, and each layer may include any other suitable number of neurons, depending on, for example, the type of task that the artificial neural network 200 is trained to perform. In the illustrated embodiment, the sparse weight matrix partitioned according to the task 110 may be a weight matrix associated with the connection 207 between the input layer 201 and the hidden layer 203 or the connection 208 between the hidden layer 203 and the output layer 205.
[0073] Figure 3 Show according to Figure 1The sparse weight matrix 300 is partitioned according to the task 110 shown. In the illustrated embodiment, the sparse weight matrix 300 has a size of 1024×1024, but in one or more embodiments, depending on the number of neurons in each layer of the artificial neural network 200, the sparse weight matrix 300 may have any other suitable size. In addition, in the illustrated embodiment, after task 110, the sparse weight matrix 300 has been partitioned into three first sub-blocks 301, each of which includes only zero-valued weight coefficients, and each of which has a size of 512×512, but in one or more embodiments, depending on the total size of the sparse weight matrix 300 and the sparsity of the sparse weight matrix 300, the sparse weight matrix 300 may be partitioned into any other number of first sub-blocks 301 and each of the first sub-blocks 301 may have any other suitable size.
[0074] In the illustrated embodiment, after task 110, the sparse weight matrix 300 has been partitioned into a second sub-block 302, the second sub-block 302 including non-zero weight coefficients and having a size of 512×512. In one or more embodiments, the second sub-block 302 may have any other suitable size depending on the total size of the sparse weight matrix 300 and the sparsity of the sparse weight matrix 300. In addition, in the illustrated embodiment, after task 110 of partitioning the sparse weight matrix 300, the second sub-block 302 has been sub-partitioned (i.e., further sub-partitioned) into two sub-blocks 303, seven sub-blocks 304, and four sub-blocks 305, the two sub-blocks 303 each including non-zero weight coefficients and each having a size of 256×256, the seven sub-blocks 304 each including non-zero weight coefficients and each having a size of 128×128, and the four sub-blocks 305 each including non-zero weight coefficients and each having a size of 64×64. The sub-regions 303, 304, 305 together form the second sub-block 302. In one or more embodiments, the second sub-block 302 may be further partitioned into any other suitable number of sub-regions having any suitable size. As described in more detail below, the size of the sub-regions 303, 304, 305 may be selected based on the size of the clusters within the accelerator (i.e., the size of the array of memristor crossbars) and the energy cost associated with storing non-zero valued weights of the sub-regions 303, 304, 305 in these clusters.
[0075] Continue to refer to Figure 1In the illustrated embodiment, method 100 also includes a task 120 of storing the weights in the second sub-block 302 in one or more clusters, each of which includes an array of memristor crossbars (i.e., task 120 includes mapping the non-zero weight coefficients in the second sub-block 302 to one or more arrays of memristor crossbars). In one or more embodiments, task 120 includes mapping the non-zero valued weights of the entire second sub-block 302 to a single cluster (i.e., a single array of memristor crossbars) having the same size as the second sub-block (e.g., 512×512). In one or more embodiments, task 110 includes further partitioning the second sub-block 302 into two or more sub-blocks (e.g., Figure 3 In one or more embodiments of the sub-regions 303, 304, 305 shown in FIG. 1 , task 120 may include storing the non-zero valued weights of the sub-regions 303, 304, 305 in separate clusters (i.e., separate arrays of memristor crossbar switches) having the same size as the sub-regions (e.g., 256×256; 128×128 and / or 64×64). That is, task 120 may include, for example: storing the non-zero valued weights of the two sub-regions 303 in two separate clusters, each of which has the same size as the sub-region 303; storing the non-zero valued weights of the seven sub-regions 304 in seven separate clusters, each of which has the same size as the sub-region 304; and storing the non-zero valued weights of the four sub-regions 305 in four separate clusters, each of which has the same size as the sub-region 305.
[0076] In one or more embodiments, the method 100 also includes a task of identifying a cluster (i.e., an array of memristor crossbar switches) that is not assigned at least one non-zero valued weight during the task 120 of storing weights in the second sub-block 302 in one or more clusters. In one or more embodiments, before or during the execution of artificial neural network inference by the cluster, the method 100 may include completely powering off (e.g., completely power gating) the cluster that is not assigned at least one non-zero valued weight during the task 120. Completely powering off the cluster that is not assigned at least one non-zero valued weight during the task 120 can reduce the energy required to perform artificial neural network inference by approximately 10 times compared to a method that does not power off clusters that only have zero valued weights.
[0077] Figure 4A series of clusters 400 of a hardware accelerator each including an array of memristor crossbars 401 are shown, which can be used in task 120 to store non-zero valued weights for a second sub-block (e.g., sub-blocks 303, 304, 305) according to one embodiment of the present disclosure. The weights for the second sub-block 302 can be stored in the memristor crossbars 401 in the array. The weights stored by each memristor crossbar 401 are proportional to the conductance (i.e., the inverse of the resistance) of the memristor crossbar 401. Thus, in one or more embodiments, the task 120 of storing the non-zero weights for the second sub-block 302 includes programming the resistance (i.e., the inverse of the conductance) of the memristor crossbars 401 of one or more clusters 400 (i.e., one or more arrays of memristor crossbars 401) to correspond to the non-zero weights for the second sub-block 302 (e.g., task 120 includes programming each memristor crossbar 401 to a resistance level that is inversely proportional to the value of one or more bits of the corresponding weight matrix coefficient). The conductance of each memristor crossbar 401 (and thus the weight stored in the memristor crossbar 401) can be programmed in task 120 using a selector such as, for example, a diode or field effect transistor (FET) 402 connected in series to the memristor crossbar 401. The memristor crossbar 401 can be any suitable type or kind of memristor such as, for example, a resistive random access memory (RRAM), a conductive bridge random access memory (CBRAM), a phase change memory (PCM), a ferroelectric field effect transistor (FerroFET), or a spin transfer torque random access memory (STT RAM). RRAM can be used as an analog memory or a digital memory and STT RAM can be used as a single-bit per cell digital memory. Thus, the memristor crossbar 401 acts as a two or three terminal non-volatile synaptic weight in analog form or digital form (single bit or multiple bits). In one or more embodiments, cluster 400 includes an array of memristor crossbar switches 401 having two or more different sizes (eg, an accelerator includes cluster 400 having two or more different sizes).
[0078] like Figure 4 As shown in FIG. 4 , each cluster 400 further includes a sample and hold array (S / H array) 403 connected to each array of the memristor crossbar 401 and an analog-to-digital converter (ADC) 404 configured to convert the dot product calculated by each array of the memristor crossbar 401 from analog to digital.
[0079] In one or more embodiments, the task 120 of storing weights may include storing the weights in two or more clusters 400 having different sizes (i.e., the task 120 may include storing the non-zero weights in two or more arrays of memristor crossbars having different sizes). Figure 3 In the embodiment of the partitioned sparse weight matrix 300 shown, the task 120 of storing weights may include: the task of storing the weights in the sub-region 303 in two clusters 400, each of which has an array of memristor cross switches 401 of size 256×256; the task of storing the weights in the sub-region 304 in seven clusters 400, each of which has an array of memristor cross switches 401 of size 128×128; and the task of storing the weights in the sub-region 305 in four clusters 400, each of which has an array of memristor cross switches 401 of size 64×64.
[0080] In one or more embodiments, the task 110 of partitioning the sparse weight matrix 300 includes the task of recursively comparing the size of the second sub-block 302 of the partitioned sparse weight matrix with the size of the smallest cluster 400 of the hardware accelerator (i.e., the smallest array of memristor crossbar switches 401). If the size of the second sub-block 302 is the same as the size of the smallest cluster 400 implemented by the hardware accelerator, the method 100 includes the following tasks: using an unblocked element cluster of the hardware accelerator to calculate a non-zero energy cost for processing within the second sub-block 302; comparing the calculated energy cost with the energy cost of operating the smallest cluster 400; and assigning the weight of the second sub-block 302 to the cluster 400 (e.g., analog or unblocked) that exhibits a lower energy cost. In one or more embodiments, the unblocked element cluster includes an unblocked element buffer and one or more digital arithmetic logic units (ALUs). If the size of the second sub-block 302 is larger than the size of the smallest cluster 400 (i.e., the smallest array of memristor crossbars 401) implemented on the hardware accelerator, the method 100 includes further splitting the second sub-block 302 into smaller sub-blocks (e.g., Figure 3303, the fourth sub-block 304, and the fifth sub-block 305 shown in ). In addition, if the size of the second sub-block 302 is larger than the size of the smallest cluster 400 implemented on the hardware accelerator, the method 100 includes the task of calculating the total energy cost of operating each of the clusters 400 corresponding to the sub-blocks (e.g., summing the non-zero energy costs of processing the sub-blocks 303, 304, 305 using separate clusters 400 having arrays of memristors 401 corresponding to the sizes of the sub-blocks 303, 304, 305) and comparing the sum of the energy costs with the energy cost of mapping the entire second sub-block 302 to a single cluster 400 having an array of memristor crossbars 401 of the same size as the second sub-block 302, and partitioning the second sub-block 302 and storing the weights according to the method that exhibits the lower energy cost. That is, if the second sub-block 302 is larger than the smallest cluster 400 (i.e., if the second sub-block 302 is larger than the cluster 400 having the smallest array of memristor crossbars 401), the method 100 may include dividing the second sub-block 302 into smaller sub-blocks (e.g., sub-blocks 303, 304, 305) and mapping the non-zero weight coefficients in the sub-blocks to a memristor crossbar array 401 having the same size as the sub-blocks, or the method may include not dividing the second sub-block 302 and mapping the non-zero weight coefficients in the second sub-block 302 to a single array of memristor crossbars 401 having the same size as the second sub-block 302, depending on which method achieves a lower energy cost.
[0081] Continue to refer to Figure 1 In the illustrated embodiment, method 100 also includes a task 130 of performing a digital or analog matrix-vector multiplication (MVM) between the weight matrix and the corresponding input vector. In one or more embodiments, task 130 is performed using Figure 4 The MVM task 130 may be performed by clusters 400 of arrays of memristor crossbar switches 401, wherein each cluster 400 is configured to perform an MVM operation between a weight matrix and an input vector. Figure 4The array of voltage signals is performed by applying an array of voltage signals to the clusters shown. The array of voltage signals corresponds to the inputs of the neurons 202, 204, 206 in the artificial neural network 200. When the array of voltage signals is applied to the rows of the array of memristor crossbar switches 401, the current measured at each column of the array is the sum of the products of the input and the conductance of the corresponding memristor crossbar switch 401 in this column (for example, the current at each column is a weighted sum of the input voltages). In this way, the array of memristor cross switches 401 automatically performs multiplication and accumulate (MAC) operations between input vectors (which represent the inputs of neurons 202, 204, 206 in a given layer 201, 203, 205) and the conductances of the memristor cross switches 401 (which represent the weights of the connections 207, 208 between neurons 202, 204, 206) and calculates output vectors representing the outputs of neurons 202, 204, 206 in a given layer 201, 203, 205 of the artificial neural network 200.
[0082] Still refer to Figure 1 In the illustrated embodiment, method 100 further includes adding the output vector obtained by task 130 executing MVM using one cluster 400 to a task 140 of running vector sum sent from a different cluster 400 (e.g., task 140 includes a task 140 of adding the output vector obtained by task 130 executing MVM using one cluster 400 to a ...). Figure 4 In the illustrated embodiment, the method 100 further includes a task 150 of calculating a set of neural network activation functions based on the task 140 of calculating the running vector sum and sending the result to the destination cluster (e.g., task 150 includes calculating the activation function of the artificial neural network 200 based on the running vector sum, wherein the running vector sum is calculated based on the result of the MVM performed by the array of memristor crossbar switches 401 in the illustrated cluster 400). Figure 4 4. The calculation is based on the results of MVM performed by the array of memristor crossbar switches 401 in the cluster 400 shown.
[0083] Figure 5 A network on chip (NoC) 500 according to one embodiment of the present disclosure is shown, and the network on chip 500 includes a series of computing nodes 501, a series of routing nodes 502 connected to the computing nodes 501, and a series of data lines and address lines 503 connected to the computing nodes 501 and the routing nodes 502. In one or more embodiments, the computing nodes 501 may be connected via a modular NoC fabric and may be configured using mature direct memory access (DMA) configuration technology. Figure 6 Draw Figure 5One of the computing nodes 501 of the NoC 500 is shown connected to one of the routing nodes 502. Figure 6 As shown, each of the computing nodes 501 includes a series of clusters 400, each of which has an array of memristor crossbar switches 401 (eg, each of the computing nodes 501 includes Figure 4 401) and a series of selectors (e.g., Figure 4 The series of selectors are used to program the conductance (i.e., weight) stored by each memristor crossbar switch 401. Figure 6 As shown, the NoC 500 further includes a pair of address buses 504 and data buses 505 for each pair of computing nodes 501 and routing nodes 502, respectively. Input vector coefficients and output vector coefficients and their destination addresses are broadcasted by the cluster 400 in the computing node 501 or by the routing node 502 of the NoC serving the computing node 501 on the pair of address buses 504 and data buses 505. Figure 6 In the illustrated embodiment, each computing node 501 includes a bus interface unit 506 configured to receive input vector coefficients and addresses of the input vector coefficients. The bus interface unit 506 is configured (e.g., programmed using control logic) to direct the input vector coefficients to the appropriate memristor crossbar 401 inputs.
[0084] Available Figure 4 to Figure 5 The NoC 500 is shown to perform task 140 of calculating the running vector sum according to the result of the MVM performed by the cluster 400 in task 130. In addition, the method 100 may include sending the calculated running vector sum calculated in task 140 and the activation function calculated in task 150 to the destination cluster 400 located on the same computing node 501 or a different computing node 501.
[0085] The method of the present disclosure may be performed by a processor that executes instructions stored in a non-volatile memory. The term "processor" used herein includes any combination of hardware, firmware, and software for processing data or digital signals. The hardware of the processor may include, for example, an application specific integrated circuit (ASIC), a general or dedicated central processor (CPU), a digital signal processor (DSP), a graphics processor (GPU), and a programmable logic device such as a field programmable gate array (FPGA). In the processor used herein, each function is performed by hardware configured (i.e., hardwired) to perform the function, or by more general hardware (e.g., CPU) configured to execute instructions stored in a non-temporary storage medium. The processor may be fabricated on a single printed wiring board (PWB) or distributed on several interconnected PWBs. The processor may include other processors; for example, the processor may include two processors FPGA and CPU interconnected on a PWB.
[0086] Although the present invention has been described in detail with particular reference to exemplary embodiments of the present invention, the exemplary embodiments described herein are not intended to be exhaustive or to limit the scope of the present invention to the exact forms disclosed. It should be understood by those skilled in the art and technology to which the present invention belongs that the described structures, assemblies and methods of operation may be modified and altered without substantially departing from the principles, spirit and scope of the present invention as set forth in the above claims and their equivalents.
Claims
1. A method for storing a sparse weight matrix for a trained artificial neural network in a circuit comprising a plurality of clusters, the method comprising: partitioning the sparse weight matrix into at least one first sub-block and at least one second sub-block, the at least one first sub-block comprising only zero-valued weights and the at least one second sub-block comprising non-zero-valued weights; as well as assigning the non-zero valued weight in the at least one second sub-block to at least one cluster of the plurality of clusters of the circuit, wherein the circuit is configured to perform a matrix-vector multiplication between the non-zero valued weights of the at least one second sub-block and an input vector, and Wherein the at least one first sub-block is not assigned to the plurality of clusters of the circuit.
2. The method according to claim 1, further comprising: A cluster of the plurality of clusters that was not assigned at least one non-zero-valued weight during the assigning of the non-zero-valued weights is identified.
3. The method according to claim 2, further comprising: The clusters not assigned at least one non-zero valued weight are completely powered off. The method of claim 1 , wherein each cluster of the plurality of clusters comprises an array of memristors.
5. The method of claim 4, wherein the memristor is selected from the group consisting of a resistive random access memory, a conductive bridging random access memory, a phase change memory, a ferroelectric field effect transistor, a spin transfer torque random access memory, and combinations thereof. 6 . The method of claim 4 , wherein assigning the non-zero valued weight comprises setting a resistance of each of the memristors using a plurality of selectors connected in series with the memristors.
7. The method according to claim 1, wherein: The sparse weight matrix has a size of 512×512, The at least one first sub-block has a size of 256×256, and The at least one second sub-block has a size selected from the group consisting of 128×128, 64×64, and 32×32.
8. The method of claim 1, wherein partitioning the sparse weight matrix comprises recursively comparing a size of the at least one second sub-block with a size of a smallest cluster among the plurality of clusters.
9. The method according to claim 8, wherein: If the size of the at least one second sub-block is equal to the size of the smallest cluster, the method further comprises: calculating a first energy cost for processing the non-zero valued weight using an unblocked element cluster including an unblocked element buffer and at least one digital arithmetic logic unit; calculating a second energy cost of processing the non-zero valued weight using the minimum cluster; determining a lower energy cost of the first energy cost and the second energy cost; and The non-zero valued weight is assigned to the non-blocked element cluster or the minimum cluster depending on the lower energy cost.
10. The method according to claim 8, wherein: If the size of the at least one second sub-block is greater than the size of the smallest cluster, the method further comprises: further partitioning the at least one second sub-block into a plurality of sub-blocks, the plurality of sub-blocks having sizes matching sizes of a first plurality of clusters in the plurality of clusters; calculating a first total energy cost for processing the non-zero valued weights for each of the plurality of sub-regions using the first plurality of clusters; calculating a second total energy cost for processing the non-zero-valued weights of each of the at least one second sub-block using a single cluster having the same size as each of the at least one second sub-block; determining a lower total energy cost of the first total energy cost and the second total energy cost; and The non-zero valued weights of the plurality of sub-regions are assigned to the first plurality of clusters or the non-zero valued weight of the at least one second sub-block is assigned to the single cluster according to the lower total energy cost.
11. A system for performing inference using an artificial neural network having a sparse weight matrix, the system comprising: A network on chip comprising a plurality of clusters, each cluster of the plurality of clusters comprising an array of memristor crossbar switches; processor; as well as A non-transitory computer-readable storage medium having stored therein instructions that, when executed by the processor, cause the processor to: partitioning the sparse weight matrix into at least one first sub-block and at least one second sub-block, the at least one first sub-block comprising only zero-valued weights and the at least one second sub-block comprising non-zero-valued weights; as well as assigning the non-zero valued weight in the at least one second sub-block to at least one cluster of the plurality of clusters of circuits, wherein the circuit is configured to perform a matrix-vector multiplication between the non-zero valued weights of the at least one second sub-block and an input vector, and Wherein the at least one first sub-block is not assigned to the plurality of clusters of the circuit.
12. The system of claim 11, wherein the instructions, when executed by the processor, further cause the processor to identify clusters of the plurality of clusters that are not assigned at least one non-zero valued weight.
13. The system of claim 12, wherein the instructions, when executed by the processor, further cause the processor to completely power off the clusters that are not assigned at least one non-zero valued weight.
14. The system of claim 11, wherein each memristor of the array of memristor crossbars is selected from the group consisting of a resistive random access memory, a conductive bridge random access memory, a phase change memory, a ferroelectric field effect transistor, and a spin transfer torque random access memory.
15. The system of claim 11, wherein the on-chip network further comprises a plurality of selectors connected in series with the memristor cross switch, and wherein the instructions, when executed by the processor, further cause the processor to assign the non-zero value weight by setting the resistance of the memristor cross switch using the plurality of selectors.
16. The system of claim 11, wherein: The sparse weight matrix has a size of 512×512, The at least one first sub-block has a size of 256×256, and The at least one second sub-block has a size selected from the group consisting of 128×128, 64×64, and 32×32.
17. The system of claim 11, wherein the instructions, when executed by the processor, further cause the processor to recursively compare the size of the at least one second sub-block with a size of a smallest cluster among the plurality of clusters.
18. The system of claim 17, wherein if the size of the at least one second sub-block is equal to the size of the smallest cluster, the instructions further cause the processor to: calculating a first energy cost for processing the non-zero valued weight using an unblocked element cluster including an unblocked element buffer and at least one digital arithmetic logic unit; calculating a second energy cost of processing the non-zero valued weight using the minimum cluster; determining a lower energy cost of the first energy cost and the second energy cost; as well as The non-zero valued weight is assigned to the non-blocked element cluster or the minimum cluster depending on the lower energy cost.
19. The system of claim 17, wherein: If the size of the at least one second sub-block is greater than the size of the smallest cluster, the instructions, when executed by the processor, further cause the processor to: further partitioning the at least one second sub-block into a plurality of sub-blocks, the plurality of sub-blocks having sizes matching sizes of a first plurality of clusters in the plurality of clusters; calculating a first total energy cost for processing the non-zero valued weights for each of the plurality of sub-regions using the first plurality of clusters; calculating a second total energy cost for processing the non-zero-valued weights of each of the at least one second sub-block using a single cluster having the same size as each of the at least one second sub-block; determining a lower total energy cost of the first total energy cost and the second total energy cost; as well as The non-zero valued weights of the plurality of sub-regions are assigned to the first plurality of clusters or the non-zero valued weight of the at least one second sub-block is assigned to the single cluster according to the lower total energy cost.
20. A non-transitory computer-readable storage medium having software instructions stored therein, the software instructions, when executed by a processor, causing the processor to: partitioning a sparse weight matrix of the artificial neural network into at least one first sub-block and at least one second sub-block, the at least one first sub-block comprising only zero-valued weights and the at least one second sub-block comprising non-zero-valued weights; and assigning the non-zero valued weights in the at least one second sub-block to at least one cluster of a network-on-chip comprising an array of memristor crossbar switches, The at least one first sub-block is not assigned to the network on chip.
Citation Information
Patent Citations
System and method for signal processing
CN109214508A