Performing XNOR equivalent operations by adjusting column thresholds of in-memory computing arrays
By adjusting the column threshold of the calculation array in memory and calculating the conversion bias current reference, the problem of XNOR operation in the prior art increases area and power consumption is solved, and a smaller and lower power consumption in-memory computing system is realized.
Patent Information
- Application Number
- CN202080055739.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-09
- Filing Date
- 2020-09-08
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-09-08
AI Technical Summary
When performing XNOR operations of binary neural networks, existing in-memory computing systems increase large area and power consumption, and eliminating XNOR operations still requires maintaining the basic transformation of the binary network.
The activation threshold of each column is adjusted by adjusting the column threshold of the calculation array in memory, adjusting the activation threshold of each column based on the functions of weight and activation values, and computing the conversion bias current reference based on the input value in the input vector to determine the threshold of the output value.
It realizes the basic transformation of binary neural network without using XNOR function, reduces the size and power consumption of memory, and improves area efficiency.
Smart Images

Figure CN114207628B_ABST
Abstract
Description
[0001] Priority Claim under 35 U.S.C.§119
[0002] This patent application claims priority to U.S. Non - Provisional Application No. 16 / 565,308, filed on September 9, 2019, entitled "PERFORMING XNOR EQUIVALENT OPERATIONS BY ADJUSTING COLUMN THRESHOLDS OF A COMPUTE - IN - MEMORY ARRAY", which is assigned to the assignee of this application and is hereby incorporated by reference in its entirety.
[0003] Field
[0004] Aspects of the present disclosure generally relate to performing XNOR equivalent operations by adjusting column thresholds of a compute - in - memory array of an artificial neural network. Background Art
[0005] Ultra - low bit - width neural networks, such as binary neural networks (BNNs), are a powerful new approach in deep neural networking (DNN). Binary neural networks can significantly reduce data traffic and save power. For example, compared to floating - point / fixed - point precision, the memory storage for binary neural networks is significantly reduced because both weights and neuron activations are binarized to - 1 or + 1.
[0006] However, digital complementary metal - oxide - semiconductor (CMOS) processing uses a [0,1] basis. To perform the binary implementations associated with these binary neural networks, the [-1,+1] basis of the binary network should be transformed into the CMOS [0,1] basis. This transformation employs computationally intensive exclusive - NOR (XNOR) operations.
[0007] Compute - in - memory systems can implement ultra - low bit - width neural networks, such as binary neural networks (BNNs). A compute - in - memory system has memory with some processing capabilities. For example, each intersection of a bit line and a word line represents a filter weight value, and this weight value is multiplied by the input activation on the word line to generate a product. Subsequently, the individual products along each bit line are summed to generate the corresponding output value of the output tensor. This implementation can be regarded as a multiply - accumulate (MAC) operation. These MAC operations can transform the [-1,+1] basis of the binary network into the CMOS [0,1] basis.
[0008] Conventionally, this transformation under an in-memory computing system is achieved by performing an XNOR operation at each bit cell. Subsequently, the results along each bit line are summed to generate the corresponding output value. Unfortunately, including the XNOR function in each bit cell consumes a relatively large area and increases power consumption.
[0009] In a conventional implementation, each bit cell includes the basic memory functions of read and write plus additional logic functions that perform XNOR between the input and the cell state. As a result of including the XNOR capability, the number of transistors in each cell in a memory (e.g., static random access memory (SRAM)) increases from 6 or 8 to 12, which significantly increases the cell size and power consumption. It would be desirable to eliminate the XNOR operation while still being able to transform from a binary neural network [-1, +1] basis to a CMOS [0, 1] basis.
[0010] Overview
[0011] In one aspect of the present disclosure, an apparatus includes an in-memory computing array including columns and rows. The in-memory computing array is configured to: adjust an activation threshold generated for each column of the in-memory computing array based on a function of a weight value and an activation value. The in-memory computing array is further configured to: calculate a conversion bias current reference based on an input value in an input vector to the in-memory computing array. The in-memory computing array is programmed with a set of weight values. The adjusted activation threshold and the conversion bias current reference are used as thresholds for determining an output value of the in-memory computing array.
[0012] Another aspect discloses a method for performing an XNOR-equivalent operation by adjusting column thresholds of an in-memory computing array having rows and columns. The method includes: adjusting an activation threshold generated for each column of the in-memory computing array based on a function of a weight value and an activation value. The method further includes: calculating a conversion bias current reference based on an input value in an input vector to the in-memory computing array. The in-memory computing array is programmed with a set of weight values. The adjusted activation threshold and the conversion bias current reference are used as thresholds for determining an output value of the in-memory computing array.
[0013] In another aspect, a non-transitory computer-readable medium records non-transitory program code. The non-transitory program code, when executed by one or more processors, causes the one or more processors to adjust an activation threshold generated for each column of a memory-based computing array having rows and columns based on a function of weight values and activation values. The program code also causes the one or more processors to calculate a conversion bias current reference based on input values in an input vector to the memory-based computing array. The memory-based computing array is programmed with a set of weight values. The adjusted activation threshold and the conversion bias current reference are used as thresholds for determining output values of the memory-based computing array.
[0014] In another aspect, an apparatus for performing XNOR-equivalent operations by adjusting column thresholds of a memory-based computing array having rows and columns is disclosed. The apparatus includes: means for adjusting an activation threshold generated for each column of the memory-based computing array based on a function of weight values and activation values. The apparatus also includes: means for calculating a conversion bias current reference based on input values in an input vector to the memory-based computing array. The memory-based computing array is programmed with a set of weight values. The adjusted activation threshold and the conversion bias current reference are used as thresholds for determining output values of the memory-based computing array.
[0015] The features and technical advantages of the present disclosure have been outlined broadly in order that the detailed description that follows may be better understood. Additional features and advantages of the present disclosure will be described hereinafter. Those skilled in the art should appreciate that the present disclosure can be readily used as a basis for modifying or designing other structures for carrying out the same purposes as the present disclosure. Those skilled in the art should also recognize that such equivalent constructions do not depart from the teachings of the present disclosure as set forth in the appended claims. The novel features believed to be characteristic of the present disclosure, both as to its organization and method of operation, together with further objects and advantages, will be better understood from the following description when considered in connection with the accompanying drawings. It is to be expressly understood, however, that each of the drawings is provided for the purpose of illustration and description only and is not intended as a definition of the limits of the present disclosure. Brief Description of the Drawings
[0017] The features, nature, and advantages of the present disclosure will become more apparent upon consideration of the following detailed description, taken in conjunction with the accompanying drawings, in which like reference numerals refer to like elements throughout.
[0018] Figure 1 An example implementation of designing a neural network using a system-on-chip (SOC), including a general-purpose processor, in accordance with certain aspects of the present disclosure is illustrated.
[0019] Figure 2A 、 2B and 2C are diagrams illustrating neural networks in accordance with aspects of the present disclosure.
[0020] Figure 2D FIG. is a diagram illustrating an exemplary deep convolutional network (DCN) in accordance with aspects of the present disclosure.
[0021] Figure 3 FIG. is a diagram illustrating an exemplary deep convolutional network (DCN) in accordance with aspects of the present disclosure.
[0022] Figure 4 Illustrates an in-memory computing (CIM) array architecture of an artificial neural network in accordance with aspects of the present disclosure.
[0023] Figure 5 Illustrates an architecture for performing XNOR equivalent operations by adjusting column thresholds of an in-memory computing array of an artificial neural network in accordance with aspects of the present disclosure.
[0024] Figure 6 Illustrates a method for performing XNOR equivalent operations by adjusting column thresholds of an in-memory computing array of an artificial neural network in accordance with aspects of the present disclosure.
[0025] DETAILED DESCRIPTION
[0026] The detailed description set forth below in connection with the appended drawings is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described herein may be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of various concepts. However, it will be apparent to those skilled in the art that these concepts may be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form to avoid obscuring such concepts.
[0027] Based on this teaching, those skilled in the art will appreciate that the scope of the present disclosure is intended to cover any aspect of the present disclosure, whether implemented independently of or in combination with any other aspect of the present disclosure. For example, any number of the aspects described may be used to implement an apparatus or practice a method. Additionally, the scope of the present disclosure is intended to cover such apparatus or methods practiced using other structures, functionality, or a combination of structures and functionality that supplement or are different from the various aspects of the present disclosure as described. It should be understood that any aspect of the present disclosure disclosed may be implemented by one or more elements of a claim.
[0028] The term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" is not necessarily to be construed as superior to or better than other aspects.
[0029] Although specific aspects are described herein, numerous variations and permutations of these aspects fall within the scope of the present disclosure. While some benefits and advantages of the preferred aspects are mentioned, the scope of the present disclosure is not intended to be limited to specific benefits, uses, or objectives. Rather, the aspects of the present disclosure are intended to be broadly applicable to different technologies, system configurations, networks, and protocols, some of which are illustrated by way of example in the figures and the following description of the preferred aspects. The detailed description and the figures merely illustrate the present disclosure and do not limit the present disclosure, the scope of which is defined by the appended claims and their equivalent technical solutions.
[0030] In-memory computing (CIM) is a method of performing multiply-accumulate (MAC) operations in a memory array. In-memory computing can improve parallelism within the memory array by activating multiple rows and performing multiplication and summation operations using analog column currents. For example, SRAM bit cells can be customized to implement XNOR and bit counting operations of a binary neural network.
[0031] Conventionally, in-memory computing binary neural network implementations are achieved by performing XNOR operations at each bit cell and summing the results for each bit cell. Adding the XNOR function in each bit cell increases the layout area and increases the power consumption. For example, the number of transistors in each cell of a memory (e.g., SRAM) increases from 6 or 8 to 12.
[0032] Aspects of the present disclosure relate to performing XNOR-equivalent operations by adjusting the column thresholds of an in-memory computing array of an artificial neural network (e.g., a binary neural network). In one aspect, the activation threshold of each column of the memory array is adjusted based on a function of the weight value and the activation value. A conversion bias current reference is calculated based on the input values from the input vector.
[0033] In one aspect, the total bit line count is compared with the sum of the conversion bias current reference and the adjusted activation threshold to determine the output of the bit line. The total bit line count is the sum of the outputs of each bit cell corresponding to the bit lines of the memory array. For example, the sum (or total count) of the outputs of each bit cell associated with a first bit line is provided as a first input to a comparator. Subsequently, the total count is compared with the sum of the conversion bias current reference and the adjusted activation threshold to determine the output of the bit line. In some aspects, the activation threshold is less than half of the number of rows of the memory array. The number of rows corresponds to the size of the input vector. In some aspects, the conversion bias current reference is less than half of the number of rows of the memory array.
[0034] The artificial neural networks of the present disclosure can be binary neural networks, multi-bit neural networks, or extremely low-bitwidth neural networks. Aspects of the present disclosure can be applicable to devices that specify extremely low memory processing and power (e.g., edge devices), or large networks that can benefit from memory savings due to binary formats. Aspects of the present disclosure reduce the memory size and improve the memory power consumption by eliminating the XNOR operations in the in-memory computing system for implementing binary neural networks. For example, a fundamental transformation is done to avoid using the XNOR function and its corresponding transistor(s) in each bit cell, thereby reducing the memory size.
[0035] Figure 1 An example implementation of a system-on-chip (SOC) 100 is illustrated, which may include a central processing unit (CPU) 102 or a multi-core CPU configured for transforming multiplication and accumulation operations of an in-memory computing (CIM) array for an artificial neural network according to certain aspects of the present disclosure. Variables (e.g., neural signals and synaptic weights), system parameters associated with the computing device (e.g., neural network with weights), latency, frequency slot information, and task information may be stored in the memory block associated with the neural processing unit (NPU) 108, the memory block associated with the CPU 102, the memory block associated with the graphics processing unit (GPU) 104, the memory block associated with the digital signal processor (DSP) 106, the memory block 118, or may be distributed across multiple blocks. Instructions executed at the CPU 102 may be loaded from the program memory associated with the CPU 102 or may be loaded from the memory block 118.
[0036] The SOC 100 may further include additional processing blocks customized for specific functions such as the GPU 104, the DSP 106, the connectivity block 110 (which may include fifth-generation (5G) connectivity, fourth-generation long-term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and a multimedia processor 112 that can detect and recognize poses, for example. In one implementation, the NPU is implemented in the CPU, DSP, and / or GPU. The SOC 100 may further include a sensor processor 114, an image signal processor (ISP) 116, and / or a navigation module 120 (which may include a global positioning system).
[0037] The SOC 100 may be based on the ARM instruction set. In one aspect of the present disclosure, the instructions loaded into the general-purpose processor 102 may include code for adjusting the activation threshold of each column of the array based on a function of the weight values (e.g., weight matrix) and activation values. The general-purpose processor 102 may further include code for calculating a conversion bias current reference based on the input values from the input vector.
[0038] Deep learning architectures can perform object recognition tasks by learning to represent the input at successively higher levels of abstraction in each layer, thereby constructing useful feature representations of the input data. In this way, deep learning addresses the main bottlenecks of conventional machine learning. Before the advent of deep learning, machine learning approaches to object recognition problems might have relied heavily on human-engineered features, perhaps combined with shallow classifiers. Shallow classifiers can be two-class linear classifiers, for example, where the weighted sum of the feature vector components is compared to a threshold to predict which class the input belongs to. Human-engineered features can be templates or kernels customized by engineers with domain expertise for a specific problem domain. In contrast, deep learning architectures can learn to represent features similar to those that human engineers might design, but it learns through training. Additionally, deep networks can learn to represent and recognize new types of features that humans might not have considered yet.
[0039] Deep learning architectures can learn a hierarchy of features. For example, if visual data is presented to the first layer, the first layer can learn to identify relatively simple features (such as edges) in the input stream. In another example, if auditory data is presented to the first layer, the first layer can learn to identify spectral power in specific frequencies. A second layer that takes the output of the first layer as input can learn to identify combinations of features, such as identifying simple shapes for visual data or combinations of sounds for auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Even higher layers can learn to identify common visual objects or spoken phrases.
[0040] Deep learning architectures may perform particularly well when applied to problems with a natural hierarchical structure. For example, the classification of motor vehicles can benefit from first learning to identify wheels, windshields, and other features. These features can be combined in different ways at higher levels to identify cars, trucks, and airplanes.
[0041] Neural networks can be designed with a variety of connectivity patterns. In a feedforward network, information is passed from lower layers to higher layers, where each neuron in a given layer communicates to neurons in the higher layer. As described above, hierarchical representations can be constructed in successive layers of a feedforward network. Neural networks can also have recurrent or feedback (also known as top-down) connections. In a recurrent connection, the output from a neuron in a given layer can be communicated to another neuron in the same layer. Recurrent architectures can help identify patterns that span more than one chunk of input data presented sequentially to the neural network. Connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. Networks with many feedback connections can be beneficial when the recognition of high-level concepts can assist in discriminating specific low-level features of the input.
[0042] The connections between the layers of a neural network can be fully connected or locally connected. Figure 2A An example of a fully connected neural network 202 is illustrated. In the fully connected neural network 202, a neuron in the first layer can convey its output to every neuron in the second layer, such that every neuron in the second layer will receive input from every neuron in the first layer. Figure 2B An example of a locally connected neural network 204 is illustrated. In the locally connected neural network 204, a neuron in the first layer can be connected to a limited number of neurons in the second layer. More generally, the locally connected layer of the locally connected neural network 204 can be configured such that each neuron in a layer will have the same or similar connectivity pattern, although the connection strengths can have different values (e.g., 210, 212, 214, and 216). The locally connected connectivity pattern can give rise to spatially distinct receptive fields in higher layers, in that neurons in higher layers in a given region can receive input that has been tuned through training to be the nature of a restricted portion of the total input to the network.
[0043] An example of a locally connected neural network is a convolutional neural network. Figure 2C An example of a convolutional neural network 206 is illustrated. The convolutional neural network 206 can be configured such that the connection strengths associated with the input to each neuron in the second layer are shared (e.g., 208). Convolutional neural networks can be well-suited for problems where the spatial location of the input is meaningful.
[0044] One type of convolutional neural network is a deep convolutional network (DCN). Figure 2D A detailed example of a DCN 200 designed to identify visual features of an image 226 input from an image capture device 230 (such as an on-vehicle camera) is illustrated. The DCN 200 of this example can be trained to identify traffic signs and the numbers provided on traffic signs. Of course, the DCN 200 can be trained for other tasks, such as identifying lane markings or identifying traffic signals.
[0045] The DCN 200 can be trained using supervised learning. During training, an image (such as image 226 of a speed limit sign) can be presented to the DCN 200, and then a "forward pass" can be computed to produce an output 222. The DCN 200 can include a feature extraction section and a classification section. Upon receiving the image 226, the convolutional layer 232 can apply a convolutional kernel (not shown) to the image 226 to generate a first set of feature maps 218. As an example, the convolutional kernel of the convolutional layer 232 can be a 5x5 kernel that generates 28x28 feature maps. In this example, since four different feature maps are generated in the first set of feature maps 218, four different convolutional kernels are applied to the image 226 at the convolutional layer 232. The convolutional kernel can also be referred to as a filter or a convolutional filter.
[0046] The first set of feature maps 218 can be subsampled by a max pooling layer (not shown) to generate a second set of feature maps 220. The max pooling layer reduces the size of the first set of feature maps 218. That is, the size of the second set of feature maps 220 (such as 14x14) is smaller than the size of the first set of feature maps 218 (such as 28x28). The reduced size provides similar information to subsequent layers while reducing memory consumption. The second set of feature maps 220 can be further convolved via one or more subsequent convolutional layers (not shown) to generate one or more subsequent sets of feature maps (not shown).
[0047] In Figure 2D the example of, the second set of feature maps 220 is convolved to generate a first feature vector 224. Additionally, the first feature vector 224 is further convolved to generate a second feature vector 228. Each feature of the second feature vector 228 can include a number corresponding to a possible feature of the image 226 (such as "sign", "60", and "100"). A softmax function (not shown) can convert the numbers in the second feature vector 228 into probabilities. Thus, the output 222 of the DCN 200 is the probability that the image 226 includes one or more features.
[0048] In this example, the probabilities of "sign" and "60" in the output 222 are higher than the probabilities of other numbers (such as "30", "40", "50", "70", "80", "90", and "100") in the output 222. Before training, the output 222 produced by the DCN 200 is likely to be incorrect. Thus, the error between the output 222 and the target output can be computed. The target output is the ground truth of the image 226 (e.g., "sign" and "60"). The weights of the DCN 200 can then be adjusted so that the output 222 of the DCN 200 aligns more closely with the target output.
[0049] To adjust the weights, the learning algorithm can compute the gradient vector for the weights. This gradient can indicate the amount by which the error will increase or decrease if the weights are adjusted. At the top layer, this gradient can directly correspond to the value of the weights connecting the activated neurons in the penultimate layer to the neurons in the output layer. In the lower layers, the gradient can depend on the value of the weights and the computed error gradients of the higher layers. The weights can then be adjusted to reduce the error. This way of adjusting the weights can be called "backpropagation" because it involves a "backward pass" in the neural network.
[0050] In practice, the error gradient of the weights may be computed on a small number of examples, so that the computed gradient approximates the true error gradient. This approximation method can be called stochastic gradient descent. Stochastic gradient descent can be repeated until the error rate that the whole system can achieve has stopped decreasing or until the error rate has reached the target level. After learning, the DCN can be presented with a new image (e.g., the speed limit sign of Image 226) and the forward pass in the network can produce an output 222, which can be considered as the inference or prediction of the DCN.
[0051] A deep belief network (DBN) is a probabilistic model that includes multiple layers of hidden nodes. The DBN can be used to extract a hierarchical representation of the training dataset. The DBN can be obtained by stacking multiple layers of restricted Boltzmann machines (RBMs). An RBM is a type of artificial neural network that can learn the probability distribution on the input set. Since an RBM can learn the probability distribution without information about which class each input should be classified into, RBMs are often used in unsupervised learning. Using a hybrid unsupervised and supervised paradigm, the bottom RBM of the DBN can be trained in an unsupervised manner and can be used as a feature extractor, while the top RBM can be trained in a supervised manner (on the joint distribution of the input from the previous layer and the target classes) and can be used as a classifier.
[0052] A deep convolutional network (DCN) is a network of convolutional networks, which is configured with additional pooling and normalization layers. The DCN has achieved state-of-the-art performance on many tasks. The DCN can be trained using supervised learning, where both the input and the output targets are known for many paradigms and are used to modify the weights of the network by using the gradient descent method.
[0053] The DCN can be a feedforward network. Additionally, as described above, the connections from the neurons in the first layer of the DCN to the groups of neurons in the next higher layer are shared across the neurons in the first layer. The feedforward and shared connections of the DCN can be utilized for fast processing. The computational burden of the DCN can be much smaller than that of, for example, a neural network of a similar size that includes recurrent or feedback connections.
[0054] The processing of each layer of a convolutional network can be considered as a spatially invariant template or basis projection. If the input is first decomposed into multiple channels, such as the red, green, and blue channels of a color image, then a convolutional network trained on that input can be considered three-dimensional, having two spatial dimensions along the axes of the image and a third dimension capturing color information. The output of a convolutional connection can be considered to form a feature map in subsequent layers, where each element in the feature map (e.g., 220) receives input from a certain range of neurons in the previous layer (e.g., feature map 218) and from each of the multiple channels. The values in the feature map can be further processed with a non-linearity such as rectified max(0,x). Values from adjacent neurons can be further pooled (which corresponds to downsampling) and can provide additional local invariance as well as dimensionality reduction. Normalization can also be applied through lateral inhibition between neurons in the feature map, which corresponds to whitening.
[0055] The performance of deep learning architectures can improve as more labeled data points become available or as computing power increases. Modern deep neural networks are routinely trained with thousands of times more computing resources than were available to a typical researcher just fifteen years ago. New architectures and training paradigms can further boost the performance of deep learning. Rectified linear units can reduce a training problem known as vanishing gradients. New training techniques can reduce over-fitting and thus enable larger models to achieve better generalization. Encapsulation techniques can abstract the data within a given receptive field and further enhance overall performance.
[0056] Figure 3 is a block diagram illustrating a deep convolutional network 350. The deep convolutional network 350 can include multiple different types of layers based on connectivity and weight sharing. As Figure 3 shown, the deep convolutional network 350 includes convolutional blocks 354A, 354B. Each of the convolutional blocks 354A, 354B can be configured with a convolutional layer (CONV) 356, a normalization layer (LNorm) 358, and a max pooling layer (Max Pool) 360.
[0057] The convolutional layer 356 can include one or more convolutional filters, which can be applied to the input data to generate a feature map. Although only two convolutional blocks 354A, 354B are shown, the present disclosure is not limited thereto, but rather, any number of convolutional blocks 354A, 354B can be included in the deep convolutional network 350 according to design preferences. The normalization layer 358 can normalize the output of the convolutional filters. For example, the normalization layer 358 can provide whitening or lateral inhibition. The max pooling layer 360 can provide spatially downsampled aggregation to achieve local invariance and dimensionality reduction.
[0058] For example, the parallel filter bank of the deep convolutional network can be loaded onto the CPU 102 or GPU 104 of the SOC 100 to achieve high performance and low power consumption. In an alternative embodiment, the parallel filter bank can be loaded onto the DSP 106 or ISP 116 of the SOC 100. Additionally, the deep convolutional network 350 can access other processing blocks that may be present on the SOC 100, such as the sensor processor 114 and navigation module 120 dedicated to sensors and navigation, respectively.
[0059] The deep convolutional network 350 may further include one or more fully connected layers 362 (FC1 and FC2). The deep convolutional network 350 may further include a logistic regression (LR) layer 364. Between each layer 356, 358, 360, 362, 364 of the deep convolutional network 350 are weights (not shown) that are to be updated. The output of each layer (e.g., 356, 358, 360, 362, 364) can be used as the input to a subsequent layer (e.g., 356, 358, 360, 362, 364) in the deep convolutional network 350 to learn hierarchical feature representations from the input data 352 (e.g., images, audio, video, sensor data, and / or other input data) provided at the first convolutional block 354A. The output of the deep convolutional network 350 is a classification score 366 for the input data 352. The classification score 366 can be a set of probabilities, where each probability is the probability that the input data includes a feature from the feature set.
[0060] When weights and neuron activations are binarized to -1 or +1 ([−1, +1] space), the memory storage of an artificial neural network (e.g., a binary neural network) can be significantly reduced. However, digital complementary metal-oxide-semiconductor (CMOS) logic operates in the [0, 1] space. Thus, during binary implementation, a transformation occurs between [0, 1]-based digital CMOS devices and [-1, +1]-based binary neural networks.
[0061] Memory cells can be configured to support the exclusive NOR (XNOR) function. For example, Table 1-3 (e.g., a truth table) illustrates the mapping of a binary neural network in the [0, 1] space to binary multiplication in the binarized [-1, +1] space. A two-input logic function is illustrated in Truth Table 1-3.
[0062] Table 1 illustrates an example of binary multiplication in the binarized [-1, +1] space. For example, multiplication in the binarized [-1, +1] space produces a 4-bit output -1 * -1, -1 * +1, +1 * -1, +1 * +1 (e.g., "1, -1, -1, 1").
[0063] Multiply -1 +1 -1 +1 -1 +1 -1 +1
[0064] Table 1
[0065] Table 2 illustrates an example of the XNOR implementation. A memory cell can be configured to perform an XNOR function on a first input value (e.g., a binary neuron activation) and a second input value (e.g., a binary synaptic weight) to generate a binary output. For example, the XNOR function is true only if all input values are true or all input values are false. If some inputs are true and others are false, the output of the XNOR function is false. Thus, when both of these inputs (e.g., the first input and the second input) are false (e.g., the first input is 0 and the second input is 0) (as shown in Table 2), the output is true (e.g., 1). When the first input is false (0) and the second input is true (1), the output is false (0). When the first input is true (1) and the second input is false (0), the output is false (0). When the first input is true (1) and the second input is true (1), the output is true (1). Thus, the truth table of the XNOR function for two inputs has a binary output of "1, 0, 0, 1".
[0066] Thus, binary multiplication in the binary [-1, +1] space maps to the binary output of XNOR in the [0, 1] space. For example, a "1" in the 4-bit output of the binary [-1, +1] space maps to a "1" in the binary output of the XNOR function, while a "-1" in the 4-bit output of the binary [-1, +1] space maps to a "0" in the binary output of the XNOR function.
[0067] XNOR 0 1 0 1 0 1 0 1
[0068] Table 2
[0069] Multiply 0 1 0 0 0 1 0 1
[0070] Table 3
[0071] In contrast, binary multiplication in the binary [-1, +1] space does not map to binary multiplication in the binary [0, 1] space shown in Table 3. For example, Table 3 illustrates multiplying in the binary [0, 1] space to produce a 4-bit output of "0, 0, 0, 1", which does not map to the 4-bit output of "1, -1, -1, 1" in the binary [-1, +1] space. For example, the 4-bit output of "0, 0, 0, 1" includes only one true bit (e.g., the last bit), while the 4-bit output of "1, -1, -1, 1" includes two true bits (e.g., the first bit and the last bit).
[0072] Conventionally, a binary neural network implemented with an in-memory computing system is implemented by computing XNOR at each bit cell and summing the results along each bit line to generate an output value. However, adding the XNOR function in each bit cell is expensive. For example, the number of transistors per cell in a memory (e.g., SRAM) increases from 6 or 8 to 12, which significantly increases the cell size and power consumption.
[0073] Aspects of the present disclosure relate to reducing the size of a memory and improving the power consumption of the memory by eliminating XNOR in an in-memory computing array of a binary neural network. In one aspect, the activation threshold of each column of the in-memory computing array is adjusted to avoid using the XNOR function and its corresponding transistor(s) in each bit cell. For example, a smaller memory with smaller memory bit cells (e.g., eight-transistor SRAM) can be used for in-memory computing binary neural networks.
[0074] Figure 4 An exemplary architecture 400 for an in-memory computing (CIM) array for an artificial neural network according to aspects of the present disclosure is illustrated. In-memory computing is a way of performing multiplication and accumulation operations in a memory array. The memory array includes word lines 404 (or WL 1 ), 405 (or WL 2 ),..., 406 (or WL M ) and bit lines 401 (or BL 1 ), 402 (or BL 2 ),..., 403 (or BL N ). Weights (e.g., binary synaptic weight values) are stored in the bit cells of the memory array. Input activations (e.g., input values, which can be input vectors) are on the word lines. Multiplication occurs at each bit cell, and the multiplication results are output via the bit lines. For example, multiplication includes multiplying the weights by the input activations at each bit cell. A summing device (not shown) associated with a bit line or column (e.g., 401), such as a voltage / current summing device, sums the outputs (e.g., charge, current, or voltage) of the bit lines and passes the result (e.g., the output) to an analog-to-digital converter (ADC). For example, the sum of each bit line is calculated according to the corresponding outputs of the bit cells of each bit line.
[0075] In one aspect of the present disclosure, activation threshold adjustments are made at each column (corresponding to the bit lines) rather than at each bit cell to improve area efficiency.
[0076] Conventionally, starting from the implementation of a binary neural network where weights and neuron activations are binarized to -1 or +1 ([-1, +1] space), multiplication and accumulation operations become XNOR operations (in bit cells), which have an XNOR result and an overall count of the XNOR results. For example, the overall count of a bit line includes the sum of the positive (e.g., "1") results of each bit cell in that bit line. The ADC acts like a comparator by using an activation threshold. For example, the ADC compares the overall count with the activation threshold or criterion. If the overall count is greater than the activation threshold, the output of the bit line is "1". Otherwise, if the overall count is less than or equal to the activation threshold, the output of the bit line is "0". However, it is desirable to eliminate the XNOR function and its corresponding transistor(s) in each bit cell to reduce the memory size. The specific techniques for adjusting the activation threshold are as follows:
[0077] The convolutional layer of a neural network (e.g., a binary neural network) may include units (e.g., bit cells) organized into an array (e.g., an in-memory computing array). These units include gating devices, and the charge levels presented in these gating devices represent the stored weights of the array. A trained XNOR binary neural network with an array having M rows (e.g., the size of the input vector) and N columns includes MxN binary synaptic weights W ij (which are weights in the form of binary values [-1, +1]) and N activation thresholds C j . The inputs of the array may correspond to word lines, and the outputs may correspond to bit lines. For example, the input activations X i (which are inputs in the form of binary values [-1, +1]) are X 1 , X 2 ...X M . The sum of the products of the inputs and the corresponding weights is referred to as the weighted sum
[0078] For example, when the weighted sum is greater than the activation threshold C j , the output is equal to one (1). Otherwise, the output is equal to zero (0). The XNOR binary neural network can be mapped to a non-XNOR binary neural network in the [0, 1] space while eliminating the XNOR function and its corresponding transistor(s) in each bit cell to reduce the memory size. In one aspect, the XNOR binary neural network can be mapped to a non-XNOR binary neural network by the following adjustment of the activation threshold C j for each column:
[0079] Equation 1 illustrates the relationship between the sum of the products of the inputs and the corresponding weights and the activation threshold C j for the XNOR binary neural network:
[0080]
[0081] wherein:
[0082] -X i (e.g., X 1 , X 2 ...X M ) is an input in the form of a binary value [-1, +1]
[0083] -W ij is a weight in the form of a binary value [-1, +1] (e.g., an MxN matrix);
[0084] -Y j (e.g., Y 1 , Y 2 ...Y N ) is an output of the bit line in the form of a binary value [-1, +1]
[0085] Conventionally, in-memory computing binary neural network implementations are achieved by performing XNOR operations at each bit cell and summing the results for each bit line. However, adding the XNOR function in each bit cell increases the layout area and power consumption. For example, the number of transistors in each cell in a memory (e.g., SRAM) increases from 6 or 8 to 12. Accordingly, aspects of the present disclosure relate to using activation threshold adjustment to transform the multiply-accumulate operations of an in-memory computing array of an artificial neural network (e.g., a binary neural network) from the [-1, +1] space to the [0, 1] space.
[0086] Transforming the binary neural network from the [-1, +1] space (in the case) to the [0, 1] space, and comparing the output of each bit line in the [0, 1] space with different thresholds as follows (e.g., the derivation thereof is discussed below):
[0087]
[0088] - is an input in the form of a binary value [0, 1]
[0089] - is a weight in the form of a binary value [0, 1]
[0090] where corresponds to an adjustable activation (or adjusted activation threshold) in the [0, 1] space, and represents a conversion bias in the [0, 1] space (e.g., a conversion bias current reference).
[0091] Adjusting the activation threshold as described herein allows for avoiding / omitting the implementation of XNOR functionality in each bit cell while achieving the same result as in the case where XNOR functionality is implemented. Adjusting the activation threshold enables the use of simpler and smaller memory bit cells (e.g., eight-transistor (8T) SRAM).
[0092] The following equations (Equations 3 and 4) include variables for mapping a network in the [-1, +1] space to the [0, 1] space:
[0093]
[0094]
[0095] Inserting the values of Equation three (3) and Equation four (4) into Equation 1, the adjusted activation threshold can be determined through a conversion or transformation between the [-1, +1] space and the [0, 1] space. The following is the derivation of the adjusted activation threshold:
[0096]
[0097] Expand Equation 5:
[0098]
[0099] The output Y of each bit line in the form of binary values [-1, +1] j is compared with the activation threshold C in the [-1, +1] space j as shown in Equation 1
[0100] Y j > C j
[0101] Insert the output value Y of the bit line in Equation 6 j into Equation 1 to obtain the overall count of each bit line in the [0, 1] space
[0102]
[0103]
[0104]
[0105] The overall count of each bit line in the [0, 1] space in Equation 9 is mapped to the overall count of each bit line in the [-1, +1] space Then the N activation thresholds C in the [-1, +1] space j are mapped to the adjustable activation and the conversion bias in the [0, 1] space For example, to achieve a non-XNOR binary neural network in the [0,1] space while eliminating the XNOR function and its corresponding transistor(s), an adjustable activation is used and a conversion bias in the [0, 1] space
[0106] Referring to Equation 9, the function is used to determine the adjustable activation For example, the element corresponds to the activation value of the adjustable activation The adjustable activation does not depend on the activation input X i The adjustable activation also includes a predetermined parameter (e.g., the binary synaptic weight W j ) that can be changed to adjust the activation threshold C ij .
[0107] The function is used to determine the conversion bias The conversion bias depends only on the input activation. This means that the conversion bias changes continuously as new inputs are received.
[0108] For example, when C j = 0; and , then the values of and are calculated as follows with respect to Equation 9:
[0109]
[0110] and
[0111]
[0112] Similarly, when and , then the values of and are calculated as follows with respect to Equation 9:
[0113]
[0114] and
[0115]
[0116] In some aspects, the function can be used to convert a binary value The weight of the form is set equal to 1 to generate a conversion bias from a reference column (e.g., a bit line that is not part of these N bit lines and is used as a reference). The one equal to 1 value is inserted into Equation 9, and the resulting equation is as follows:
[0117]
[0118] The overall count is equal to which can be used to offset the activation current. Thus, only a single column is specified to determine the reference bias value or conversion bias for the entire memory array.
[0119] Figure 5 Exemplary architecture 500 for performing XNOR equivalent operations by adjusting column thresholds of an in-memory computing array of an artificial neural network in accordance with aspects of the present disclosure is illustrated. The architecture includes a memory array that includes a reference column 502, comparators 504, bit lines 506 (corresponding to the columns of the memory array), and word lines 508 (corresponding to the rows of the memory array). Input activations are received via the word lines. In one example, the binary value of the input activation X i is 10101011.
[0120] Multiplication (e.g., X i W ij ) occurs at each bit cell, and the multiplication results from each bit cell are output via the bit lines 506. Summing devices 510 associated with each bit line 506 sum the outputs of each bit cell of the memory array and pass the result to the comparators 504. The comparator 504 can be part of an analog-to-digital converter (ADC) as shown in Figure 4 . In one aspect, the outputs of each bit cell associated with the first bit line are summed separately from the outputs of each bit cell associated with the second bit line.
[0121] Activation threshold adjustment is performed at each column (corresponding to these bit lines) rather than at each bit cell to improve area efficiency. For example, the sum (or overall count) of the outputs of each bit cell associated with the first bit line is provided as a first input to the comparator 504. The second input to the comparator includes the adjustable activation plus the conversion bias (e.g., a conversion bias current reference). For example, the conversion bias can be programmed for each bit line 506 according to a criterion (e.g., the reference column 502). When the overall count is greater than the adjustable activation plus the conversion bias When the sum is reached, the output of comparator 504 (which corresponds to the output of the first bit line) is "1". Otherwise, if the overall count is less than or equal to the adjustable activation and the conversion bias sum, the output of this comparator (which corresponds to the output of the first bit line) is "0". Thus, each bit line overall count is compared with the adjustable activation and the conversion bias sum.
[0122] Figure 6 Method 600 for performing XNOR equivalent operations by adjusting column thresholds of an in-memory computing array of an artificial neural network according to aspects of the present disclosure is illustrated. As Figure 6 shown, at block 602, the activation thresholds generated for each column of the in-memory computing array can be adjusted based on a function of weight values and activation values. At block 604, a conversion bias current reference is calculated based on input values from an input vector to the in-memory computing array, which is programmed with a set of weight values. Each of the adjusted activation threshold and the conversion bias current reference is used as a threshold for determining the output value of the in-memory computing array. The in-memory computing array has both columns and rows.
[0123] According to a further aspect of the present disclosure, an apparatus for performing XNOR equivalent operations by adjusting column thresholds of an in-memory computing array of an artificial neural network is described. The apparatus includes: means for adjusting the activation thresholds of each column of the array based on a function of weight values and activation values. The adjusting means: includes deep convolutional network 200, deep convolutional network 350, convolutional layer 232, SoC 100, CPU 102, architecture 500, architecture 400, and / or convolutional block 354A. The apparatus further includes: means for calculating a conversion bias current reference based on input values from an input vector. The calculating means includes: deep convolutional network 200, deep convolutional network 350, convolutional layer 232, SoC 100, CPU 102, architecture 500, architecture 400, and / or convolutional block 354A.
[0124] The apparatus further includes: means for comparing the bit line overall count with the sum of the conversion bias current reference and the adjusted activation threshold to determine the output of the bit line. The comparing means includes: Figure 5 comparator 504, and / or Figure 4 an analog-to-digital converter (ADC). In another aspect, the foregoing means can be any module or any apparatus configured to perform the functions recited by the foregoing means.
[0125] The various operations of the methods described above can be performed by any suitable means capable of performing the corresponding functions. These means may include various hardware and / or software components and / or modules, including but not limited to circuitry, application specific integrated circuits (ASICs), or processors. In general, where operations are illustrated in the figures, those operations may have corresponding paired means plus function components with similar numbers.
[0126] As used herein, the term "determine" encompasses a variety of actions. For example, "determine" may include computing, calculating, processing, deriving, researching, looking up (e.g., looking up in a table, database, or other data structure), ascertaining, and the like. Additionally, "determine" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), and similar actions. Further, "determine" may include parsing, selecting, choosing, establishing, and similar actions.
[0127] As used herein, the phrase reciting "at least one of" a list of items refers to any combination of those items, including a single member. By way of example, "at least one of a, b, or c" is intended to cover: a, b, c, a - b, a - c, b - c, and a - b - c.
[0128] The various illustrative logical blocks, modules, and circuits described in connection with the present disclosure can be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array signal (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general purpose processor may be a microprocessor, but in the alternative, the processor may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0129] The steps of the methods or algorithms described in connection with the present disclosure may be implemented directly in hardware, in software modules executed by a processor, or in a combination of both. The software modules may reside in any form of storage medium known in the art. Some examples of storage media that may be used include random access memory (RAM), read only memory (ROM), flash memory, erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), registers, hard disks, removable disks, CD-ROMs, and the like. The software modules may include a single instruction or many instructions and may be distributed over several different code segments, distributed among different programs, and across multiple storage media. The storage medium may be coupled to the processor such that the processor can read information from, and write information to, the storage medium. In an alternative, the storage medium may be integrated into the processor.
[0130] The methods disclosed herein include one or more steps or acts for implementing the described methods. These method steps and / or acts may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of the steps or acts is specified, the order and / or use of the specific steps and / or acts may be altered without departing from the scope of the claims.
[0131] The described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in hardware, an example hardware configuration may include a processing system in a device. The processing system may be implemented with a bus architecture. Depending on the particular application and overall design constraints of the processing system, the bus may include any number of interconnected buses and bridges. The bus may link together various circuits including a processor, a machine-readable medium, and a bus interface. The bus interface may be used to connect, among other things, a network adapter to the processing system via the bus. The network adapter may be used to implement signal processing functions. For some aspects, a user interface (e.g., keypad, display, mouse, joystick, etc.) may also be connected to the bus. The bus may also link various other circuits such as a timing source, peripherals, voltage regulators, power management circuits, and similar circuits that are well known in the art and will not be described further herein.
[0132] The processor may be responsible for managing the bus and general processing, including the execution of software stored on a machine-readable medium. The processor may be implemented with one or more general-purpose and / or special-purpose processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuitry that can execute software. Software should be construed broadly to mean instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. As an example, the machine-readable medium may include random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, magnetic disks, optical disks, hard drives, or any other suitable storage medium, or any combination thereof. The machine-readable medium may be embodied in a computer program product. The computer program product may include packaging material.
[0133] In a hardware implementation, the machine-readable medium may be a part of the processing system separate from the processor. However, as will be readily appreciated by those skilled in the art, the machine-readable medium or any part thereof may be external to the processing system. As an example, the machine-readable medium may include transmission lines, carrier waves modulated with data, and / or computer products separate from the device, all of which may be accessed by the processor via a bus interface. Alternatively or additionally, the machine-readable medium or any part thereof may be integrated into the processor, such as may be the case with a cache and / or a general register file. Although the various components discussed may be described as having a particular location, such as local components, they may also be configured in various ways, such as some components being configured as part of a distributed computing system.
[0134] The processing system can be configured as a general-purpose processing system that has one or more microprocessors providing processor functionality, as well as an external memory providing at least a portion of the machine-readable medium, all linked together by an external bus architecture to other support circuitry. Alternatively, the processing system can include one or more neuromorphic processors for implementing the neuron models and nervous system models described herein. As another alternative, the processing system can be implemented with an application specific integrated circuit (ASIC) with a processor, bus interface, user interface, support circuitry, and at least a portion of the machine-readable medium integrated in a single chip, or with one or more field programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gated logic, discrete hardware components, or any other suitable circuitry, or any combination of circuits capable of performing the various functions described throughout this disclosure. Depending on the particular application and the overall design constraints imposed on the system, those of ordinary skill in the art will recognize how best to implement the functionality described with respect to the processing system.
[0135] The machine-readable medium can include several software modules. These software modules include instructions that, when executed by the processor, cause the processing system to perform various functions. These software modules can include a transmission module and a reception module. Each software module can reside in a single storage device or be distributed across multiple storage devices. As an example, when a triggering event occurs, the software module can be loaded from a hard drive into RAM. During the execution of the software module, the processor can load some of the instructions into a cache to improve access speed. One or more cache lines can then be loaded into the general register file for execution by the processor. When referring to the functionality of the software modules below, it will be understood that such functionality is implemented by the processor when the processor executes instructions from the software module. Additionally, it should be appreciated that aspects of the present disclosure result in improvements to the capabilities of a processor, computer, machine, or other system implementing such aspects.
[0136] If implemented in software, each function can be stored on or transmitted via a computer-readable medium as one or more instructions or code. The computer-readable medium includes both computer storage media and communication media, including any medium that facilitates the transfer of a computer program from one place to another. The storage media can be any available medium that can be accessed by a computer. By way of example and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a web site, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology such as infrared (IR), radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology such as infrared, radio, and microwave is included in the definition of the medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and disc, where disk often magnetically reproduces data, while disc optically reproduces data with a laser. Thus, in some aspects, the computer-readable medium may include non-transitory computer-readable media (e.g., tangible media). Additionally, for other aspects, the computer-readable medium may include transitory computer-readable media (e.g., signals). Combinations of the above should also be included within the scope of computer-readable media.
[0137] Accordingly, some aspects may include a computer program product for performing the operations given herein. For example, such a computer program product may include a computer-readable medium having (and / or encoded with) instructions that can be executed by one or more processors to perform the operations described herein. For some aspects, the computer program product may include packaging material.
[0138] Furthermore, it should be appreciated that modules and / or other suitable means for performing the methods and techniques described herein can be downloaded and / or otherwise obtained by a user terminal and / or a base station where applicable. For example, such devices can be coupled to a server to facilitate the transfer of means for performing the methods described herein. Alternatively, the various methods described herein can be provided via a storage device (e.g., RAM, ROM, a physical storage medium such as a compact disc (CD) or floppy disk, etc.) such that once the storage device is coupled to or provided to the user terminal and / or the base station, the device can obtain the various methods. Additionally, any other suitable technology can be utilized that is adapted to provide the methods and techniques described herein to a device.
[0139] It will be understood that the claims are not limited to the exact configurations and components illustrated above. Various modifications, substitutions, and variations can be made in the layout, operation, and details of the methods and apparatuses described above without departing from the scope of the claims.
Claims
1. A device, comprising: a memory - in - computing array including rows and columns, the memory - in - computing array being configured to: adjust an activation threshold generated for each column of the memory - in - computing array based on a function of weight values and activation values; and calculate a conversion bias current reference based on input values in an input vector to the memory - in - computing array, the memory - in - computing array being programmed with a set of weight values, wherein the adjusted activation threshold and the conversion bias current reference are used as thresholds for determining output values of the memory - in - computing array.
2. The device according to claim 1, further comprising: a comparator configured to compare a bit - line overall count with the sum of the conversion bias current reference and the adjusted activation threshold to determine the output of the bit line.
3. The device according to claim 1, wherein the artificial neural network including the memory - in - computing array comprises a binary neural network.
4. The device according to claim 1, wherein the activation threshold is less than half of the number of rows of the memory - in - computing array, and the number of rows corresponds to the size of the input vector.
5. The device according to claim 1, wherein the conversion bias current reference is less than half of the number of rows of the memory - in - computing array, and the number of rows corresponds to the size of the input vector.
6. A method, comprising: adjusting an activation threshold generated for each column of a memory - in - computing array having rows and columns based on a function of weight values and activation values; calculating a conversion bias current reference based on input values in an input vector to the memory - in - computing array, the memory - in - computing array being programmed with a set of weight values, wherein the adjusted activation threshold and the conversion bias current reference are used as thresholds for determining output values of the memory - in - computing array.
7. The method according to claim 6, further comprising: comparing a bit - line overall count with the sum of the conversion bias current reference and the adjusted activation threshold to determine the output of the bit line.
8. The method according to claim 6, wherein the artificial neural network including the memory - in - computing array comprises a binary neural network.
9. The method according to claim 6, wherein the activation threshold is less than half of the number of rows of the memory - in - computing array, and the number of rows corresponds to the size of the input vector.
10. The method according to claim 6, wherein the conversion bias current reference is less than half of the number of rows of the memory - in - computing array, and the number of rows corresponds to the size of the input vector.
11. A non - transient computer - readable medium having non - transient program code recorded thereon, the program code comprising: program code for adjusting an activation threshold generated for each column of a memory - in - computing array having rows and columns based on a function of weight values and activation values; and Program code for calculating a conversion bias current reference based on input values in an input vector to a computing array within a memory, the computing array within the memory being programmed with a set of weight values, wherein an adjusted activation threshold and the conversion bias current reference are used as thresholds for determining an output value of the computing array within the memory.
12. The non-transitory computer-readable medium of claim 11, further comprising: Program code for comparing a total bit line count with the sum of the conversion bias current reference and the adjusted activation threshold to determine an output of the bit line.
13. The non-transitory computer-readable medium of claim 11, wherein the artificial neural network subject to the adjustment and the calculation comprises a binary neural network.
14. The non-transitory computer-readable medium of claim 11, wherein the activation threshold is less than half of the number of rows of the computing array within the memory, the number of rows corresponding to the size of the input vector.
15. The non-transitory computer-readable medium of claim 11, wherein the conversion bias current reference is less than half of the number of rows of the computing array within the memory, the number of rows corresponding to the size of the input vector.
16. An apparatus, comprising: means for adjusting an activation threshold generated for each column of a computing array within a memory having rows and columns based on a function of weight values and activation values; and means for calculating a conversion bias current reference based on input values in an input vector to the computing array within the memory, the computing array within the memory being programmed with a set of weight values, wherein an adjusted activation threshold and the conversion bias current reference are used as thresholds for determining an output value of the computing array within the memory.
17. The apparatus of claim 16, further comprising: means for comparing a total bit line count with the sum of the conversion bias current reference and the adjusted activation threshold to determine an output of the bit line.
18. The apparatus of claim 16, wherein the artificial neural network comprising the computing array within the memory comprises a binary neural network.
19. The apparatus of claim 16, wherein the activation threshold is less than half of the number of rows of the computing array within the memory, the number of rows corresponding to the size of the input vector.
20. The apparatus of claim 16, wherein the conversion bias current reference is less than half of the number of rows of the computing array within the memory, the number of rows corresponding to the size of the input vector.
Citation Information
Patent Citations
Programmable Neuron For Analog Non-Volatile Memory In Deep Learning Artificial Neural Network
US20190205729A1