Folded column adder architecture for digital compute-in-memory

JP2024530610A5Active Publication Date: 2025-06-27QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024505074
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-08-02
Filing Date
2022-07-18
Publication Date
2025-06-27
Estimated Expiration
2042-07-18

AI Technical Summary

Technical Problem

Existing machine learning processing systems face inefficiencies in power and space constraints, particularly in edge devices, and traditional Computation in Memory (CIM) processes using analog signals lead to inaccuracies in neural network calculations.

Method used

A digital CIM architecture with a folded column adder circuit configuration, utilizing memory cells to store neural network weights and employing adder trees and accumulators for accurate in-memory computation, allowing for configurable bit sizes and reduced power consumption.

Benefits of technology

The solution enhances processing efficiency and accuracy in machine learning tasks by reducing power usage and latency, particularly suitable for edge devices and mobile applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Some aspects provide an apparatus for performing machine learning tasks, particularly for computation in memory architectures. One aspect provides a circuit for in-memory computation. The circuit generally includes a plurality of memory cells on each of a plurality of columns of a memory, the plurality of memory cells configured to store a plurality of bits representing weights of a neural network, where the plurality of memory cells on each of the plurality of columns are on different word lines of the memory, a plurality of summing circuits, each coupled to a respective one of the plurality of columns, a first adder circuit coupled to outputs of at least two of the plurality of summing circuits, and an accumulator coupled to an output of the first adder circuit.
Need to check novelty before this filing date? Find Prior Art

Description

Claiming priority

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Application No. 17 / 391,718, filed August 2, 2021, which is assigned to the assignee of this application and is incorporated by reference in its entirety herein. [Technical field]

[0002] Introduction Aspects of the present disclosure relate to performing machine learning tasks, and in particular to computation-in-memory architectures. [Background technology]

[0003]

[0003] Machine learning is generally the process of creating a trained model (e.g., an artificial neural network, tree, or other structure) that represents a generalized fit to a set of training data that is known a priori. Applying the trained model to new data produces inferences, which can be used to gain insight into the new data. In some cases, applying a model to new data is described as "performing inference" on the new data.

[0004]

[0004] As the use of machine learning has proliferated to enable various machine learning (or artificial intelligence) tasks, a need has arisen for more efficient processing of machine learning model data. In some cases, dedicated hardware such as machine learning accelerators may be used to improve the capacity of a processing system to process machine learning model data. However, such hardware requires space and power, which is not always available on the processing device. For example, "edge processing" devices, such as mobile devices, always-on devices, and Internet of Things (IoT) devices, must generally balance processing power with power and packaging constraints. Furthermore, accelerators may move data across a common data bus, which can cause significant power usage and introduce latency to other processes sharing the data bus. Therefore, other aspects of the processing system have been considered to process machine learning model data.

[0005]

[0005] A memory device is an example of another aspect of a processing system that can be utilized to perform processing of machine learning model data through a so-called computation-in-memory (CIM) process. Conventional CIM processes perform calculations using analog signals, which can cause inaccuracies in the calculation results and adversely affect neural network calculations. Therefore, a system and method for performing computation-in-memory with increased accuracy is needed. Summary of the Invention

[0006]

[0006] Certain aspects provide apparatus and techniques for performing machine learning tasks, particularly for computation-in-memory architectures.

[0007] One aspect provides a circuit for in-memory computation that generally includes a plurality of memory cells on each of a plurality of columns of a memory, the plurality of memory cells configured to store a plurality of bits representing weights of a neural network, where the plurality of memory cells on each of the plurality of columns are on different word lines of the memory; a plurality of adder circuits, each coupled to a respective one of the plurality of columns; a first adder circuit coupled to outputs of at least two of the plurality of adder circuits; and an accumulator coupled to an output of the first adder circuit.

[0008] One aspect provides a method for in-memory computation that generally includes: summing output signals on a respective one of a plurality of columns of a memory via each of a plurality of summing circuits, where a plurality of memory cells are on each of the plurality of columns, the plurality of memory cells storing a plurality of bits representing weights of a neural network, where the plurality of memory cells on each of the plurality of columns are on different word lines of the memory; summing output signals of at least two of the plurality of summing circuits via a first adder circuit; and accumulating output signals of the first adder circuit via an accumulator.

[0009] One aspect provides an apparatus for in-memory computation that generally includes first means for adding output signals on a respective one of a plurality of columns of a memory, where a plurality of memory cells are on each of the plurality of columns, the plurality of memory cells storing a plurality of bits representing weights of a neural network, where the plurality of memory cells on each of the plurality of columns are on different word lines of the memory; second means for adding at least two output signals of the first means for adding; and means for accumulating the output signals of the second means for adding.

[0010]

[0010] Other aspects provide a processing system configured to perform the above-mentioned methods as well as methods described herein; a non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of the processing system, cause the processing system to perform the above-mentioned methods as well as methods described herein; a computer program product embodied on a computer-readable storage medium comprising code for performing the above-mentioned methods as well as methods further described herein; and a processing system comprising means for performing the above-mentioned methods as well as methods further described herein.

[0011] The following description and the annexed drawings set forth in detail certain illustrative features of the one or more aspects.

[0012]

[0001] So that the above-recited features of the present disclosure may be understood in detail, a more particular description briefly summarized above may be had by reference to the embodiments, some of which are illustrated in the accompanying drawings. However, it should be noted that the accompanying drawings illustrate only some typical embodiments of the present disclosure, and therefore should not be considered as limiting the scope of the present disclosure, since the description may lead to other equally effective embodiments. [Brief description of the drawings]

[0013] [Figure 1A]

[0012] FIG. 1 illustrates examples of various types of neural networks that may be implemented according to aspects of the present disclosure. [Figure 1B] FIG. 1 illustrates examples of various types of neural networks that may be implemented according to aspects of the present disclosure. [Figure 1C] FIG. 1 illustrates examples of various types of neural networks that may be implemented according to aspects of the present disclosure. [Figure 1D] FIG. 1 illustrates examples of various types of neural networks that may be implemented according to aspects of the present disclosure. [Diagram 2]

[0013] FIG. 1 illustrates an example of a traditional convolution operation that may be implemented according to aspects of the present disclosure. [Figure 3A]

[0014] FIG. 1 illustrates an example of a depthwise separable convolution operation that may be implemented by aspects of the present disclosure. [Figure 3B] FIG. 1 illustrates an example of a depthwise separable convolution operation that may be implemented by aspects of the present disclosure. [Figure 4]

[0015] FIG. 2 illustrates an exemplary memory cell implemented as an eight-transistor (8T) static random access memory (SRAM) cell for compute-in-memory (CIM) circuitry. [Figure 5A]

[0016] FIG. 2 illustrates a circuit for a CIM in accordance with some aspects of the present disclosure. [Figure 5B]

[0017] FIG. 2 illustrates an example implementation of an adder circuit. [Figure 5C]

[0018] FIG. 2 illustrates an example implementation of an accumulator. [Figure 6]

[0019] FIG. 2 illustrates a circuit for a CIM implemented using a bit string adder tree in accordance with certain aspects of the disclosure. [Figure 7]

[0020] 7 is a timing diagram illustrating signals associated with the circuit of FIG. 6 in accordance with some aspects of the present disclosure. [Figure 8A]

[0021] FIG. 13 is a block diagram illustrating a CIM circuit with configurable bit size of weights in accordance with some aspects of the present disclosure. [Figure 8B] FIG. 13 is a block diagram illustrating a CIM circuit with configurable bit size of weights in accordance with some aspects of the present disclosure. [Figure 8C] FIG. 13 is a block diagram illustrating a CIM circuit with configurable bit size of weights in accordance with some aspects of the present disclosure. [Figure 9]

[0022] 1 is a flow diagram illustrating example operations for in-memory computing in accordance with certain aspects of the present disclosure. [Figure 10]

[0023] FIG. 1 illustrates an example electronic device configured to perform operations for signal processing in a neural network in accordance with some aspects of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0014]

[0024] For ease of understanding, wherever possible, identical reference numbers have been used to designate identical elements that are common to the figures. It is contemplated that elements and features of one embodiment may be beneficially incorporated in other embodiments without further recitation.

[0015]

[0025] Aspects of the present disclosure provide apparatuses, methods, processing systems, and computer-readable media for implementing computation-in-memory (CIM) to handle data-intensive processing, such as implementing machine learning models. Some aspects provide techniques for implementing digital CIM using adder circuits, each adder circuit adding (e.g., accumulating) output signals on a respective one of multiple columns of a memory after multiple computation cycles. As used herein, a "adder circuit" generally refers to any circuit that adds (or accumulates over successive computation cycles) output signals of memory cells on a column. In some cases, the adder circuit may be an accumulator. An accumulator generally refers to a circuit used to accumulate output signals over multiple cycles. In other cases, the adder circuit may be an adder tree. An "adder circuit" or "adder tree" generally refers to a digital adder used to add output signals of multiple memory cells (e.g., memory cells across word lines or columns). One exemplary implementation of an adder circuit is described herein with respect to FIG. 5B, and one exemplary implementation of an accumulator is described herein with respect to FIG. 5C. The adder circuit may be implemented as an adder tree having multiple adder circuits, or as an accumulator. In some aspects, the word lines of the CIM circuit are activated in succession, and the accumulators perform accumulation simultaneously to provide an accumulation result after two or more of the word lines are activated in succession.

[0016]

[0026] Some aspects provide a folding architecture that allows configurability of the bit size of the weights used for the calculation. For example, one or more processing paths (also called "wings") of the CIM architecture can be disabled to adjust the bit size of the weights being used. For example, eight processing paths (including, for example, columns and associated processing circuits) can be used to implement 8-bit weights, or four processing paths can be used to implement 4-bit weights (with the other four processing paths temporarily disabled).

[0017]

[0027] CIM-based machine learning (ML) / artificial intelligence (AI) may be used for a wide variety of tasks, including image and audio processing and making wireless communication decisions (e.g., to optimize or at least increase throughput and signal quality). Furthermore, CIMs may be based on various types of memory architectures, such as dynamic random access memory (DRAM), static random access memory (SRAM) (e.g., based on the SRAM cell described in FIG. 4), magnetoresistive random access memory (MRAM), resistive random access memory (ReRAM or RRAM®), and may be attached to various types of processing units, including central processing units (CPUs), digital signal processors (DSPs), graphics processing units (GPUs), field programmable gate arrays (FPGAs), AI accelerators, and the like. In general, CIMs may beneficially reduce the "memory wall" problem, which is when moving data into and out of memory consumes more power than computing the data. Thus, significant power savings may be realized by implementing computation-in-memory. This is particularly useful for various types of electronic devices, such as low power edge processing devices, mobile devices, etc.

[0018]

[0028] For example, a mobile device may include a memory device configured to store data and perform compute-in-memory operations. The mobile device may be configured to perform ML / AI operations based on data generated by the mobile device, such as image data generated by a camera sensor of the mobile device. Thus, a memory controller unit (MCU) of the mobile device may load weights from another on-board memory (e.g., flash or RAM) into a CIM array of the memory device and allocate an input feature buffer and an output (e.g., output activation) buffer. The processing device may then begin processing the image data, for example, by loading a layer in the input buffer and processing that layer with the weights loaded into the CIM array. This process may be repeated for each layer of the image data, and the outputs (e.g., output activations) may be stored in an output buffer and then used by the mobile device for ML / AI tasks such as face recognition. A brief background on neural networks, deep neural networks, and deep learning

[0029] Neural networks are organized into layers of interconnected nodes. In general, a node (or neuron) is where computation occurs. For example, a node may combine input data with a set of weights (or coefficients) that either amplify or attenuate the input data. Thus, the amplification or attenuation of an input signal can be viewed as an assignment of relative importance to various inputs with respect to the task the network is trying to learn. In general, input-weight products are added (or accumulated) and then the sum is passed through the node's activation function to determine whether and how far the signal should proceed further through the network.

[0019]

[0030] In its most basic implementation, a neural network may have an input layer, a hidden layer, and an output layer. A "deep" neural network generally has two or more hidden layers.

[0020]

[0031] Deep learning is a method of training deep neural networks. Broadly speaking, deep learning maps inputs to the network to outputs from the network, and is therefore sometimes called a "universal approximator" because it can learn to approximate an unknown function f(x)=y between any input x and any output y. In other words, deep learning finds the correct f to transform x to y.

[0021]

[0032] More specifically, deep learning trains each layer of nodes based on a distinct set of features, which are the output from the previous layer. Thus, in each successive layer of a deep neural network, the features become more complex. Deep learning is therefore powerful because it learns to represent the input at successively higher levels of abstraction in each layer, thereby accumulating useful feature representations of the input data, allowing it to progressively extract higher level features from the input data and perform complex tasks such as object recognition.

[0022]

[0033] For example, when presented with visual data, a first layer of a deep neural network may learn to recognize relatively simple features in the input data, such as edges. In another example, when presented with auditory data, a first layer of a deep neural network may learn to recognize spectral power at specific frequencies in the input data. A second layer of the deep neural network may then learn to recognize combinations of features, such as simple shapes in the case of visual data, or combinations of sounds in the case of auditory data, based on the output of the first layer. A higher layer may then learn to recognize complex shapes in the visual data, or words in the auditory data. An even higher layer may learn to recognize common visual objects or spoken phrases. Thus, deep learning architectures may work particularly well when applied to problems with natural hierarchical structures. Layer Connectivity in Neural Networks

[0034] Neural networks, such as deep neural networks (DNNs), can be designed with a variety of connectivity patterns between layers.

[0023]

[0035] 1A illustrates an example of a fully-connected neural network 102, in which each node in the first layer communicates its output to every node in the second layer, such that each node in the second layer receives input from every node in the first layer.

[0024]

[0036] 1B shows an example of a locally connected neural network 104. In the locally connected neural network 104, a node in a first layer may be connected to a limited number of nodes in a second layer. More generally, the locally connected layer of the locally connected neural network 104 may be configured such that each node in the layer has the same or similar connectivity pattern, but with connection strengths (or weights) that may have different values ​​(e.g., values ​​associated with local areas 110, 112, 114, and 116 of the first layer nodes). The connectivity pattern of the local connections may result in spatially distinct receptive fields in the upper layers, since the upper layer nodes in a given region may receive inputs that are conditioned through training to the properties of a limited portion of the total inputs to the network.

[0025]

[0037] One type of locally connected neural network is a convolutional neural network (CNN). Figure 1C shows an example of a convolutional neural network 106. The convolutional neural network 106 can be configured such that the connection strengths associated with the inputs for each node in the second layer are shared (e.g., for local areas 108 that overlap with another local area of ​​the first layer node). Convolutional neural networks are suitable for problems where the spatial location of the inputs is meaningful.

[0026]

[0038] One type of convolutional neural network is the deep convolutional network (DCN), which is a network of multiple convolutional layers, which may be further composed of, for example, pooling layers and normalization layers.

[0027]

[0039] 1D shows an example of a DCN 100 designed to recognize visual features in an image 126 generated by an image capture device 130. For example, if the image capture device 130 is a camera mounted in or on a vehicle (or otherwise moving with the vehicle), the DCN 100 can be trained using a variety of supervised learning techniques to identify traffic signs, and even numbers on traffic signs. The DCN 100 can be trained for other tasks as well, such as identifying lane markings, or identifying traffic signals. These are just a few example tasks, and many other tasks are possible.

[0028]

[0040] In the example of FIG. 1D, the DCN 100 includes a feature extraction section and a classification section. Upon receiving an image 126, a convolutional layer 132 applies a convolutional kernel (e.g., as shown and described in FIG. 2) to the image 126 to generate a first set of feature maps (or intermediate activations) 118. Generally, a "kernel" or "filter" comprises a multi-dimensional array of weights designed to emphasize different aspects of the input data channels. In various examples, "kernel" and "filter" may be used interchangeably to refer to a set of weights applied in a convolutional neural network.

[0029]

[0041] The first set of feature maps 118 may then be subsampled by a pooling layer (e.g., a max pooling layer, not shown) to generate a second set of feature maps 120. The pooling layer may reduce the size of the first set of feature maps 118 while retaining most of the information to improve model performance. For example, the second set of feature maps 120 may be downsampled by the pooling layer from a 28×28 matrix to a 14×14 matrix.

[0030]

[0042] This process may be repeated through many layers. In other words, the second set of feature maps 120 may be further convolved through one or more subsequent convolutional layers (not shown) to generate one or more subsequent sets of feature maps (not shown).

[0031]

[0043] 1D example, the second set of feature maps 120 is provided to a fully connected layer 124, which generates an output feature vector 128. Each feature in the output feature vector 128 may include a number corresponding to a possible feature of the image 126, such as "sign," "60," and "100." In some cases, a softmax function (not shown) may convert the numbers in the output feature vector 128 into probabilities. In such a case, the output 122 of the DCN 100 is the probability that the image 126 contains one or more features.

[0032]

[0044] A softmax function (not shown) may convert the individual elements of the output feature vector 128 into probabilities such that the output 122 of the DCN 100 is one or more probabilities that the image 126 includes one or more features, such as a sign with the number "60" thereon, as in the case of the image 126. Thus, in this example, the probabilities in the output 122 for "sign" and "60" should be higher than the probabilities of other elements of the output 122, such as "30", "40", "50", "70", "80", "90", and "100".

[0033]

[0045] Prior to training the DCN 100, the output 122 produced by the DCN 100 may be inaccurate. Thus, an error may be calculated between the output 122 and a target output known a priori. For example, here the target output is an indication that the image 126 contains a "sign" and the number "60." Using the known target output, the weights of the DCN 100 may then be adjusted through training such that subsequent outputs 122 of the DCN 100 achieve (with high probability) the target output.

[0034]

[0046] To adjust the weights of the DCN 100, the learning algorithm may calculate a gradient vector for the weights. The gradient vector may indicate the amount that the error would increase or decrease if the weights were adjusted in a particular way. The weights may then be adjusted to reduce the error. This manner of adjusting the weights is sometimes called "backpropagation" because the adjustment process involves a "backward pass" through the layers of the DCN 100.

[0035]

[0047] In practice, the error gradient of the weights may be calculated over a small number of examples such that the calculated gradient approximates the true error gradient. This approximation method is sometimes called stochastic gradient descent. Stochastic gradient descent may be repeated until the achievable error rate of the overall system no longer decreases or until the error rate reaches a target level.

[0036]

[0048] After training, the DCN 100 can be presented with new images and the DCN 100 can generate inferences, such as classifications or probabilities that various features are present in the new image. Convolutional Techniques for Convolutional Neural Networks

[0049] Convolution is generally used to extract useful features from an input data set. For example, in a convolutional neural network, such as the one described above, convolution allows the extraction of different features using kernels and / or filters whose weights are automatically learned during training. The extracted features are then combined to make inferences.

[0037]

[0050] Activation functions may be applied before and / or after each layer of a convolutional neural network. Activation functions are generally mathematical functions that determine the output of a node in a neural network. Thus, activation functions determine whether a node should pass information or not based on whether the node's input is relevant to the model's prediction. In an example where y=conv(x) (i.e., y is a convolution of x), both x and y may generally be considered "activations." However, with respect to a particular convolution operation, x may also be referred to as a "pre-activation" or "input activation" since x exists before the particular convolution, and y may be referred to as an output activation or feature map.

[0038]

[0051] 2 shows an example of classical convolution where an input image of 12 pixels by 12 pixels by 3 channels is convolved using a 5×5×3 convolution kernel 204 and a stride (or step size) of 1. The resulting feature map 206 is 8 pixels by 8 pixels by 1 channel. As can be seen in this example, classical convolution may change the dimensionality of the input data compared to the output data (here, from 12×12 pixels to 8×8 pixels), including the channel dimensionality (here, from 3 channels to 1 channel).

[0039]

[0052] One way to reduce the computational burden (e.g., measured in floating-point operations per second (FLOPs)) and number of parameters associated with neural networks with convolutional layers is to factorize the convolutional layers. For example, a spatially separable convolution, such as the one shown in FIG. 2, can be factorized into two components: (1) a depthwise convolution (e.g., spatial fusion), where each spatial channel is independently convolved with a depthwise convolution, and (2) a pointwise convolution (e.g., channel fusion), where all spatial channels are linearly combined. An example of a depthwise separable convolution is shown in FIG. 3A and FIG. 3B. In general, during spatial fusion, the network learns features from the spatial plane, and during channel fusion, the network learns the relationships between these features across channels.

[0040]

[0053] In one example, deep unit separable convolution can be implemented using a 5×5 kernel for spatial fusion and a 1×1 kernel for channel fusion. In particular, channel fusion can use a 1×1×d kernel that iterates through every single point in the input image of depth d, where the depth d of the kernel generally matches the number of channels of the input image. Channel fusion via pointwise convolution is useful for dimensionality reduction for efficient computation. Applying a 1×1×d kernel and adding an activation layer after the kernel can give the network added depth, which can increase the performance of the network.

[0041]

[0054] In particular, in FIG. 3A, a 12 pixel by 12 pixel by 3 channel input image 302 is convolved with a filter comprising three separate kernels 304A-304C, each having a 5 by 5 by 1 dimensionality, to generate an 8 pixel by 8 pixel by 3 channel feature map 306, where each channel is generated by a respective kernel in kernels 304A-304C.

[0042]

[0055] The feature map 306 is then further convolved using a point-wise convolution operation with a kernel 308 having dimensionality 1×1×3 to generate an 8 pixel×8 pixel×1 channel feature map 310. As shown in this example, the feature map 310 has reduced dimensionality (1 channel vs. 3 channels), which together allows for more efficient computation.

[0043]

[0056] The results of the deeply separable convolution in FIGS. 3A and 3B are substantially similar to the classical convolution in FIG. 2, but the number of computations is significantly reduced, and thus the deeply separable convolution provides significant efficiency gains when the network design allows it.

[0044]

[0057] Although not shown in FIG. 3B, multiple (e.g., m) pointwise convolution kernels 308 (e.g., individual components of a filter) may be used to increase the channel dimensionality of the convolution output. Thus, for example, m=256 1×1×3 kernels 308 may be generated, each output being an 8 pixel×8 pixel×1 channel feature map (e.g., feature map 310), which may be stacked to obtain a resulting feature map of 8 pixels×8 pixels×256 channels. The resulting increase in channel dimensionality provides more parameters for training, which may improve the ability of the convolutional neural network to identify features (e.g., in the input image 302). Exemplary Compute-in-Memory (CIM) Architecture

[0058] 4 illustrates an exemplary static random access memory (SRAM) memory cell 400 that may be implemented in a CIM array. Memory cell 400 is sometimes referred to as an eight-transistor (8T) SRAM cell because memory cell 400 is implemented with eight transistors.

[0045]

[0059] As shown, the memory cell 400 may include a cross-coupled inverter pair 424 having an output 414 and an output 416. As shown, the cross-coupled inverter pair output 414 is selectively coupled to a write bit line (WBL) 406 via a pass gate transistor 402, and the cross-coupled inverter pair output 416 is selectively coupled to a complementary write bit line (WBLB) 420 via a pass gate transistor 418. The WBL 406 and WBLB 420 are configured to provide complementary digital signals to be written (e.g., stored) to the cross-coupled inverter pair 424. The WBL and WBLB may be used to store bits for neural network weights in the memory cell 400. The gates of the pass gate transistors 402, 418 may be coupled to a write word line (WWL) 404 as shown. For example, a digital signal to be written may be provided to the WBL (and the complement of the digital signal is provided to the WBLB). The pass gate transistors 402, 418, here implemented as n-type field effect transistors (NFETs), are then turned on by providing a logic high signal to the WWL 404, causing the digital signal to be stored in the cross-coupled inverter pair 424.

[0046]

[0060] As shown, the cross-coupled inverter pair output 414 may be coupled to the gate of transistor 410. The source of transistor 410 may be coupled to a reference potential node (VSS or electrical ground) and the drain of transistor 410 may be coupled to the source of transistor 412. The drain of transistor 412 may be coupled to a read bit line (RBL) 422 as shown. The gate of transistor 412 may be controlled via a read word line (RWL) 408. The RWL 408 may be controlled via an activation input signal.

[0047]

[0061] During a read cycle, RBL 422 may be precharged to a logic high. If both the activation input and the weight bit stored at cross-coupled inverter pair output 414 are logic high, then transistor 410 and transistor 412 are both turned on, electrically coupling RBL 422 to VSS at the source of transistor 410 and discharging RBL 422 to a logic low. If either the activation input or the weight bit stored at cross-coupled inverter pair output 414 are logic low, then at least one of transistors 410, 412 will be turned off, and thus RBL 422 will remain logic high. Thus, the output of memory cell 400 at RBL 422 is logic low only when both the weight bit and the activation input are logic high, and logic high otherwise, effectively implementing a NAND gate operation.

[0048]

[0062] 5A illustrates a circuit 500 for a CIM in accordance with some aspects of the present disclosure. The circuit 500 includes word lines (also called rows) 504. 0 ~504 31 and column 506 0 ~506 7 The CIM array 501 includes a word line 504. 0 ~504 31 are collectively referred to as word lines (WL) 504, and columns 506 0 ~506 7 506. As shown, the CIM array 501 may include an activation circuit 590 configured to provide activation signals to the word lines 504. The CIM array 501 is implemented with 32 word lines and 8 columns for ease of understanding, although the CIM array may be implemented with any number of word lines or columns. As shown, memory cells 502 (collectively referred to as memory cells 502) 0-0 ~502 31-7 is implemented at the intersection of WL 504 and column 506.

[0049]

[0063] Each of the memory cells 502 may be implemented using the memory cell architecture described with respect to FIG. 4. As shown, activation inputs a(0,0) through a(31,0) may be provided to respective word lines 504, and the memory cells 502 may store neural network weights w(0,0) through w(31,7). For example, memory cells 502 may 0-0 ~502 0-7 can store weight bits w(0,0) to w(0,7), and memory cell 502 1-0 ~502 1-7 may store weight bits w(1,0) through w(1,7), and so on. Each word line may store multi-bit weights. For example, weight bits w(0,0) through w(0,7) represent 8 bits of weights for a neural network.

[0050]

[0064] As shown, circuit 500 includes adder trees 510 (collectively referred to as adder trees 510), each implemented for a respective one of columns 506. 0 ~510 7 Each of the adder trees 510 sums the output signals from the memory cells 502 on a respective one of the columns 506. Each adder tree is implemented using a tree of adder circuits, such as adder circuit 511. The outputs of the adder trees 510 are coupled to weight-shifting adder tree circuit 512 as shown. The weight-shifting adder tree circuit 512 includes multiple weight-shifting adders (e.g., weight-shifting adder 514), each including bit-shifting and addition circuitry to facilitate implementation of a bit-shift and addition operation. In other words, the output signals from the memory cells 502 on a respective one of the columns 506 are summed. Each adder tree is implemented using a tree of adder circuits, such as adder circuit 511. The outputs of the adder trees 510 are coupled to weight-shifting adder tree circuit 512 as shown. The weight-shifting adder tree circuit 512 includes multiple weight-shifting adders (e.g., weight-shifting adder 514), each including bit-shifting and addition circuitry to facilitate implementation of a bit-shift and addition operation. 0 The upper memory cell may store the most significant bit (MSB) for each weight, and column 506 7 The upper memory cell may store the least significant bit (LSB) for each weight. Thus, when performing summation across columns 506, a bit shifting operation is performed to shift the bits to take into account the importance of the bits on the associated column.

[0051]

[0065] The output of the weight shift adder tree circuit 512 is provided to an activation shift accumulator circuit 516. The activation shift accumulator circuit 516 includes a bit shift circuit 518 and an accumulator 520. The activation shift accumulator circuit 516 may also include a flip-flop (FF) 522 and a FF 591.

[0052]

[0066] During operation of the circuit 500, the activation circuit 590 provides a first set 599 of activation inputs a(0,0) to a(31,0) to the memory cells 502 for calculation during a first activation cycle. The first set of activation inputs a(0,0) to a(31,0) represent the most significant bits of the activation parameter. The outputs of the calculation on each column are summed using a respective one of the adder trees 510. The outputs of the adder trees 510 are summed using the weight shift adder tree circuit 512, and the result is provided to the activation shift accumulator. The same operation is performed for other sets of activation inputs during subsequent activation cycles, such as activation inputs a(0,1) to a(31,1) representing the second most significant bits of the activation parameter, and so on, until the activation inputs representing the least significant bits of the activation parameter are processed. The bit shift circuit 518 performs a bit shift operation based on the activation cycle. For example, for an 8-bit activation parameter that is processed using eight activation cycles, the bit shift circuit may perform an 8-bit shift for the first activation cycle, a 7-bit shift for the second activation cycle, etc. After an activation cycle, the output of the bit shift circuit 518 is accumulated using an accumulator 520 and stored in FFs 522, 591, which may implement a transfer register.

[0053]

[0067] The architecture of circuit 500 is referred to as a "folded" architecture due to the symmetrical structure of the processing circuits, such as weight shifting adder tree circuit 512. The folded architecture allows for configurability of the number of bits associated with the weights used during a calculation. For example, instead of a calculation using 8-bit weights, a calculation using 4-bit weights may be implemented by deactivating (deactivating) four of the columns 506, as described in more detail herein.

[0054]

[0068] The embodiment described with respect to FIG. 5A provides bitwise storage and bitwise multiplication. The adder trees 510 perform population count addition for the columns 506. That is, each of the adder trees 510 adds the output signals of the memory cells for the columns. The weight shift adder tree circuit 512 (e.g., having three stages as shown for eight columns) combines the weighted sums generated for the eight columns (e.g., provides an accumulation result for a given activation bit position during an activation cycle). The activation shift accumulator circuit 516 combines the results from multiple (e.g., eight) activation cycles and outputs a final accumulation result. For example, the bit shift circuit 518 shifts the bits at the output of the weight shift adder tree circuit 512 based on the associated activation cycle. The serial accumulator 520 accumulates the shifted adder outputs generated by the bit shift circuit 518. A transfer register implemented using FFs 522, 591 copies the output of the serial accumulator 520 after the calculation for the last activation cycle is completed.

[0055]

[0069] Parallel addition across columns increases the processing performance (in Tera Operations Per Second (TOPS)) associated with circuit 500, provides a more compact full adder cell, reduces parasitic penalties since the adders are implemented next to the bit multiplication memory cells, reduces switching activity since fewer rows of memory have higher activation amplitudes compared to conventional implementations, and provides easy tiling allowing easy macro generation since the cells are placed side-by-side in an abutting configuration for realization of an adder tree. The aspects described with respect to FIG. 5A may be implemented at a single clock frequency.

[0056]

[0070] The circuit 500 provides linear energy scaling across computations using different bit sizes of activation or weight parameters. In other words, using the adder tree 510 and weight shift adder tree circuit 512 as described herein provides bit size configurability, allowing n-bit activation with m-bit weight accumulation, where n and m are positive integers. The energy consumption associated with the circuit 500 scales linearly based on the configured bit sizes for the activation parameters and weights.

[0057]

[0071] 5B illustrates an exemplary implementation of an adder circuit 585 configured to perform an add operation. The adder circuit 585 may correspond to any of the adder circuits described herein, such as the adder circuit 511. As shown, the adder circuit includes an exclusive-OR (XOR) gate 573 that receives inputs 570, 571 (labeled A and B). The output of the XOR gate 573 is provided to an input of an XOR gate 574, the other input of which receives a carry-in (Cin) signal 572. The output of the XOR gate 574 provides an output (labeled SUM) of the adder circuit. As shown, the adder circuit may also include an AND gate 575 that receives the Cin signal 572 and the output of the XOR gate 573. The AND gate 576 receives the inputs 570, 571. The outputs of AND gates 575, 576 are provided to inputs of OR gate 578, which generates a carry out signal for the addition operation. Although Figure 5B shows one example implementation of an adder circuit for ease of understanding, the aspects described herein may be implemented using any suitable adder circuit architecture.

[0058]

[0072] 5C is an exemplary implementation of an accumulator 587. The accumulator 587 may correspond to any of the accumulators described herein, such as accumulator 520. As shown, the accumulator 520 includes an adder circuit 580 that receives an input signal, as shown. An output of the adder circuit 580 is provided to a register configured to store the output of the adder circuit at each cycle of a clock signal provided to register 581. An output 582 of the register 581 is used as an output of the accumulator 587 and is fed back to an input of the adder circuit 580, as shown. Although FIG. 5C illustrates one exemplary implementation of an accumulator for ease of understanding, aspects described herein may be implemented using any suitable accumulator architecture.

[0059]

[0073] 6 illustrates a circuit 600 for a CIM implemented using a bit string adder tree circuit 650 and a column accumulator circuit 652 in accordance with some aspects of the disclosure. The bit string adder tree circuit 650 includes a number of sense amplifiers 602 connected to a number of columns 506. 0 , 602 1 , ~602 7 5, where each column has a number of bit lines (e.g., RBL). For example, each of the columns 506 may have four bit lines, each of which is connected to four sense amplifiers (e.g., sense amplifiers 602, 603, 604, 605, 606, 607, 608, 609, 610, 611, 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 64 0 ) are coupled to the inputs of one of the four bit lines on each column. The word lines of the CIM array 501 may include multiple word line groups (e.g., eight groups), each group having four word lines. Each word line of each group of four word lines is coupled to a respective one of the four bit lines on each column. Each group of four word lines is activated by a corresponding activation signal in a given computation cycle, and the activation signals of the remaining word line groups are set to logic low. The word line groups (e.g., eight word line groups for a total of 32 word lines in this example) are processed in a total of eight clock cycles. Sense amplifiers 602 0 , 602 1 , ~602 7 are collectively referred to as sense amplifiers 602. Multiple sense amplifiers (e.g., four) are included for each of the columns 506, allowing multiple columns to be sensed simultaneously. For example, sense amplifiers 602 0 is column 506 0 Each of the memory cells 502 0-0 ~502 3-0 and simultaneously sense the outputs of sense amplifier 602. 1 is column 506 1 Each of the memory cells 502 0-1 ~502 3-1 Detect the output of column 506 at the same time. 7 Each of the memory cells 502 0-7 ~502 3-7 Sense amplifier 602 which simultaneously senses the outputs of 7The output of the sense amplifier for each column is fed to an adder tree (e.g., adder trees 604, collectively referred to as adder trees 604). 0 , 604 1 , ~604 7 ) Each of the adder circuits used to implement each of the adder trees 604 may be implemented as described with respect to FIG.

[0060]

[0074] For simplicity, each of the sense amplifiers 602 is shown as having an input coupled to the output of a single memory cell. However, the input of each of the sense amplifiers 602 may be coupled to the outputs of multiple memory cells, which may be activated in a serial manner. In other words, if there are four sense amplifiers for each column, then four word lines may be activated at one time on each column. As an example, the sense amplifiers 602 0 The inputs are connected to a first group of word lines as shown (e.g., word lines 504 0 ~504 3 ), but also coupled to the outputs of the respective memory cells of a second group of word lines (e.g., word lines 504 4 ~504 7 ) and a third group of word lines (e.g., word lines 504 8 ~504 11 ) and coupled to the outputs of the respective memory cells for the last group of word lines (e.g., word line 504 28 ~504 31 ), and so on. Thus, for 32 word lines and four sense amplifiers per column, eight calculation cycles may be used to complete the calculations for a set of activation inputs (e.g., activation inputs a(0,0) through a(31,0)).

[0061]

[0075] As illustrated, the outputs of the adder trees 604 are coupled to a column accumulator circuit 652. For example, the output of each of the adder trees 604 may be coupled to an accumulator 606 (collectively referred to as accumulators 606) of the column accumulator circuit 652.0 , 606 1 , ~606 7 Each of the accumulators 606 may be implemented as described with respect to FIG. 5C. Each of the accumulators 606 performs an accumulation of an output signal of a respective one of the adder trees 604 over multiple calculation cycles. For example, during each calculation cycle, a calculation is performed for four word lines, and the output signals of the calculation for the four word lines are added using the adder tree 604 of the bit string adder tree circuit 650. After multiple calculation cycles (e.g., eight cycles for 32 word lines when using four sense amplifiers), each of the accumulators 606 performs an accumulation of an output signal of a respective one of the adder trees 604.

[0062]

[0076] Upon completion of multiple computation cycles, the output of the accumulator 606 is provided to the weight shift adder tree circuit 512 for summation across columns, and the output of the weight shift adder tree circuit 512 is provided to the activation shift accumulator circuit 516 for accumulation across activation cycles, as described with respect to FIG. 5A. In other words, bitwise accumulation is performed in each of the accumulators 606 over multiple computation cycles (e.g., eight computation cycles, each computation cycle for four word lines until computation for 32 word lines is completed). The weight shift adder tree circuit 512 combines the weighted sums of the eight columns (e.g., provides an accumulation result for a given activation bit position during each activation cycle), and the activation shift accumulator circuit 516 combines the results from multiple (e.g., eight) activation cycles to output a final accumulation result. In some aspects, the CIM array 501, bit string adder tree circuit 650, and column accumulator circuit 652 operate at a higher frequency (e.g., 8 times higher when implemented using 8 computation cycles, or less than 8 times higher as determined by critical path delay limits while still using 8 computation cycles) than the weight shift adder tree circuit 512 and activation shift accumulator circuit 516. As shown, half latch circuits 608 (collectively referred to as half latch circuits 608) 0 , 608 1 , ~608 7may be coupled to the outputs of the respective accumulators 606. Each half-latch circuit holds the output of a respective one of the accumulators 606 and provides an output to a respective input of the weight-shifting adder tree circuits 512 upon completion of multiple computation cycles. In other words, a half-latch circuit generally refers to a latch circuit that holds a digital input (e.g., the output of one of the accumulators 606) at the beginning of a clock cycle and provides the digital input to the output of the latch circuit at the end of the clock cycle. The half-latch circuits 608 facilitate a transition from a higher frequency operation of the column accumulator circuits 652 (e.g., at 8× as shown) to a lower frequency operation of the weight-shifting adder tree circuits 512 (e.g., at 1× as shown).

[0063]

[0077] FIG. 7 is a timing diagram 700 illustrating signals associated with the circuit 600 according to some aspects of the disclosure. The circuit 600 may run on a digital compute in memory (DCIM) clock. The DCIM clock may be used as the main clock on which the circuits 500, 600 run. After eight cycles of the DCIM clock, the final accumulated output may be provided for weight multiplication with the 8-bit activation input. As shown, a higher frequency clock signal, called a local clock, may be generated from the lower frequency DCIM clock. For example, the local clock may have a frequency eight times greater than the frequency of the DCIM clock.

[0064]

[0078] As shown, one bit of each of the activation inputs is provided during each of eight cycles of the DCIM clock. For example, bits a(0,0) through a(31,0) (e.g., the MSBs of different activation inputs) are provided to the memory cells during a first activation cycle (e.g., the first cycle of the DCIM clock), bits a(0,1) through a(31,1) (e.g., the second MSB (MSB-1) of the different activation input) are provided to the memory cells during a second activation cycle (e.g., the second cycle of the DCIM clock), and so on.

[0065]

[0079] During each cycle of the local clock, the output of sense amplifier 602 (labeled "SA Out") and the output of adder tree 604 (labeled "Col Add Out") are provided for a calculation cycle. During each cycle of the local clock, SA Out and Col Add Out provide outputs for memory cells of a subset of word lines 504 (e.g., for four word lines in the example described with respect to FIG. 6). For example, during the first cycle of the local clock, adder tree 604 0 About Col Add Out Memory Cell 502 0-0 ~502 3-0 During the second cycle of the local clock, Col Add Out is stored in memory cell 502. 4-0 ~502 7-0 , where Col Add Out is the current value of memory cell 502 after eight cycles of the local clock. 28-0 ~502 31-0 and so on until calculations performed by

[0066]

[0080] As shown, the output of the column accumulator circuit 652 (labeled "Col Acc Latch") and the output of the weight shift adder tree circuit 512 (labeled "Weight Shift Add Out") are provided after eight local clock cycles (e.g., after a single DCIM clock cycle). The activation shift accumulator circuit 516 accumulates the Weight Shift Add Out over the eight DCIM clock cycles and provides an output (labeled "Acc Out") at the end of the eight DCIM clock cycles.

[0067]

[0081] In some aspects, the number of bits associated with the activation inputs and / or weights may be configurable. The bit string adder tree circuit 650 allows for configurability of the number of bits for the weights, down to a single bit. For example, to implement a 4-bit weight, the bit string adder tree circuit 650 may add up to 16 bits for the weights, as described in more detail herein. 4 , 506 5 , 506 6 , 506 7 The circuitry associated with may be deactivated.

[0068]

[0082] 8A, 8B, and 8C are block diagrams illustrating CIM circuits with configurable bit sizes of weights according to some aspects of the disclosure. For example, as shown in FIG. 8A, 8-bit weights may be stored in memory cells 502 and processed using bit string adder tree circuit 650, string accumulator circuit 652, weight shift adder tree circuit 512, and activation shift accumulator circuit 516 as described herein.

[0069]

[0083] As shown, the clock generator circuit 870 may include a clock generator 871 configured to generate a DCIM clock. The clock generator 871 may be implemented using any suitable clock generation circuit, such as a phase-locked loop (PLL) or a ring oscillator. The weight shift adder tree circuit 512 may receive and operate on the DCIM clock described with respect to FIG. 7. In some aspects, the clock generator circuit 870 may include a frequency multiplier 802 that may be used to generate a local clock on which the activation circuit 590, the bit string adder tree circuit 650, and the column accumulator circuit 652 operate. Although the frequency multiplier 802 is shown as being part of the clock generator circuit 870, the frequency multiplier 802 may be separate from the clock generator 871 in some implementations. A frequency multiplier generally refers to any circuit that receives a clock signal having a first frequency and generates a second clock signal having a second, different frequency, where the second frequency is a multiple of the first frequency.

[0070]

[0084] Some aspects provide a computation technique that uses wing-serial operation, as described with respect to Figures 8B and 8C. In the case of a CIM circuit, "wing-serial operation" as used herein generally refers to operating on one wing (one processing path of the CIM circuit) and then operating on another wing (another processing path of the CIM circuit). For example, if a 4-bit weight is to be used to perform a first 4-bit weight calculation, a set of four columns (e.g., columns 506, 507, 509, 510, 511, 512, 513, 514, 515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584 4 , 506 5 , 506 6 , 506 7 ), and the 4-bit weights may be stored in memory cells on another set of four columns (e.g., columns 506, 0 , 506 1 , 506 2 , 506 3) in memory cells on the left side of the four columns. The two sets of four columns may be independent sets. In the example provided, the four-bit weight calculations are stored in columns 506 and 507. 0 , 506 1 , 506 2 , 506 3 Before being performed on the 4-bit weight calculation, the 4-bit weight calculation is performed on column 506 4 , 506 5 , 506 6 , 506 7 For example, a 4-bit weight calculation is performed on column 506 4 , 506 5 , 506 6 , 506 7 Before being performed on the 4-bit weight calculation, the 4-bit weight calculation is performed on column 506 0 , 506 1 , 506 2 , 506 3 The method may be carried out on the basis of the above.

[0071]

[0085] FIG. 8B illustrates a first cycle during which a first 4-bit weight calculation is performed. During the first cycle, column 506 4 , 506 5 , 506 6 , 506 7 The circuits of the bit string adder tree circuit 650 and the column accumulator circuit 652 used for processing signals for are deactivated. For example, the clock gating circuit 804 may deactivate the accumulator 606 during the first cycle to reduce power consumption. 4 , 606 5 , 606 6 , 606 7 A clock gating circuit, as used herein, generally refers to any circuit that receives a clock signal (e.g., an AND gate having a first input receives the clock signal) and provides a clock signal at an output of the circuit in response to a control signal (e.g., the control signal provided to a second input of the AND gate is a logic high). At the end of the first cycle, the activation shift accumulator circuit 516 provides a result for the first 4-bit weight calculation.

[0072]

[0086] FIG. 8C illustrates a second cycle during which a second 4-bit weight calculation is performed. During the second cycle, column 506 0 , 506 1 , 506 2 , 506 3 The circuits of the bit string adder tree circuit 650 and the column accumulator circuit 652 used for processing signals for are deactivated. For example, the clock gating circuit 804 may deactivate the accumulator 606 during the second cycle to reduce power consumption. 0 , 606 1 , 606 2 , 606 3 Although the clock gating technique is only shown with respect to the clock signal to the column accumulator circuit 652 for ease of understanding, the clock gating technique may be used to deactivate clock signals to other circuits that are unused, such as the circuitry of a bit string adder tree. Exemplary Operations for Digital Computation-in-Memory (CIM)

[0087] 9 is a flow diagram illustrating an example operation 900 for in-memory computing according to some aspects of the disclosure. The operation 900 may be performed by a circuit for a CIM, such as the circuit 500 described with respect to FIG. 5A or the circuit 600 described with respect to FIG.

[0073]

[0088] The operation 900 begins at block 905 by the circuit summing output signals on a respective one of a number of columns (e.g., columns 506) of a memory through each of a number of summing circuits (e.g., adder tree 510 or accumulator 606). A number of memory cells are on each of the number of columns, and the number of memory cells store a number of bits representing weights of a neural network (e.g., w(0,0) through w(31,7) shown in FIG. 5A). The number of memory cells on each of the number of columns are on different word lines (e.g., word line 504) of the memory.

[0074]

[0089] At block 910, the circuit adds output signals of at least two of the plurality of adder circuits via a first adder circuit (e.g., weight shifting adder tree circuit 512). At block 915, the circuit accumulates output signals of the first adder circuit via an accumulator (e.g., accumulator 520 or activation shift accumulator circuit 516). In some aspects, the circuit selectively disables one or more portions of the first adder circuit and / or one or more of the plurality of adder circuits based on the number of bits associated with each of the weights.

[0075]

[0090] In some aspects, adding the output signals on each one of the multiple columns may include accumulating (e.g., via accumulator 606) the output signals of the memory cells on each one of the multiple columns after two or more of the word lines are successively activated. In some aspects, the circuit adds the output signals of the memory cells on each one of the multiple columns and two or more of the word lines via a second adder circuit (e.g., each of adder trees 604) coupled between each of the multiple adder circuits and the respective one of the multiple columns. In some aspects, the circuit senses the output signals of the memory cells on each one of the multiple columns and two or more of the word lines via a sense amplifier (e.g., sense amplifier 602) coupled between the second adder circuit and the respective one of the multiple columns. In this case, the summing via the second adder circuit is based on the sensed output signals.

[0076]

[0091] In some aspects, the circuitry disables a first portion of the first adder circuit and / or at least one of the adder circuits during a first calculation cycle and disables a second portion of the first adder circuit and at least another one of the adder circuits during a second calculation cycle.

[0077]

[0092] In some aspects, the circuitry sequentially activates two or more of the word lines, in which case summing the output signals on a respective one of the multiple columns via each of the multiple summing circuits includes accumulating the output signals of memory cells on a respective one of the multiple columns via each of the multiple summing circuits (e.g., accumulator 606) after two or more of the word lines are sequentially activated.

[0078]

[0093] In some aspects, the summing of the output signals of at least two of the plurality of summing circuits includes performing a bit shift and add operation on at least two of the plurality of summing circuits. In some aspects, the circuit generates a first clock signal, where the plurality of summing circuits operate based on the first clock signal (e.g., the local clock shown in FIG. 7), and the circuit generates a second clock signal, where the first adder circuit operates based on the second clock signal (e.g., the DCIM clock shown in FIG. 7), and the second clock signal has a different frequency than the first clock signal. In some aspects, the circuit generates the second clock signal based on the first clock signal via a frequency multiplier (e.g., frequency multiplier 802).

[0079]

[0094] In some aspects, the circuit sequentially activates the plurality of memory cells based on different activation inputs, and accumulating the output signal of the first adder circuit occurs after the plurality of memory cells are sequentially activated. For example, sequentially activating the plurality of memory cells may include receiving a first set of activation inputs (e.g., activation inputs a(0,0) through a(31,0)) during a first activation cycle and receiving a second set of activation inputs (e.g., activation inputs a(0,1) through a(31,1)) during a second activation cycle, where accumulating the output signal of the first adder circuit occurs after the first activation cycle and the second activation cycle.

[0080]

[0095] In some aspects, the plurality of columns may include a first subset of the plurality of columns (e.g., columns 506 0 ~506 3 ), and a second subset of columns (for example, column 506 4 ~506 7 ). The first subset may be activated during a first computation cycle (e.g., cycle 1 shown in FIG. 8B). The second subset may be activated during a second computation cycle (e.g., cycle 2 shown in FIG. 8C), the second computation cycle being after the first computation cycle.

[0081]

[0096] In some aspects, the memory cells on each of the word lines are configured to store one of the weights of the neural network, and the amount of the first subset of the plurality of columns (e.g., four in the example shown in FIG. 8B) is related to the amount of bits of one of the weights. In some aspects, the circuit deactivates a clock signal associated with processing signals from a second subset of the plurality of columns via a clock gating circuit (e.g., clock gating circuit 804). Exemplary Processing System for Computation in Memory

[0097] 10 illustrates an example electronic device 1000. The electronic device 1000 may be configured to perform the methods described herein, including the operations 900 described with respect to FIG.

[0082]

[0098] The electronic device 1000 includes a central processing unit (CPU) 1002, which in some aspects may be a multi-core CPU. Instructions executed in the CPU 1002 may be loaded, for example, from a program memory associated with the CPU 1002 or may be loaded from the memory 1024.

[0083]

[0099] The electronic device 1000 also includes additional processing blocks adapted to specific functions, such as a graphics processing unit (GPU) 1004, a digital signal processor (DSP) 1006, a neural processing unit (NPU) 1008, a multimedia processing block 1010, and a wireless connectivity processing block 1012. In one implementation, the NPU 1008 may be implemented in one or more of the CPU 1002, the GPU 1004, and / or the DSP 1006.

[0084]

[0100] In some aspects, the wireless connectivity processing block 1012 may include components for, for example, third generation (3G) connectivity, fourth generation (4G) connectivity (e.g., 4G LTE), fifth generation connectivity (e.g., 5G or NR), Wi-Fi connectivity, Bluetooth connectivity, and wireless data transmission standards. The wireless connectivity processing block 1012 is further connected to one or more antennas 1014 to facilitate wireless communication.

[0085]

[0101] The electronic device 1000 may also include one or more sensor processors 1016 associated with any type of sensor, one or more image signal processors (ISPs) 1018 associated with any type of image sensor, and / or a navigation processor 1020, which may include satellite-based positioning system components (e.g., GPS or GLONASS) as well as inertial positioning system components.

[0086]

[0102] The electronic device 1000 may also include one or more input and / or output devices 1022, such as a screen, a touch-sensitive surface (including a touch-sensitive display), physical buttons, a speaker, a microphone, etc. In some aspects, one or more of the processors of the electronic device 1000 may be based on the ARM instruction set.

[0087]

[0103] The electronic device 1000 also includes memory 1024, which represents one or more static and / or dynamic memories, such as dynamic random access memories, flash-based static memories, etc. In this example, the memory 1024 includes computer-executable components that may be executed by one or more of the above-mentioned processors or CIM controllers 1032 (also referred to as control circuits) of the electronic device 1000. For example, the electronic device 1000 may include a CIM circuit 1026, such as circuit 500, as described herein. The CIM circuit 1026 may be controlled via the CIM controller 1032. For example, in some aspects, the memory 1024 may include code 1024A for storing (e.g., storing weights in memory cells) and code 1024B for computing (e.g., performing neural network computations by applying activation inputs). As shown, the CIM controller 1032 may include circuitry 1028A for storing (e.g., storing weights in memory cells) and circuitry 1028B for computing (e.g., performing neural network computations by applying activation inputs). The components shown, as well as other components not shown, may be configured to implement various aspects of the methods described herein.

[0088]

[0104] In some aspects, such as when the electronic device 1000 is a server device, various aspects may be omitted from the example shown in FIG. 10, such as one or more of the multimedia processing block 1010, the wireless connectivity processing block 1012, the antenna 1014, the sensor processor 1016, the ISP 1018, or the navigation processor 1020. Example clauses

[0105] Clause 1. A circuit for in-memory computation comprising: a plurality of memory cells on each of a plurality of columns of a memory, the plurality of memory cells configured to store a plurality of bits representing weights of a neural network, where the plurality of memory cells on each of the plurality of columns are on different word lines of the memory; a plurality of adder circuits, each coupled to a respective one of the plurality of columns; a first adder circuit coupled to outputs of at least two of the plurality of adder circuits; and an accumulator coupled to an output of the first adder circuit.

[0089]

[0106] Clause 2. The circuit of clause 1, wherein one or more portions of the first adder circuit are configured to be selectively disabled.

[0090]

[0107] Clause 3. The circuit of any one of clauses 1-2, wherein each of the plurality of adder circuits comprises an adder tree coupled to a plurality of memory cells on a respective one of the plurality of columns.

[0091]

[0108] Clause 4. The circuit of any one of clauses 1 to 3, wherein each of the plurality of summing circuits comprises a separate accumulator.

[0092]

[0109] Clause 5. The circuit of any one of clauses 1-4, wherein a first portion of the first adder circuit is configured to be selectively disabled during a first calculation cycle, and a second portion of the first adder circuit is configured to be selectively disabled during a second calculation cycle.

[0093]

[0110] Clause 6. The circuit of any one of clauses 1-5, further comprising a second adder circuit coupled between each of the plurality of adder circuits and a respective one of the plurality of columns.

[0094]

[0111] Clause 7. The circuit of clause 6, wherein the second adder circuit comprises an adder tree coupled to two or more of the word lines.

[0095]

[0112] Clause 8. The circuit of clause 7, wherein the adder tree is configured to add output signals of memory cells on respective ones of the plurality of columns and two or more of the word lines.

[0096]

[0113] Clause 9. The circuit of clause 6, further comprising a sense amplifier coupled between the second summer circuit and each one of the plurality of columns.

[0097]

[0114] Clause 10. The circuit of any one of clauses 1-9, wherein the first adder circuit comprises an adder tree configured to add output signals of at least two of the plurality of adder circuits.

[0098]

[0115] Clause 11. The circuit of clause 10, wherein one or more adders of the adder tree comprise bit shifting and adding circuits.

[0099]

[0116] Clause 12. The circuit of any one of clauses 1-11, further comprising a clock generator circuit having a first output configured to output a first clock signal and a second output configured to output a second clock signal, wherein a plurality of adder circuits are coupled to the first output of the clock generator and configured to operate based on the first clock signal, and a first adder circuit is coupled to the second output of the clock generator and configured to operate based on the second clock signal, the second clock signal having a different frequency than the first clock signal.

[0100]

[0117] Clause 13. The circuit of clause 12, wherein the clock generator circuit comprises a frequency multiplier configured to generate the second clock signal based on the first clock signal.

[0101]

[0118] Clause 14. The circuit of any one of clauses 1-13, further comprising a plurality of half-latch circuits, each half-latch circuit coupled between the first adder circuit and one of the plurality of adder circuits.

[0102]

[0119] Clause 15. The circuit of any one of clauses 1-14, wherein the plurality of memory cells are configured to be successively activated based on different activation inputs, and the accumulator is configured to accumulate an output signal of the first adder circuit after the plurality of memory cells are successively activated.

[0103]

[0120] Clause 16. The circuit of any one of clauses 1-15, wherein the accumulator is the only accumulator coupled to the output of the first adder circuit.

[0104]

[0121] Clause 17. The circuit of any one of clauses 1-16, wherein the plurality of columns comprises a first subset of the plurality of columns and a second subset of the plurality of columns, the first subset being activated during a first computation cycle.

[0105]

[0122] Clause 18. The circuit of clause 17, wherein the second subset is activated during a second computation cycle, the second computation cycle being after the first computation cycle.

[0106]

[0123] Clause 19. The circuit of any one of clauses 17-18, wherein at least some of the memory cells on each of the word lines are configured to store one of the weights of a neural network, and a quantity of a first subset of the plurality of columns is related to a quantity of bits of one of the weights.

[0107]

[0124] Clause 20. The circuit of any one of clauses 17-19, further comprising a clock gating circuit having an output coupled to the plurality of summing circuits and configured to deactivate a clock signal associated with processing signals from a second subset of the plurality of columns.

[0108]

[0125] Clause 21. A method for in-memory computation comprising: summing output signals on a respective one of a plurality of columns of a memory via each of a plurality of summing circuits, wherein a plurality of memory cells are on each of the plurality of columns, the plurality of memory cells storing a plurality of bits representing weights of a neural network, and wherein the plurality of memory cells on each of the plurality of columns are on different word lines of the memory; summing at least two output signals of the plurality of summing circuits via a first adder circuit; and accumulating the output signals of the first adder circuit via an accumulator.

[0109]

[0126] Clause 22. The method of clause 21, further comprising selectively disabling one or more portions of the first adder circuit based on a number of bits associated with each of the weights.

[0110]

[0127] Clause 23. The method of any one of clauses 21-22, wherein adding the output signals on each one of the plurality of columns comprises accumulating output signals of memory cells on each one of the plurality of columns after two or more of the word lines are successively activated.

[0111]

[0128] Clause 24. The method of clause 23, further comprising summing output signals of memory cells on respective ones of the plurality of columns and two or more of the word lines via a second adder circuit coupled between each of the plurality of summing circuits and the respective one of the plurality of columns.

[0112]

[0129] Clause 25. The method of clause 24, further comprising sensing output signals of memory cells on each one of the plurality of columns and two or more of the word lines via a sense amplifier coupled between a second adder circuit and each one of the plurality of columns, wherein the summing via the second adder circuit is based on the sensed output signals.

[0113]

[0130] Clause 26. The method of any one of clauses 21-25, wherein adding output signals of at least two of the plurality of summation circuits comprises performing a bit shift and add operation on at least two of the plurality of summation circuits.

[0114]

[0131] Clause 27. The method of any one of clauses 21-26, further comprising: generating a first clock signal, wherein a plurality of adder circuits operate based on the first clock signal; and generating a second clock signal, wherein the first adder circuits operate based on the second clock signal, and the second clock signal has a different frequency than the first clock signal.

[0115]

[0132] Clause 28. The method of any one of clauses 21-27, further comprising sequentially activating a plurality of memory cells based on different activation inputs, wherein accumulating the output signal of the first adder circuit occurs after the plurality of memory cells have been sequentially activated.

[0116]

[0133] Clause 29. The method of clause 28, wherein successively activating the plurality of memory cells comprises receiving a first set of activation inputs during a first activation cycle and receiving a second set of activation inputs during a second activation cycle, wherein accumulating the output signal of the first adder circuit occurs after the first activation cycle and the second activation cycle.

[0117]

[0134] Clause 30. An apparatus for in-memory computation comprising: first means for adding output signals on a respective one of a plurality of columns of a memory, wherein a plurality of memory cells are on each of the plurality of columns, the plurality of memory cells storing a plurality of bits representing weights of a neural network, wherein the plurality of memory cells on each of the plurality of columns are on different word lines of the memory; second means for adding at least two output signals of the first means for adding; and means for accumulating output signals of the second means for adding. Additional Considerations

[0135] The above description is provided to enable those skilled in the art to practice the various aspects described herein. The examples described herein are not intended to limit the scope, applicability, or aspects described in the claims. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. For example, changes may be made in the function and configuration of the elements described without departing from the scope of the disclosure. Various examples may omit, substitute, or add various procedures or components, as appropriate. For example, the methods described may be performed in an order different from that described, and various steps may be added, omitted, or combined. Also, features described with respect to some examples may be combined in some other examples. For example, an apparatus may be implemented, or a method may be practiced, using any number of the aspects described herein. Furthermore, the scope of the disclosure is intended to cover such apparatus or methods implemented using other structures, functions, or structures and functions in addition to or other than the various aspects of the disclosure described herein. It should be understood that any aspect of the disclosure disclosed herein may be implemented by one or more elements of a claim.

[0118]

[0136] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects.

[0119]

[0137] As used herein, a phrase referring to "at least one of" a list of items refers to any combination of those items, including single members. As an example, "at least one of a, b, or c" is intended to encompass a, b, c, ab, ac, bc, and abc, as well as any combination with multiples of the same element (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc, and ccc, or any other permutation of a, b, and c).

[0120]

[0138] The term "determining" as used herein encompasses a wide variety of actions. For example, "determining" may include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, database, or another data structure), ascertaining, and the like. Also, "determining" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), and the like. Also, "determining" may include resolving, selecting, choosing, establishing, and the like.

[0121]

[0139] The methods disclosed herein comprise one or more steps or actions for achieving the method. The steps and / or actions of the methods may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions may be modified without departing from the scope of the claims. Furthermore, various operations of the methods described above may be performed by any suitable means capable of performing the corresponding functions. Those means may include various (one or more) hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs), or processors. Generally, where there are operations illustrated in the figures, those operations may have corresponding counterpart means-plus-function components with similar numbers. For example, a means for adding may include an adder tree, such as adder tree 510 or weight shift adder tree 512, or an accumulator, such as accumulator 606. A means for accumulating may include an accumulator, such as activation shift accumulator 516. The means for detecting may include an SA, such as SA 602.

[0122]

[0140] The following claims are not limited to the embodiments set forth herein, but are to be accorded the full scope consistent with the language of the claims. In the claims, reference to an element in the singular does not mean "the one and only" unless expressly stated as such, but means "one or more." Unless otherwise expressly stated, the term "some" refers to one or more. No claim element is to be construed under the provisions of 35 U.S.C. 112(f) unless the element is expressly recited using the phrase "means for" or, in the case of a method claim, unless the element is recited using the phrase "step for." All structural and functional equivalents of the elements of the various embodiments described throughout this disclosure that are known or later become known to those skilled in the art are expressly incorporated herein by reference and are encompassed by the claims. Moreover, nothing disclosed herein is made public, regardless of whether such disclosure is expressly recited in the claims.

Claims

1. A plurality of memory cells on each of a plurality of columns of a memory, wherein the plurality of memory cells are configured to store a plurality of bits representing weights of a neural network, and wherein the plurality of memory cells on each of the plurality of columns are on different word lines of the memory, A plurality of adder circuits, each coupled to a respective one of the plurality of columns, and wherein each of the plurality of adder circuits comprises an adder tree coupled to the plurality of memory cells on the respective one of the plurality of columns, A first adder circuit coupled to outputs of at least two of the plurality of adder circuits, An accumulator coupled to an output of the first adder circuit, A circuit for in-memory computing, comprising.

2. The circuit according to claim 1, wherein one or more portions of the first adder circuit are configured to be selectively disabled.

3. The circuit according to claim 1, wherein each of the plurality of adder circuits comprises a separate accumulator.

4. The circuit according to claim 1, wherein a first portion of the first adder circuit is configured to be selectively disabled during a first calculation cycle, and a second portion of the first adder circuit is configured to be selectively disabled during a second calculation cycle.

5. The circuit according to claim 1, further comprising a second adder circuit coupled between each of the plurality of adder circuits and the respective one of the plurality of columns.

6. The second adder circuit comprises an adder tree coupled to two or more of the word lines, Preferably, the adder tree is configured to add output signals of the memory cells on the respective one of the plurality of columns and the two or more of the word lines, the circuit according to claim 5.

7. The circuit according to claim 5, further comprising a sense amplifier coupled between the second adder circuit and the respective one of the plurality of columns.

8. The first adder circuit comprises an adder tree configured to add the output signals of at least two of the plurality of adder circuits, Preferably, one or more adders of the adder tree comprise a bit shift and add circuit, the circuit according to claim 1.

9. further comprising a clock generator circuit having a first output configured to output a first clock signal and a second output configured to output a second clock signal, wherein the plurality of adder circuits are coupled to the first output of the clock generator and are configured to operate based on the first clock signal, the first adder circuit is coupled to the second output of the clock generator and is configured to operate based on the second clock signal, and the second clock signal has a frequency different from that of the first clock signal, preferably, the clock generator circuit comprises a frequency multiplier configured to generate the second clock signal based on the first clock signal, The circuit according to claim 1.

10. further comprising a plurality of half-latch circuits, each half-latch circuit being coupled between the first adder circuit and one of the plurality of adder circuits, or the plurality of memory cells are configured to be continuously activated based on different activation inputs, the accumulator is configured to accumulate the output signal of the first adder circuit after the plurality of memory cells are continuously activated, or the accumulator is the only accumulator coupled to the output of the first adder circuit. The circuit according to claim 1.

11. the plurality of columns comprise a first subset of the plurality of columns and a second subset of the plurality of columns, the first subset is activated during a first calculation cycle, The circuit according to claim 1.

12. the second subset is activated during a second calculation cycle, and the second calculation cycle is after the first calculation cycle, or at least some of the memory cells on each of the word lines are configured to store one of the weights of the neural network, the amount of the first subset of the plurality of columns is related to the amount of bits of the one of the weights, or The circuit according to claim 11, further comprising a clock gating circuit coupled to the plurality of adder circuits and configured to deactivate a clock signal related to processing signals from the second subset of the plurality of columns.

13. Adding the output signals on each of a plurality of columns of a memory via each of a plurality of adder circuits, wherein a plurality of memory cells are on each of the plurality of columns, the plurality of memory cells store a plurality of bits representing weights of a neural network, and the plurality of memory cells on each of the plurality of columns are on different word lines of the memory. Adding at least two output signals of the plurality of adder circuits via a first adder circuit. Accumulating the output signal of the first adder circuit via an accumulator. Here, each of the plurality of adder circuits includes an adder tree coupled to the plurality of memory cells on each of the respective ones of the plurality of columns. A method for in-memory computing, comprising the above steps.

14. Further comprising continuously activating the plurality of memory cells based on different activation inputs, wherein the accumulating of the output signal of the first adder circuit is performed after the plurality of memory cells are continuously activated. The method according to claim 13.

15. The continuously activating of the plurality of memory cells includes: Receiving a first set of the activation inputs during a first activation cycle; Receiving a second set of the activation inputs during a second activation cycle; And the accumulating of the output signal of the first adder circuit is performed after the first activation cycle and the second activation cycle. The method according to claim 14. The method according to claim 14. ​