Accumulator for digital memory-in-computation architectures

The digital CIM architecture addresses inaccuracies and inefficiencies in conventional CIM by using digital counters and accumulators for high-speed, energy-efficient in-memory computation, optimizing performance and reducing power consumption on devices like mobile devices.

KR102993699B1Active Publication Date: 2026-07-21QUALCOMM INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
QUALCOMM INC
Filing Date
2022-10-12
Publication Date
2026-07-21

Smart Images

  • Figure 112024037191966-PCT00013_ABST
    Figure 112024037191966-PCT00013_ABST
Patent Text Reader

Abstract

Certain embodiments provide apparatuses for performing machine learning tasks, specifically, for in-memory computation (CIM) architectures. One embodiment provides a method for in-memory computation. The method generally comprises the step of accumulating output signals on each of a number of columns of memory through each of a number of digital counters, wherein a number of memory cells are on each of the number of columns, and the number of memory cells store a number of bits representing the weights of a neural network, and the number of memory cells on each of the number of columns correspond to different word lines of memory; the step of adding the output signals of the number of digital counters through an adder circuit; and the step of accumulating the output signals of the adder circuit through an accumulator.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] Cross-reference regarding related applications

[0002] This application claims priority to U.S. Application No. 17 / 450,815 filed on October 13, 2021, which is assigned to the assignee of this application and is incorporated herein by reference in its entirety. Background Technology

[0003] Aspects of the present disclosure relate to performing machine learning tasks, specifically to computation-in-memory (CIM) architectures.

[0004] Machine learning is generally the process of generating trained models (e.g., artificial neural networks, trees, or other structures) that represent a generalized fit to a set of training data known a priori. Applying a trained model to new data generates inferences, which can be used to gain insights into the new data. In some cases, applying a model to new data is described as "executing inference" on the new data.

[0005] As the use of machine learning has surged to enable various machine learning (or artificial intelligence) tasks, the need for more efficient processing of machine learning model data has arisen. In some cases, dedicated hardware, such as machine learning accelerometers, can be used to enhance the capacity of processing systems for processing machine learning model data. However, such hardware requires space and power, and is not always available on processing devices. For example, "edge processing" devices, such as mobile devices, always-on devices, and Internet of Things (IoT) devices, typically must balance processing capabilities with power and packaging constraints. Furthermore, accelerometers can move data across common data buses, which can result in significant power consumption and introduce latency to other processes sharing the data bus. Consequently, other modes of processing systems are being considered to process machine learning model data.

[0006] Memory devices are an example of another form of processing system that can be leveraged to perform the processing of machine learning model data through so-called In-Memory Computation (CIM) processes. Conventional CIM processes perform computation using analog signals, which leads to inaccuracies in computation results and can adversely affect neural network computations. Therefore, systems and methods are needed to perform In-Memory Computation with increased accuracy.

[0007] Certain embodiments provide devices and techniques for performing machine learning tasks, specifically, for in-memory computation architectures.

[0008] One embodiment provides a circuit for in-memory computation. The circuit generally comprises: a memory having a plurality of columns; a plurality of memory cells on each column of the memory—the plurality of memory cells are configured to store a plurality of bits representing weights of a neural network, and the plurality of memory cells on each of the plurality of columns correspond to different word lines of the memory—; a plurality of digital counters—each of the plurality of digital counters is coupled to each of the columns of the memory—; an adder circuit coupled to the outputs of the plurality of digital counters; and an accumulator coupled to the output of the adder circuit.

[0009] One embodiment provides a method for in-memory computation. The method generally comprises the steps of: accumulating output signals on each of a plurality of columns of memory through each of a plurality of digital counters—where a plurality of memory cells are on each of a plurality of columns, and the plurality of memory cells store a plurality of bits representing weights of a neural network, and the plurality of memory cells on each of a plurality of columns correspond to different word lines of memory—; adding the output signals of a plurality of digital counters through an adder circuit; and accumulating the output signals of the adder circuit through an accumulator.

[0010] One embodiment provides an apparatus for in-memory computation. The apparatus generally comprises means for counting the amount of logic highs of output signals on each of a plurality of columns of memory—where a plurality of memory cells are on each of the plurality of columns, and the plurality of memory cells are configured to store a plurality of bits representing weights of a neural network, and the plurality of memory cells on each of the plurality of columns correspond to different word lines of memory—; means for adding output signals of a plurality of digital counters; and means for accumulating output signals of the means for adding.

[0011] Other embodiments provide a processing system configured to perform the methods described herein as well as those described herein; non-transient computer-readable media comprising instructions that, when executed by one or more processors of the processing system, cause the processing system to perform the methods described herein as well as those described herein; a computer program product implemented on a computer-readable storage medium comprising code for performing the methods described above as well as those further described herein; and a processing system comprising means for performing the methods described above as well as those further described herein.

[0012] The following description and related drawings describe in detail specific exemplary features of one or more embodiments. Brief explanation of the drawing

[0013] In a manner that allows the features of the above disclosure to be understood in detail, a more specific description briefly summarized above may be made with reference to embodiments, among which portionThis is illustrated in the attached drawings. However, it should be noted that the attached drawings merely illustrate specific typical embodiments of the present disclosure and should therefore not be construed as limiting the scope of the present disclosure, as the present description may allow for other equally valid embodiments. Fig. 1a inside Fig. 1d Is Examples of various types of neural networks that can be implemented by the embodiments of the present disclosure are described. Fig. 2 Is An example of a traditional convolution operation that can be implemented by the embodiments of the present disclosure is described. Fig. 3a and Fig. 3b Is Examples of depth-separable convolution operations that can be implemented by the embodiments of the present disclosure are described. Fig. 4 It exemplifies an exemplary memory cell implemented as an 8-transistor (8T) static random access memory (SRAM) cell for in-memory-computation (CIM) circuits. Fig. 5 This exemplifies a circuit for CIM according to specific embodiments of the present disclosure. Fig. 6 This exemplifies a digital CIM (DCIM) circuit having an accumulator implemented using a pulse generator and a digital counter, according to specific embodiments of the present disclosure. Fig. 7 This is a block diagram illustrating a CIM circuit implemented using a delay circuit for reducing electromagnetic interference (EMI) according to specific embodiments of the present disclosure. Fig. 8a is a block diagram illustrating an exemplary implementation of a frequency converter according to specific aspects of the present disclosure. Fig. 8b is a graph showing the input and output signals of an edge-to-pulse converter according to specific embodiments of the present disclosure. Fig. 9aand Fig. 9b This illustrates an integrated circuit (IC) layout for implementing a DCIM circuit according to specific embodiments of the present disclosure. Fig. 10 This exemplifies a DCIM circuit implemented using a multiplexer to facilitate the reuse of an adder tree and an accumulator according to specific embodiments of the present disclosure. Fig. 11 is a flowchart illustrating exemplary operations for in-memory computation according to specific aspects of the present disclosure. Fig. 12 exemplifies an exemplary electronic device configured to perform operations for signal processing in a neural network according to specific embodiments of the present disclosure. For ease of understanding, the same reference numbers have been used where possible to designate identical elements common to the drawings. It is considered that elements and features of one embodiment may be beneficially incorporated into other embodiments without further description. Specific details for implementing the invention

[0014] Aspects of the present disclosure provide apparatuses, methods, processing systems, and computer-readable media for performing in-memory computation (CIM) to handle data-intensive processing, such as implementing machine learning models. Some aspects provide techniques for performing digital CIM (DCIM) using digital counters, each digital counter accumulating output signals in its respective column of a number of columns of memory. As used herein, "accumulator" generally refers to a circuit used to accumulate output signals over a number of cycles. "Adder circuit" or "adder tree" generally refers to digital adders used to add output signals of a number of memory cells (e.g., memory cells over word lines or columns).

[0015] The embodiments described herein provide a high-speed and energy-efficient accumulator for digital CIM applications. Word lines of a CIM circuit can be activated sequentially, and after two or more word lines are activated sequentially, a digital counter can be used to perform accumulation and provide the accumulation result. For example, a digital counter can be used to count the amount of logic highs generated on a column of memory after a number of computation cycles.

[0016] The DCIM circuit provided herein has energy consumption per bit that is not scaled by the number of active rows of the memory array, and thus can reduce the total energy consumption of the DCIM circuit compared to conventional implementations. The embodiments described herein can also reduce area consumption for the entire CIM system compared to conventional implementations by reusing the adder tree and accumulator for multiple weight column groups. Furthermore, the DCIM circuit provided herein has a self-timed operation that enables high-speed operation by timing using signals on their respective columns, in contrast to digital counters using clock signals. Partial sums from digital counters can be accumulated on a global accumulator, which operates at a slower clock, resulting in higher energy efficiency of the CIM system. In some embodiments, as described in more detail herein, electromagnetic interference (EMI) is reduced by the self-timed operation within the DCIM circuit and the phase shift of the local clock.

[0017] CIM-based machine learning (ML) and artificial intelligence (AI) can be used for a wide variety of tasks, including image and audio processing, and making wireless communication decisions (e.g., to optimize or at least increase throughput and signal quality). Additionally, CIM supports various types of memory architectures, such as dynamic random-access memory (DRAM) and static random-access memory (SRAM) (e.g., Fig. 4 It can be based on SRAM cells (such as those in), magnetoresistive random-access memory (MRAM), and resistive random-access memory (ReRAM or RRAM), and can be attached to various types of processing units, including central processing units (CPUs), digital signal processors (DSPs), graphics processing units (GPUs), field-programmable gate arrays (FPGAs), AI accelerators, etc. Generally, CIM can advantageously reduce the "memory wall" problem, which occurs when the movement of data into and out of memory consumes more power than the computation of the data. Therefore, significant power savings can be realized by performing in-memory computation. This is particularly useful for various types of electronic devices, such as lower-power edge processing devices and mobile devices.

[0018] For example, a mobile device may include a memory device configured to store data and perform in-memory computation operations (also referred to as "in-memory compute" operations). The mobile device may be configured to perform ML / AI operations based on data generated by the mobile device, such as image data generated by the mobile device's camera sensor. Accordingly, the memory controller unit (MCU) of the mobile device may load weights from other on-board memory (e.g., flash or RAM) into the memory device's CIM array and allocate input feature buffers and output (e.g., output enable) buffers. Subsequently, the processing device may begin processing the image data by, for example, loading a certain layer into the input buffer and processing that layer with the weights loaded into the CIM array. This processing may be repeated for each channel of the image data, and outputs (e.g., output enable) are stored in output buffers and subsequently used by the mobile device for ML / AI tasks such as facial recognition.

[0019] A brief background on neural networks, deep neural networks, and deep learning

[0020] Neural networks are organized into layers of interconnected nodes. Generally, nodes (or neurons) are where computation takes place. For example, a node can combine input data with a set of weights (or coefficients) that amplify or attenuate the input data. Thus, the amplification or dampening of input signals can be regarded as assigning relative importance to various inputs related to the task the network is attempting to learn. Generally, input-weight products are summed (or accumulated), and the sum then passes through the node's activation function to determine whether and to what extent the corresponding signal should proceed further through the network.

[0021] In the most basic implementation, a neural network can have an input layer, a hidden layer, and an output layer. "Deep" neural networks generally have more than one hidden layer.

[0022] Deep learning is a method of training deep neural networks. Generally, deep learning maps inputs to the network to outputs from the network, and therefore deep learning [can handle] any input x and output y unknown function between It is sometimes referred to as a "universal approximator" because it can learn to approximate it. In other words, deep learning is x cast y To convert to the correct Finds.

[0023] More specifically, deep learning trains each layer of nodes based on a unique set of features that are the outputs from the previous layer. Consequently, the features become more complex with each successive layer of the deep neural network. Therefore, deep learning is powerful because it can perform complex tasks such as object recognition by progressively extracting higher-level features from input data, learning to represent inputs at successive levels of abstraction at each layer, and thereby constructing useful feature representations of the input data.

[0024] For example, when visual data is provided, the first layer of a deep neural network can be trained to recognize relatively simple features, such as edges, in the input data. In another example, when auditory data is provided, the first layer of the deep neural network can be trained to recognize the spectral power of specific frequencies in the input data. Subsequently, the second layer of the deep neural network can be trained to recognize combinations of features, such as simple shapes in the visual data, or combinations of sounds in the auditory data, based on the output of the first layer. Later, the upper layers can be trained to recognize complex shapes in the visual data or words in the auditory data. Even higher layers can be trained to recognize general visual objects or spoken phrases. Thus, deep learning architectures can perform particularly well when applied to problems that have a natural hierarchical structure.

[0025] Layered connections in neural networks

[0026] Neural networks, such as deep neural networks (DNNs), can be designed with various connection patterns between layers.

[0027] Fig. 1aThis illustrates an example of a fully connected neural network (102). In a fully connected neural network (102), each node of the first layer can communicate its output to all nodes of the second layer, so that each node of the second layer will receive input from all nodes of the first layer.

[0028] Fig. 1b This illustrates an example of a locally connected neural network (104). In the locally connected neural network (104), a node in the first layer may be connected to a limited number of nodes in the second layer. More generally, the locally connected layer of the locally connected neural network (104) may be configured to have connection strengths (or weights) such that each node in the layer has the same or similar connection pattern but different values ​​(e.g., values ​​associated with local regions (110, 112, 114, 116) of the first layer nodes). The locally connected connection pattern may generate spatially distinct reception fields in the upper layer, because the upper layer nodes within a given region may receive inputs that are tuned through training on attributes of a limited portion of the total input to the network.

[0029] One type of locally connected neural network is the convolutional neural network (CNN). Fig. 1c This illustrates an example of a convolutional neural network (106). The convolutional neural network (106) can be configured so that connection strengths associated with inputs to each node of the second layer are shared (e.g., for a local region (108) that overlaps with other local regions of the first layer nodes). Convolutional neural networks can be very suitable for problems where the spatial locations of the inputs are meaningful.

[0030] One type of convolutional neural network is a deep convolutional network (DCN). Deep convolutional networks are networks of multiple convolutional layers, which can be further composed, for example, of pooling and normalization layers.

[0031] Fig. 1d This illustrates an example of a DCN (100) designed to recognize visual features within an image (126) generated by an image capture device (130). For example, if the image capture device (130) is a camera mounted in or on a vehicle (or otherwise moving with it), the DCN (100) can be trained with various supervised learning techniques to identify traffic signs and even numbers on traffic signs. Similarly, the DCN (100) can be trained for other tasks, such as identifying lane markings or traffic lights. These are just some exemplary tasks, and many other tasks are possible.

[0032] Fig. 1d In the example, the DCN (100) includes a feature extraction section and a classification section. When an image (126) is received, the convolutional layer (132) is a convolutional kernel (e.g., Fig. 2 (as described and explained in) is applied to the image (126) to generate a first set of feature maps (118) (or intermediate activations). Generally, the “kernel” or “filter” comprises a multidimensional array of weights designed to highlight different aspects of the input data channel. In various examples, the terms “kernel” and “filter” may be used interchangeably to refer to sets of weights applied in a convolutional neural network.

[0033] The first set of feature maps (118) may then be subsampled by a pooling layer (e.g., a max pooling layer, not shown) to generate a second set of feature maps (120). The pooling layer may reduce the size of the first set of feature maps (118) while retaining a large amount of information to improve model performance. For example, the second set of feature maps (120) may be downsampled by the pooling layer from a 28 x 28 matrix to a 14 x 14 matrix.

[0034] This process can be repeated through many layers. In other words, the second set of feature maps (120) can be further convolved through one or more subsequent convolutional layers (not shown) to generate one or more subsequent sets of feature maps (not shown).

[0035] Fig. 1d In the example, a second set of feature maps (120) is provided to a fully connected layer (124), which then generates an output feature vector (128). Each feature of the output feature vector (128) may include a number corresponding to a possible feature of the image (126), such as "signboard", "60", and "100". In some cases, a softmax function (not shown) may convert the numbers in the output feature vector (128) into probabilities. In such cases, the output (122) of the DCN (100) is the probability that the image (126) contains one or more features.

[0036] A softmax function (not shown) can convert individual elements of an output feature vector (128) into probabilities such that the output (122) of the DCN (100) becomes one or more probabilities of an image (126) containing one or more features, such as a sign having the number "60" on it, as in the image (126). Thus, in this example, the probabilities in the output (122) for "sign" and "60" must be higher than the probabilities of other elements of the output (122), such as "30", "40", "50", "70", "80", "90", and "100".

[0037] Before training the DCN (100), the output (122) generated by the DCN (100) may be inaccurate. Therefore, the output (122) and a priori The error between known target outputs can be calculated. For example, the target output here is an indication that the image (126) contains a "sign" and the number "60". Then, using the known target output, the weights of the DCN (100) can be adjusted through training so that the subsequent output (122) of the DCN (100) achieves the target output (having high probabilities).

[0038] To adjust the weights of the DCN (100), the learning algorithm can compute a gradient vector for the weights. The gradient vector may indicate the amount by which the error increases or decreases when the weights are adjusted in a particular way. Subsequently, the weights can be adjusted to reduce the error. This method of adjusting the weights may be referred to as "backpropagation" because this adjustment process involves "backward passing" through the layers of the DCN (100).

[0039] In fact, since the error gradient of the weights can be calculated over a small number of examples, the calculated gradient approaches the actual error gradient. This approach may be referred to as "stochastic gradient descent." Stochastic gradient descent can be repeated until the achievable error rate of the entire system stops decreasing or until the error rate reaches a target level.

[0040] After training, new images may be presented to the DCN (100), and the DCN (100) may generate inferences such as classifications or probabilities of various features in the new images.

[0041] Convolution techniques for convolutional neural networks

[0042] Convolution is generally used to extract useful features from input data sets. For example, in a convolutional neural network as described above, convolution enables the extraction of different features using kernels and / or filters whose weights are automatically learned during training. Subsequently, inferences are performed by combining the extracted features.

[0043] Activation functions can be applied before and / or after each layer of a convolutional neural network. Activation functions are generally mathematical functions that determine the output of a node in a neural network. Therefore, an activation function determines whether a node should pass information through based on whether the node's input is relevant to the model's prediction. In one example, (in other words, y Is x In the case where it is a convolution of, x and y Both can generally be regarded as "activations." However, from the perspective of a specific convolution operation, xIt may also be referred to as "pre-activations" or "input activations," which is x This is because it exists before a specific convolution, y It can be referred to as "output activations" or "feature maps".

[0044] Fig. 2 This describes an example of traditional convolution in which a 12-pixel x 12-pixel x 3-channel input image is convolved using a 5-x5-x3 convolution kernel (204) and a stride (or step size) of 1. The resulting feature map (206) is 8 pixels x 8 pixels x 1 channel. As can be seen from this example, traditional convolution can change the number of dimensions of the input data relative to the output data (here, 12 x 12 to 8 x 8 pixels) by including channel dimensions (here, 3 channels to 1 channel).

[0045] One way to reduce the computational burden (measured, for example, in floating-point operations per second (FLOPs)) and the number of parameters associated with neural networks, including convolutional layers, is to factorize the convolutional layers. For example, Fig. 2 Spatially separable convolutions, as described in [figure], can be factored into two components: (1) depth-wise convolution, where each spatial channel is independently convolved by depth-wise convolution (e.g., spatial fusion); and (2) point-wise convolution, where all spatial channels are linearly combined (e.g., channel fusion). An example of a depth-wise separable convolution is Fig. 3a and Fig. 3b It is described in. Generally, during spatial fusion, the network learns features from spatial planes, and during channel fusion, the network learns the relationships between these features across channels.

[0046] In one example, depth-separable convolution can be implemented using 5x5 kernels for spatial fusion and 1x1 kernels for channel fusion. In particular, channel fusion is depth d 1x1x repeating through every single point in the input image d The kernel can be used, and here the depth of the kernel d It generally matches the number of channels in the input image. Channel fusion through point-by-point convolution is useful for dimensionality reduction for efficient computations. 1x1x d Applying kernels and adding an activation layer after the kernels can provide added depth to the network, which can increase network performance.

[0047] especially, Fig. 3a In this case, a 12 pixel x 12 pixel x 3 channel input image (302) is convolved with a filter comprising three distinct kernels (304A to 304C), each having a 5 x 5 x 1 dimension, to generate an 8 pixel x 8 pixel x 3 channel feature map (306), wherein each channel is generated by an individual kernel among the kernels (304A to 304C).

[0048] Next, the feature map (306) is further convolved using a point-by-point convolution operation with a kernel having dimensions 1 x 1 x 3 to generate an 8 pixel x 8 pixel x 1 channel feature map (310). As depicted in this example, the feature map (310) has reduced dimensions (1 channel versus 3 channels), which makes computations with it more efficient.

[0049] Fig. 3a and Fig. 3b The result of depth-separable convolution in Fig. 2Although substantially similar to traditional convolution, the number of computations is significantly reduced, and thus depth-separable convolution provides significant efficiency gains if the network design allows it.

[0050] Fig. 3b Although not depicted in, a number of (e.g., m Point-by-point convolution kernels (308) (e.g., individual components of a filter) can be used to increase the channel dimension of the convolution output. Thus, for example, m = 256 1 x 1 x 3 kernels (308) can be generated, each output being an 8 pixel x 8 pixel x 1 channel feature map (e.g., feature map (310)), and these feature maps can be stacked to obtain a resulting feature map of 8 pixels x 8 pixels x 256 channels. The resulting increase in the number of channels provides more parameters for training, which can improve the ability of the convolutional neural network to identify features (e.g., in the input image (302)).

[0051] Exemplary In-Memory Computation (CIM) Architecture

[0052] Digital In-Memory Computation (CIM) is used to resolve energy and speed bottlenecks caused by moving data between the processing unit and memory for logical operations. For example, digital CIM can be used to perform logic operations in memory, such as bit-parallel / bit-serial logic operations (e.g., AND operations). However, machine learning workloads still involve final accumulation that can be performed using circuitry outside the memory array. Since each row in memory is read sequentially, high-speed and energy-efficient accumulators can be implemented near memory to achieve better performance (e.g., to increase tera-operations per second (TOPS).

[0053] Fig. 4 This illustrates an exemplary memory cell (400) of static random access memory (SRAM) that can be implemented in a CIM array. As the memory cell (400) is implemented with eight transistors, the memory cell (400) may be referred to as an "8-transistor (8T) SRAM cell."

[0054] As illustrated, the memory cell (400) may include a cross-coupled inverter pair (424) having an output (414) and an output (416). As illustrated, the cross-coupled inverter pair output (414) is optionally coupled to a write bit-line (WBL) (406) via a pass-gate transistor (402), and the cross-coupled inverter pair output (416) is optionally coupled to a complementary write bit-line (WBLB) (420) via a pass-gate transistor (418). The WBL (406) and WBLB (420) are configured to provide complementary digital signals to be written (e.g., stored) to the cross-coupled inverter pair (424). Bits for neural network weights can be stored in the memory cell (400) using the WBL and WBLB. As illustrated, the gates of the pass-gate transistors (402, 418) can be coupled to a write word-line (WWL) (404). For example, a digital signal to be written can be provided to the WBL (and a complement of the digital signal is provided to the WBLB). Then, the pass-gate transistors (402, 418) (which are implemented here as n-type field-effect transistors (NFETs)) are turned on by providing a logic high signal to the WWL (404), and as a result, the digital signal is stored in a cross-coupled inverter pair (424).

[0055] As illustrated, the cross-coupled inverter pair output (414) can be coupled to the gate of transistor (410). The source of transistor (410) can be coupled to a reference potential node (VSS or electrical ground), and the drain of transistor (410) can be coupled to the source of transistor (412). As illustrated, the drain of transistor (412) can be coupled to a read bit-line (RBL) (422). The gate of transistor (412) can be controlled via a read word-line (RWL) (408). The RWL (408) can be controlled via an enable input signal.

[0056] During a read cycle, the RBL (422) may be pre-charged to logic high. If both the enable input on the RWL (408) and the weight bit stored in the cross-coupled inverter pair output (414) are logic high, both transistors (410, 412) are turned on, so that the RBL (422) is electrically coupled to VSS at the source of transistor (410) and the RBL (422) is discharged to logic low. If either the enable input on the RWL (408) or the weight bit stored in the cross-coupled inverter pair output (414) is logic low, at least one of the transistors (410, 412) will be turned off, and the RBL (422) will remain logic high. Therefore, the output of the memory cell (400) in the RBL (422) is logic low only when both the weight bit and the activation input are logic high, and logic high otherwise, so it effectively implements NAND gate operation.

[0057] Fig. 5 ... illustrates a circuit (500) for CIM according to specific embodiments of the present disclosure. The circuit (500) has word lines (5040 to 504 31It includes a CIM array (501) having )(also referred to as "rows") and columns (5060 to 5067). Word lines (5040 to 504 31 ) are collectively referred to as "word lines (WL) (504)", and columns (5060 to 5067) are collectively referred to as "columns (506)". As illustrated, the CIM array (501) may include an activation circuit (590) configured to provide activation signals to the word lines (504). For ease of understanding, the CIM array (501) is implemented with 32 word lines and 8 columns, but the CIM array may be implemented with any number of word lines or columns. As illustrated, memory cells (502 0-0 to 502 31-7 )(collectively referred to as “memory cells (502”)) are implemented at the intersections of WLs (504) and columns (506).

[0058] Each of the memory cells (502) is Fig. 4 It can be implemented using the memory cell architecture described in relation to. As illustrated, activation inputs a (0,0) to a (31,0) can be provided to their respective word lines (504), and memory cells (502) can store neural network weights w (0,0) to w (31,7). For example, memory cells (502 0-0 to 502 0-7 ) can store weight bits w(0,0) to w(0,7), and memory cells (502 1-0 to 502 1-7 Each word line can store weight bits w(1,0) to w(1,7), etc. Each word line can store multi-bit weights. For example, weight bits w(0,0) to w(0,7) can represent 8 bits of weights of a neural network.

[0059] As illustrated, each of the columns (506) is coupled to a sense amplifier (SA) (5030 to 5037). The sense amplifiers (5030, 5031 to 5037) are collectively referred to as "sense amplifiers (503)". As illustrated, the input of each sense amplifier (503) can be coupled to the outputs of the memory cells on their respective columns.

[0060] The outputs of the sensing amplifiers (503) are coupled to a column accumulator circuit (553). For example, each output of the sensing amplifiers (503) is coupled to one of the accumulators (5070, 5071 to 5077) (collectively referred to as "accumulators (507)") of the column accumulator circuit (553). Each accumulator (507) performs the accumulation of the output signals of the respective sensing amplifiers of the sensing amplifiers (503) over a number of computation cycles. For example, during each computation cycle, computation is performed on a single word line, and the output signal of the computation on the word line is accumulated with the output signals during other computation cycles using the respective accumulators of the accumulators (507). After a number of computation cycles (e.g., 32 cycles for 32 word lines), each of the accumulators (507) provides an accumulated result.

[0061] During the operation of the circuit (500), the activation circuit section (590) provides the activation inputs a (0,0) to a (31,0) of the first set (599) to the memory cells (502) for computation during the first activation cycle. The activation inputs a (0,0) to a (31,0) are provided for one row at a time, and the outputs of each computation for each row are accumulated using their respective accumulators of the accumulators (507) as described. The same operation is performed until the activation inputs representing the least significant bits (LSB) of the activation parameters are processed for other sets of activation inputs, such as the activation inputs a (0,1) to a (31,1) representing the second most significant bits (MSB) of the activation parameters during subsequent activation cycles.

[0062] Once multiple computation cycles have been completed during each activation cycle, the outputs of the accumulators (507) are provided to the weighted shift adder tree circuit (512) for addition across columns, and the output of the weighted shift adder tree circuit (512) is provided to the activation shift accumulator circuit (516) for accumulation across activation cycles. In other words, the activation shift accumulator circuit (516) accumulates the computation results after the activation cycles are completed.

[0063] The weighted shift adder tree circuit (512) includes a plurality of weighted shift adders (e.g., weighted shift adder (514)), each comprising a bit shift and adder circuit, to facilitate the performance of bit shift and adder operations. In other words, memory cells on column (5060) can store the most significant bits (MSBs) for their respective weights, and memory cells on column (5067) can store the least significant bits (LSBs) for their respective weights. Thus, when performing addition across columns (506), a bit shift operation is performed to shift the bits in order to take into account the importance of the bits on the associated columns. In other words, once bitwise accumulation occurs in each of the accumulators (507) over multiple computation cycles, the weighted shift adder tree circuit (512) combines the weighted sums of the eight columns (e.g., providing the accumulation result for a given activation bit position during each activation cycle), and the activation shift accumulator circuit (516) combines the results from multiple (e.g., eight) activation cycles to output the final accumulation result.

[0064] As described, the output of the weighted shift adder tree circuit (512) is provided to the activation shift accumulator circuit (516). The activation shift accumulator circuit (516) includes a bit shift circuit (518) and an accumulator (520). The bit shift circuit (518) performs a bit shift operation based on activation cycles. For example, for an 8-bit activation parameter processed using 8 activation cycles, the bit shift circuit may perform an 8-bit shift during the first activation cycle, a 7-bit shift during the second activation cycle, and so on. After activation cycles, the outputs of the bit shift circuit (518) are accumulated using the accumulator (520) to generate a DCIM output signal.

[0065] In some embodiments, the CIM array (501), the activation circuit (590), and the column accumulator circuit (553) operate at a higher frequency (e.g., 8 times or more) than the weighted shift adder tree circuit (512) and the activation shift accumulator circuit (516). As illustrated, half-latch circuits (5090, 5091 to 5097) (collectively referred to as "half-latch circuits (509)") may be coupled to the respective outputs of the accumulators (507). Each half-latch circuit holds the output of the respective accumulator of the accumulators (507) and, once a number of computation cycles have been completed, provides the output to the respective input of the weighted shift adder tree circuit (512). In other words, a half-latch circuit generally refers to a latch circuit that holds a digital input (e.g., the output of one of the accumulators (507)) at the start of a clock cycle and provides the digital input to the output of the latch circuit at the end of a clock cycle. The half-latch circuits (509) facilitate the transition from the higher frequency operation of the column accumulator circuit (553) to the lower frequency operation of the weighted shift adder tree circuit (512). The half-latch circuits (509) can synchronize between clock domains without using other components (e.g., extra buffers).

[0066] Certain modes are, Fig. 6 As described in more detail in relation thereto, a digital counter is provided that enables adder-free partial sum generation (e.g., a one's counter configured to count the amount of logic highs in the time domain). In other words, each of the accumulators (507) can be implemented using a digital counter that counts the amount of logic highs provided to the output of each of the sense amplifiers (503).

[0067] Fig. 6The DCIM circuit section having an accumulator (e.g., one of the accumulators (507)) implemented using a pulse generator (602) and a digital counter (604) according to specific embodiments of the present disclosure. As illustrated, Fig. 6 A memory cell (502) illustrated as a NAND gate 0-0 A memory cell such as ) can provide a computation result based on an activation input and a stored weight bit. Each of the activation input and the stored weight bit can have a 50% toggling probability. In other words, the activation input can have a 50% probability of being logic high and a 50% probability of being logic low. Similarly, the stored weight bit can have a 50% probability of being logic high and a 50% probability of being logic low. As a result, the output of the NAND gate can have a 25% toggling probability due to the NAND operation (e.g., a 25% probability of being logic high and a 75% probability of being logic low).

[0068] The pulse generator (602) generates a pulse for each logic high output of the associated sensing amplifier. For example, during the first computation cycle, the memory cell (502 0-0 If the output of ) is logic high, the pulse generator (602) generates a pulse, and during the second computation cycle, the memory cell (502 0-1 When the output of ) is logic high, the pulse generator (602) generates a pulse, etc. The output of the pulse generator (602) is provided to a digital counter (604). The digital counter counts the number of pulses generated by the pulse generator (602) and generates a digital counter output signal (e.g., a 6-bit digital signal including bits q(0) to q(5)).

[0069] As illustrated, the digital counter (604) comprises flip-flops (6060 to 6065) (collectively referred to as "flip-flops (606)"), wherein the output of the pulse generator (602) is provided to the clock (CLK) input of the flip-flop (606). The complementary output of each flip-flop ( ) is fed back to the data (D) input of that flip-flop, and the output (Q) of each flip-flop is provided to the CLK of the subsequent flip-flop in the flip-flop chain.

[0070] As illustrated, flip-flop (6060) has the highest energy consumption among the flip-flops (606), and flip-flop (6065) has the lowest energy consumption among the flip-flops (606). For example, if the pulse generator (602) consumes 0.5 femtojoules (fJ) per step (e.g., per computation cycle), flip-flop (6060) consumes 0.8 fJ / step, flip-flop (6061) consumes 0.4 fJ / step, flip-flop (6062) consumes 0.2 fJ / step, flip-flop (6063) consumes 0.1 fJ / step, flip-flop (6064) consumes 0.05 fJ / step, and flip-flop (6065) consumes 0.025 fJ / step. In other words, the flip-flop (6060) generates the least significant bit (LSB) q (0) of the digital counter output signal, and the flip-flop (6065) generates the most significant bit (MSB) q (5) of the digital counter output signal. Since the output of the flip-flop (6061) has half the toggling probability compared to the output of the flip-flop (6060), the energy consumption of the flip-flop (6060) is twice the energy consumption of the flip-flop (6061). Similarly, since the output of the flip-flop (6062) has half the toggling probability compared to the output of the flip-flop (6061), the energy consumption of the flip-flop (6061) is twice the energy consumption of the flip-flop (6062), and so on.

[0071] Each stage of the digital counter (604) is effectively a divide-by-two frequency divider, wherein the toggling of one flip-flop stage is controlled by the output signal of the preceding flip-flop stage. Implementing additional stages for the digital counter (e.g., increasing the number of bits of the digital counter output signal) has little effect on the energy consumption of the DCIM circuitry, because energy consumption increases asymptotically with the additional stages. A half-latch circuitry may be coupled to the output of the digital counter for each column to synchronize the output of the counter to the slow clock domain (DCIM clock). In other words, each of the bits q (0) through q (5) may be provided to a half-latch circuitry (e.g., one of the half-latch circuitrys of the half-latch circuitry (509)). In some embodiments, Fig. 7 As explained in more detail in [link], a delay circuit can be used to reduce interference between rows of memory.

[0072] Fig. 7This is a block diagram illustrating a CIM circuit implemented using a delay circuit for reducing electromagnetic interference (EMI) according to specific embodiments of the present disclosure. For example, as illustrated, 8-bit weights may be stored in memory cells (502) and processed using sense amplifiers (503), column accumulators (507), a weight shift adder tree circuit (512), and an activation shift accumulator circuit (516), as described herein. Delay circuits (750) (e.g., associated with different delays) may be implemented between the sense amplifiers (503) and the column accumulators (507) to implement an offset with respect to the phase of the signals provided to the column accumulators (507). For example, a single delay element (labeled "1D") may be coupled between the sense amplifier (5030) and the column accumulator (5070), and two delay elements (labeled "2D") may be coupled between the sense amplifier (5031) and the column accumulator (5071), etc. In this way, the falling / rising edges of the signals provided to the column accumulators (507) are offset to reduce EMI between the columns. In other words, skew (e.g., through one or more delay cells) is added to the input of the digital counter (604) to reduce EMI that would otherwise be due to simultaneous switching noise.

[0073] As described, the clock generator circuit (770) may be used to generate the DCIM clock and the local clock. For example, the clock generator circuit (770) may include a clock generator (771) configured to generate the DCIM clock. The clock generator (771) may be implemented using any suitable clock generation circuit, such as a phase-locked loop (PLL) or a ring oscillator (RO). The weighted shift adder tree circuit (512) may receive the DCIM clock and operate on it. For certain embodiments, the clock generator circuit (770) may include a frequency converter (702) that may be used to generate the local clock from the DCIM clock based on the operation of the activation circuit (590). Although the frequency converter (702) is depicted as being part of the clock generator circuit (770), in some implementations, the frequency converter (702) may be separate from the clock generator circuit (770). The frequency converter generally refers to any circuit that receives a clock signal having a first frequency and generates a second clock signal having a second different frequency.

[0074] The frequency converter can be implemented using any suitable technique. For example, the frequency converter can be implemented as a ring oscillator (RO) modulated by system clock timing (e.g., modulated using the DCIM clock), or can be implemented by generating pulses from the rising or falling edges of the system clock, which Fig. 8a and Fig. 8b As explained in more detail regarding this. In this way, the local clock can have a rising edge synchronized with the DCIM clock.

[0075] Fig. 8a is a block diagram illustrating an exemplary implementation of a frequency converter (702) according to specific embodiments of the present disclosure. Fig. 8bis a graph representing the input signal (840) and output signal (842) of an edge-pulse converter. The frequency converter (702) is one or more edge-pulse converters (8021, 8022 to 802 n It may include )(collectively referred to as "edge-pulse converters (802)"). Each of the edge-pulse converters (802) generates a pulse at each rising edge and each falling edge of the input signal presented to the edge-pulse converter. For example, Fig. 8b As illustrated in the figure, the edge-pulse converter (8021) generates a pulse (822) after detecting the rising edge (820) of the input signal (840), and generates another pulse (826) after detecting the falling edge (824) of the input signal (840). In this way, the output signal (842) of the edge-pulse converter is twice the frequency of the input signal (840) of the edge-pulse converter. Fig. 7 As explained in relation to, using multiple edge-pulse converters in series enables the frequency of the DCIM clock to be upconverted to generate a local clock. Although edge-pulse converters are provided as one example of a frequency converter, any suitable type of frequency converter may be used.

[0076] Fig. 9a and Fig. 9b This illustrates an integrated circuit (IC) layout (900) for implementing a DCIM circuit according to specific embodiments of the present disclosure. Fig. 9aAs illustrated, each SRAM column can be implemented using pins (e.g., to implement fin field-effect transistors (FinFETs)) and may have fin pitches of 10 to 14 nm. A fin pitch refers to the distance from one pin to an adjacent pin. As described herein, the memory cells of the SRAM column are coupled to a sense amplifier. As illustrated, the output of the sense amplifier is coupled to a digital counter implemented using flip-flops cascaded along each column having pins. The counter design may have fin pitches of 10 to 12 nm.

[0077] Fig. 9b As illustrated in the figure, each column is an 8-transistor (8T) SRAM column (e.g., a column of memory cells, each cell implemented using 8 transistors), a single-ended sense amplifier (e.g., sense amplifier (5030)), a pulse generator (e.g., pulse generator (602)), Fig. 6 As illustrated in [Figure], it may include one's counter circuitry for generating each of bits q(0) through q(5) (e.g., each flip-flop of the flip-flops (606)), and latch circuits for each of bits q(0) through q(5). As illustrated, a binary adder tree (e.g., weighted shift adder tree circuit (512)) and an accumulator (e.g., activated shift accumulator circuit (516)) are also coupled to the outputs of the latch circuits. In some embodiments, Fig. 10 As described in more detail in relation to this, a multiplexer can be used to facilitate the reuse of the adder tree circuit (512) and the accumulator circuit (516).

[0078] Fig. 10The present disclosure illustrates a DCIM circuit (1000) implemented using a multiplexer (1004) to facilitate the reuse of an adder tree and an accumulator according to specific embodiments of the present disclosure. As illustrated, a CIM array (501) (e.g., including 8T SRAM cells) may include a plurality of weight column groups, each weight column group having a set of columns associated with a weight parameter having a plurality of bits. For example, a weight column group (1002) may refer to columns (506) storing 8-bit weights on each row. As illustrated, each weight column group may be coupled to 8 accumulators (e.g., digital counters) and 48 half-latch circuits (e.g., 8 accumulators x 6 bits per accumulator), as described herein.

[0079] In some embodiments, the partial sum operation is implemented using a shared accumulator (e.g., accumulator circuit (516)) across weight column groups, so that a single accumulator result can be provided at the end of each multiply-and-accumulate (MAC) cycle. A binary adder tree (e.g., weight shift adder tree circuit (512)) and a 21-bit accumulator (e.g., accumulator circuit (516)) consume a significant portion of the total area of ​​the DCIM circuitry. Therefore, sharing the binary adder tree and the activation shift accumulator across multiple weight column groups reduces the total area consumption of the DCIM circuitry.

[0080] As described, during each of the 8 activation cycles (10200 to 10207) (collectively referred to as "activation cycles (1020)"), 32 computation cycles occur (C00 to C031), and one computation cycle for each of the 32 rows is Fig. 5As illustrated in [Figure]. After activation cycles (1020), computation outputs for weight column groups are available at the outputs of the half-latch circuits.

[0081] The multiplexer (1004) can be used to individually select each group of weight columns to be processed using an adder tree circuit (512) and an accumulator circuit (516). For example, during a first weight cycle, a half-latch output (e.g., a 6-bit output) on a first group of weight columns (e.g., weight column group (1002)) is selected by the multiplexer (1004), and Fig. 5 As described above, it can be provided to an adder tree circuit (512) and an accumulator circuit (516) for processing.

[0082] During the second weighting cycle, a half-latch output (e.g., a 6-bit output) on the second weighting column group is selected by a multiplexer and may be provided to an adder tree circuit (512) and an accumulator circuit (516) for processing, etc. Time multiplexing of the latched partial sums enables the reuse of the adder tree circuit (512) and the accumulator circuit (516). To select a new weighting column group, (e.g., Fig. 7 A multiplexer select signal may be generated every two clock cycles of the DCIM clock shown in the figure. In other words, after every two clock cycles, a 21-bit accumulation result for each column is provided by the accumulator circuit (516), one clock cycle is for performing addition through the adder tree circuit (512), and the other clock cycle is for accumulation through the accumulator circuit (516). In some embodiments, a divided 2 version of the DCIM clock may be used to operate the multiplexer (1004), the adder tree circuit (512), and the accumulator circuit (516) to increase the throughput of addition and accumulation operations for weight column groups.

[0083] Aspects of the present disclosure provide innovative circuit and physical designs that enable high-speed and energy-efficient accumulation for any DCIM product. The described aspects enable high-speed computation by being virtually entirely self-timing without involving routing clock signals to generate partial sums. In other words, the digital counters used to implement the accumulators (507) are timed using the outputs of associated sense amplifiers (503) instead of clock signals. Data-divided clock (e.g., Fig. 7 Due to the local clock (as described in [the embodiment]), the embodiments of the present disclosure provide a low-energy implementation of the DCIM circuitry as the energy consumption of the digital counter increases asymptotically, and the accumulation for each weight column group is performed at a rate associated with the divided clock without using fast toggling partial sum components. The embodiments provided herein also allow the accumulation (e.g., accumulation circuit (516)) and the binary adder tree (e.g., adder tree circuit (512)) to be shared across weight column groups, thereby enabling area reduction at the DCIM system level. The DCIM circuitry described herein also provides a less complex design (e.g., because it does not include any full-adder cells or generate cells used in conventional implementations), allowing for a compact implementation. The embodiments described herein can also be implemented without modifications to the SRAM array used in CIM applications. Fig. 9a and Fig. 9b As described in relation thereto, the circuits are column-matched to the SRAM array to provide full array efficiency. As described herein, self-timing clocking and skew (e.g., phase offset) between the processing of different columns reduce EMI caused by simultaneous switching noise.

[0084] Exemplary operations for Digital Memory-In-Computation (DCIM)

[0085] Fig. 11 is a flowchart illustrating exemplary operations (1100) for in-memory computation according to specific embodiments of the present disclosure. The operations (1100) Figs. 5, 6, 7, 8a, 8b, 9a, 9b and Fig. 10 This can be performed by a circuit for CIM, such as the circuit (500) described in relation to.

[0086] In block (1101), the circuit can receive activation inputs from multiple memory cells (e.g., memory cells (502)) on each of the multiple columns of memory. In block (1105), the circuit accumulates output signals on each of the multiple columns of memory through each of the multiple digital counters (e.g., digital counter (604)). The multiple memory cells store multiple bits representing the weights of the neural network, and the multiple memory cells on each of the multiple columns correspond to different word lines (e.g., word lines (504)) of memory. In some embodiments, the circuit counts the amount of specific logic values ​​(e.g., logic highs) of output signals on each of the multiple columns through the digital counter. In some embodiments, the circuit generates one or more pulses based on the output signals of the multiple memory cells on the columns through a pulse generator coupled to each of the multiple columns, wherein the accumulation of output signals through the digital counter is based on one or more pulses.

[0087] In block (1110), the circuit adds the output signals of a plurality of digital counters through an adder circuit (e.g., an adder tree circuit (512)). In block (1115), the circuit accumulates the output signals of the adder circuit through an accumulator (e.g., an accumulator circuit (516)). In block (1120), the circuit can generate a DCIM output signal based on the accumulation of the output signals of the adder circuit. In some embodiments, the circuit generates output signals during a plurality of computation cycles through a plurality of memory cells of each of a plurality of columns, and the digital counter is configured to count the amount of logic highs of the digital output signals.

[0088] In some embodiments, the digital counter comprises a set of flip-flops (e.g., flip-flops (606)), wherein the clock input of a first flip-flop (e.g., flip-flop (6060)) of the set of flip-flops is coupled to a column, and the output of the first flip-flop is coupled to the clock input of a second flip-flop (e.g., flip-flop (6061)) of the set of flip-flops. In some embodiments, the circuit generates bits of the output signal generated by the digital counter through each of the flip-flops of the set. A half-latch circuit may be coupled to the output of each of the flip-flops of the set.

[0089] In some embodiments, the circuit applies a first delay to output signals generated by multiple memory cells on a first column of multiple columns (e.g., through one of the delay circuits (750)) and applies a second delay to output signals generated by multiple memory cells on a second column of multiple columns (e.g., through another of the delay circuits (750)). The first delay may be different from the second delay.

[0090] In certain embodiments, the circuit accumulates output signals on each of a plurality of different columns of memory through each of a plurality of different digital counters, wherein a plurality of different memory cells are on each of a plurality of different columns, and the plurality of different memory cells store a plurality of bits representing weights of a neural network, wherein a plurality of different memory cells for each of a plurality of different columns correspond to different word lines of memory. The circuit can select output signals from a plurality of digital counters during a first weighting cycle through a multiplexer (e.g., multiplexer (1004)), and the output signals are added through an adder circuit based on the selection of output signals. The circuit can also select other output signals from a plurality of digital counters during a second weighting cycle through a multiplexer, and add other output signals based on the selection of other output signals through an adder circuit.

[0091] Exemplary processing systems for in-memory computation

[0092] Fig. 12 illustrates an example electronic device 1200. The electronic device 1200 is Fig. 11 It may be configured to perform the methods described in this specification, including the operations (1100) described in relation to.

[0093] The electronic device (1200) includes a central processing unit (CPU) (1202), which may be a multi-core CPU in some embodiments. Instructions executed by the CPU (1202) may be loaded, for example, from a program memory associated with the CPU (1202) or from memory (1224).

[0094] The electronic device (1200) also includes additional processing blocks tailored to specific functions, such as a graphics processing unit (GPU) (1204), a digital signal processor (DSP) (1206), a neural processing unit (NPU) (1208), a multimedia processing block (1210), a multimedia processing block (1210), and a wireless connection processing block (1212). In one implementation, the NPU (1208) is implemented in one or more of a CPU (1202), a GPU (1204), and / or a DSP (1206).

[0095] In some embodiments, the wireless connection processing block (1212) may include components for, for example, a 3rd generation (3G) connection, a 4th generation (4G) connection (e.g., 4G LTE), a 5th generation connection (e.g., 5G or NR), a Wi-Fi connection, a Bluetooth connection, and wireless data transmission standards. The wireless connection processing block (1212) is additionally connected to one or more antennas (1214) to facilitate wireless communication.

[0096] The electronic device (1200) may also include one or more sensor processors (1216) associated with any type of sensor, one or more image signal processors (1218) associated with any type of image sensor, and / or a navigation processor (1220) that may include satellite-based positioning system components (e.g., Global Positioning System (GPS) or Global Navigation Satellite System (GLONASS)) as well as inertial positioning system components.

[0097] The electronic device (1200) may also include one or more input and / or output devices (1222), such as screens, touch-sensitive surfaces (including touch-sensitive displays), physical buttons, speakers, microphones, etc. In some embodiments, one or more of the processors of the electronic device (1200) may be based on an ARM (advanced RISC (reduced instruction set computing) machine) instruction set.

[0098] The electronic device (1200) also includes a memory (1224) representing one or more static and / or dynamic memories, such as dynamic random access memory, flash-based static memory, etc. In this example, the memory (1224) includes computer executable components that can be executed by one or more of the aforementioned processors of the electronic device (1200) or the CIM controller (1232) (also referred to as the “control circuit unit”). For example, the electronic device (1200) may include a CIM circuit (1226), such as the circuit (500) as described herein. The CIM circuit (1226) may be controlled through the CIM controller (1232). For example, in some embodiments, the memory (1224) may include code (1224A) for storage (e.g., storing weights in memory cells) and code (1224B) for computing (e.g., performing neural network computation by applying activation inputs). As illustrated, the CIM controller (1232) may include circuitry (1228A) for storage (e.g., storing weights in memory cells) and circuitry (1228B) for computing (e.g., performing neural network computation by applying activation inputs). The described components and others not described may be configured to perform various embodiments of the methods described herein.

[0099] In some embodiments, such as when the electronic device (1200) is a server device, Fig. 12 From the example described in, various embodiments such as one or more of a multimedia processing block (1210), a wireless connection processing block (1212), an antenna (1214), sensor processors (1216), ISPs (1218), or a navigation processor (1220) may be omitted.

[0100] Exemplary provisions

[0101] Clause 1. A circuit for in-memory computation, comprising: a memory having a plurality of columns; a plurality of memory cells on each column of the memory—the plurality of memory cells are configured to store a plurality of bits representing weights of a neural network, and the plurality of memory cells on each of the plurality of columns are on different word lines of the memory—; a plurality of digital counters—each of the plurality of digital counters is coupled to each of the columns of the memory—; an adder circuit coupled to the outputs of the plurality of digital counters; and an accumulator coupled to the output of the adder circuit.

[0102] Clause 2. The circuit of Clause 1, further comprising a pulse generator having an input portion coupled to each of the columns, wherein the output portion of the pulse generator is coupled to the input portion of the digital counter.

[0103] Clause 3. A circuit according to Clause 1 or Clause 2, wherein a plurality of memory cells on each of the plurality of columns are configured to generate digital output signals during a plurality of computation cycles, and a digital counter is configured to count the amount of logic highs of the digital output signals.

[0104] Clause 4. In any one of Clauses 1 to 3, the digital counter comprises a circuit including a counter circuit.

[0105] Clause 5. A circuit according to any one of Clauses 1 to 4, wherein the digital counter comprises a set of flip-flops, the clock input of a first flip-flop among the set of flip-flops is coupled to the column, and the output of the first flip-flop is coupled to the clock input of a second flip-flop among the set of flip-flops.

[0106] Clause 6. A circuit according to Clause 5, wherein the output of each of the flip-flops of the set provides a bit of a digital signal generated by the digital counter.

[0107] Clause 7. A circuit according to Clause 5 or Clause 6, further comprising a half-latch circuit coupled to the output of each of the flip-flops of the set.

[0108] Clause 8. A circuit in any one of Clauses 1 to 7, further comprising delay circuits, wherein each delay circuit is coupled between a plurality of memory cells on each of the plurality of columns and each of the plurality of digital counters.

[0109] Clause 9. In Clause 8, the delay circuits on the plurality of columns have different delays.

[0110] Clause 10. A circuit further comprising, in any one of Clauses 1 through 9, other plurality of memory cells on each of a plurality of different columns of the memory—the other plurality of memory cells are configured to store a plurality of bits representing weights of the neural network, and the other plurality of memory cells on each of a plurality of different columns are on different word lines of the memory—; other plurality of digital counters—each of the other plurality of digital counters is coupled to a respective column of a plurality of different columns of the memory—; and outputs of the plurality of digital counters and a multiplexer coupled to the outputs of the other plurality of digital counters.

[0111] Clause 11. The circuit of Clause 10, wherein the multiplexer is configured to provide output signals of the plurality of digital counters to the adder circuit during a first weighting cycle; and to provide output signals of the other plurality of digital counters to the adder circuit during a second weighting cycle.

[0112] Clause 12. In any one of Clauses 1 to 11, the adder circuit comprises an adder tree configured to add the output signals of the plurality of digital counters.

[0113] Clause 13. In Clause 12, one or more adders of the adder tree comprise a bit shift and add circuit.

[0114] Clause 14. A clock generator circuit having a first output portion configured to output a first clock signal and a second output portion configured to output a second clock signal, in any one of Clauses 1 to 13; and further comprising an activation circuit configured to provide activation signals to a plurality of memory cells, wherein the activation circuit is coupled to the first output portion of the clock generator circuit and is configured to operate based on the first clock signal, and the adder circuit is coupled to the second output portion of the clock generator circuit and is configured to operate based on the second clock signal, and the second clock signal has a frequency different from the first clock signal.

[0115] Clause 15. The clock generator circuit of Clause 14, comprising an edge-pulse converter configured to generate the first clock signal by generating a pulse based on detecting the edge of the second clock signal.

[0116] Clause 16. In Clause 15, the above edge includes a rising edge, circuit.

[0117] Clause 17. A circuit comprising, in any one of Clauses 1 through 16, a sensing amplifier further coupled between each column and the digital counter.

[0118] Clause 18. A circuit according to any one of Clauses 1 through 17, wherein the plurality of memory cells are configured to be sequentially activated based on different activation inputs, and the digital counter is configured to accumulate output signals from the plurality of memory cells after the plurality of memory cells are sequentially activated.

[0119] Clause 19. A method for in-memory computation, comprising the steps of: accumulating output signals on each of a plurality of columns of memory through each of a plurality of digital counters, wherein a plurality of memory cells are on each of the plurality of columns, and the plurality of memory cells store a plurality of bits representing weights of a neural network, and the plurality of memory cells on each of the plurality of columns are on different word lines of the memory; adding the output signals of the plurality of digital counters through an adder circuit; and accumulating the output signals of the adder circuit through an accumulator.

[0120] Clause 20. The method of Clause 19, further comprising the step of generating one or more pulses based on output signals of a plurality of memory cells on the column through a pulse generator coupled to each column of the plurality of columns, and the step of accumulating the output signals through the digital counter is based on the one or more pulses.

[0121] Clause 21. A method according to Clause 19 or Clause 20, further comprising the step of counting the amount of specific logic values ​​of the output signals on each of the plurality of columns through the digital counter.

[0122] Clause 22. A method according to any one of Clauses 19 to 21, further comprising the step of generating output signals during a plurality of computation cycles through a plurality of memory cells on each of the plurality of columns, wherein the digital counter is configured to count the amount of logic highs of the digital output signals.

[0123] Clause 23. A method in which, in any one of Clauses 19 through 22, the digital counter comprises a 1 counter circuit.

[0124] Clause 24. A method in which, in any one of Clauses 19 through 23, the digital counter comprises a set of flip-flops, wherein the clock input of a first flip-flop among the set of flip-flops is coupled to the column, and the output of the first flip-flop is coupled to the clock input of a second flip-flop among the set of flip-flops.

[0125] Clause 25. The method of Clause 24, further comprising the step of generating a bit of the output signal generated by the digital counter through each of the flip-flops of the set.

[0126] Clause 26. A method according to Clause 24 or Clause 25, further comprising the step of latching the output signal of each of the respective flip-flops of the set through a half-latch circuit.

[0127] Clause 27. A method further comprising, in any one of Clauses 19 to 26, the step of applying a first delay to output signals generated by a plurality of memory cells on a first column among the plurality of columns; and the step of applying a second delay to output signals generated by a plurality of memory cells on a second column among the plurality of columns.

[0128] Clause 28. In Clause 27, the first delay is a method different from the second delay.

[0129] Clause 29. A method further comprising, in any one of Clauses 19 through 28, the step of accumulating output signals on each of a plurality of different columns of the memory through each of a plurality of other digital counters - wherein a plurality of other memory cells are on each of the plurality of different columns, and the plurality of other memory cells store a plurality of bits representing weights of the neural network, and the plurality of other memory cells on each of the plurality of different columns are on different word lines of the memory -; the step of selecting the output signals from the plurality of digital counters during a first weighting cycle through a multiplexer - wherein the output signals are added through an adder circuit based on the selection of the output signals -; the step of selecting other output signals from the plurality of other digital counters during a second weighting cycle through the multiplexer; and the step of adding the other output signals based on the selection of the other output signals through an adder circuit.

[0130] Clause 30. An apparatus for in-memory computation, comprising: means for counting the amount of specific logic values ​​of output signals on each of a plurality of columns of memory—wherein a plurality of memory cells are on each of the plurality of columns, wherein the plurality of memory cells are configured to store a plurality of bits representing weights of a neural network, and wherein the plurality of memory cells on each of the plurality of columns are on different word lines of the memory—; means for adding the output signals of the plurality of digital counters; and means for accumulating the output signals of the means for adding.

[0131] Additional considerations

[0132] The preceding description is provided to enable those skilled in the art to perform the various functions described herein. The examples discussed herein do not limit the categories, applicability, or embodiments described in the claims. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments. For example, changes may be made to the function and arrangement of the elements discussed without departing from the scope of this disclosure. Various examples may appropriately omit, substitute, or add various procedures or components. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Furthermore, features described in relation to some examples may be combined in some other examples. For instance, an apparatus may be implemented or a method may be performed using any number of the embodiments described herein. Additionally, the scope of the present disclosure is intended to cover any structure, function, or device or method implemented using any other structure and function in addition to or other than the various aspects of the present disclosure described herein. It should be understood that any aspect of the present disclosure disclosed herein may be embodied by one or more elements of the claims.

[0133] As used herein, the word "exemplary" means "serving as an example, illustration, or illustration." Any aspect described herein as "exemplary" should not be interpreted as necessarily being more desirable or advantageous than other aspects.

[0134] As used herein, the phrase referring to "at least one of" a list of items refers to any combination of the items, including single members. For example, "at least one of a, b, or c" is intended to cover a, b, c, ab, ac, bc, and abc, as well as any combination with multiples of the same element (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc, and ccc or any other ordering of a, b, and c).

[0135] As used herein, the term “determining” encompasses a wide variety of actions. For example, “determining” may include calculation, computing, processing, derivation, inspection, searching (e.g., searching in a table, database, or other data structure), verification, etc. Additionally, “determining” may include receiving (e.g., receiving information), accessing (e.g., accessing data in memory), etc. Furthermore, “determining” may include solving, selecting, selecting, setting, etc.

[0136] The methods disclosed herein include one or more steps or actions for achieving the methods. Method steps and / or actions may be interchangeable without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions may be modified without departing from the scope of the claims. Additionally, various operations of the methods described above may be performed by any suitable means capable of performing corresponding functions. These means may include various hardware and / or software component(s) and / or module(s), including but not limited to circuits, application-specific integrated circuits (ASICs), or processors. Generally, where operations illustrated in the drawings exist, such operations may have corresponding counterpart means-plus-function components having similar numbering. For example, the means for adding may include an adder tree such as adder trees (510) or a weighted shift adder tree (512), or an accumulator such as accumulators (606). The means for accumulating may include an accumulator such as an activation shift accumulator (516). The means for detecting may include an SA such as SAs (602).

[0137] The following claims are not intended to be limited to the embodiments set forth in this specification, but are to be followed in their entire scope consistent with the expression of the claims. Within the claims, references to singular elements are intended to mean "one and only one" rather than "one or more" unless specifically stated otherwise. Unless otherwise specifically stated, the term "some" refers to one or more. No claim element shall be interpreted under the provisions of 35 USC §112(f) unless the element is explicitly described using the phrase "means of doing," or, in the case of a method claim, the element is not described using the phrase "step of doing." All structural and functional equivalents to the elements of the various embodiments described throughout this disclosure, known to or to be known to those skilled in the art, are expressly incorporated by reference into this specification and are intended to be encompassed by the claims. Furthermore, nothing disclosed in this specification is intended to be made available to the public, regardless of whether such disclosure is explicitly cited in the claims.

Claims

Claim 1 A circuit for in-memory computation, comprising: a memory having a plurality of columns; a plurality of memory cells on each column of the memory, wherein the plurality of memory cells are configured to store a plurality of bits representing weights of a neural network, and the plurality of memory cells on each column of the memory correspond to different word lines of the memory; an activation circuit coupled to the different word lines; a plurality of digital counters, wherein each of the plurality of digital counters is coupled to a respective column of the plurality of columns of the memory; an adder circuit coupled to the outputs of the plurality of digital counters; an accumulator coupled to the output of the adder circuit; and a frequency converter configured to generate a first clock signal for the activation circuit based on a second clock signal for the adder circuit. Claim 2 A circuit for memory-in-computation according to claim 1, further comprising a pulse generator having an input portion coupled to each of the columns, wherein the output portion of the pulse generator is coupled to the input portion of the digital counter for each of the columns. Claim 3 A circuit for in-memory computation according to claim 1, wherein each of the plurality of columns, the plurality of memory cells are configured to generate digital output signals during a plurality of computation cycles, and the digital counter for each column is configured to count the amount of logic highs of the digital output signals. Claim 4 In claim 1, a circuit for in-memory computation, wherein each of the plurality of digital counters comprises one's counter circuit. Claim 5 A circuit for memory-in-computation according to claim 1, wherein each of the plurality of digital counters comprises a set of flip-flops, the clock input of a first flip-flop among the set of flip-flops is coupled to the column, and the output of the first flip-flop is coupled to the clock input of a second flip-flop among the set of flip-flops. Claim 6 In paragraph 5, a circuit for in-memory computation, wherein the output of each of the flip-flops of the set provides a bit of a digital signal generated by each of the digital counters of the plurality of digital counters. Claim 7 A circuit for memory-in-computation according to claim 5, further comprising a half-latch circuit coupled to the output of each of the flip-flops of the set. Claim 8 A circuit for in-memory computation according to claim 1, further comprising delay circuits, wherein each delay circuit is coupled between the plurality of memory cells of each column of the plurality of columns and each digital counter of the plurality of digital counters. Claim 9 In paragraph 8, the delay circuits of the plurality of columns are circuits for memory-in-computation having different delays. Claim 10 A circuit for in-memory computation according to claim 1, further comprising: different multiple memory cells on each of a plurality of different columns of the memory, wherein the different multiple memory cells are configured to store a plurality of bits representing weights of the neural network, and the different multiple memory cells on each of a plurality of different columns correspond to different word lines of the memory; different multiple digital counters, wherein each of the different multiple digital counters is coupled to a respective column of the plurality of different columns of the memory; and a multiplexer coupled to the outputs of the plurality of digital counters and the outputs of the different multiple digital counters. Claim 11 A circuit for in-memory computation according to claim 10, wherein the multiplexer is configured to provide output signals of the plurality of digital counters to the adder circuit during a first weighting cycle; and to provide output signals of the other plurality of digital counters to the adder circuit during a second weighting cycle. Claim 12 In claim 1, the circuit for in-memory computation comprises an adder tree configured to add the output signals of the plurality of digital counters. Claim 13 In paragraph 12, one or more adders of the adder tree are circuits for in-memory computation, comprising bit shift and add circuits. Claim 14 A circuit for memory-in-computation according to claim 1, further comprising a clock generator circuit including the frequency converter, a first output portion configured to output the first clock signal, and a second output portion configured to output the second clock signal; wherein the activation circuit is configured to provide activation signals to the plurality of memory cells, the activation circuit is coupled to the first output portion of the clock generator circuit and is configured to operate based on the first clock signal, and the adder circuit is coupled to the second output portion of the clock generator circuit and is configured to operate based on the second clock signal, and the second clock signal has a frequency different from the first clock signal. Claim 15 A circuit for in-memory computation according to claim 14, wherein the frequency converter comprises an edge-to-pulse converter configured to generate the first clock signal by generating a pulse based on detecting the edge of the second clock signal. Claim 16 In paragraph 15, the above edge includes a rising edge, a circuit for in-memory computation. Claim 17 A circuit for memory-in-computation according to claim 1, further comprising a sensing amplifier coupled between each column and the digital counter for each column. Claim 18 A circuit for in-memory computation according to claim 1, wherein the plurality of memory cells are configured to be sequentially activated based on different activation inputs, and each digital counter of the plurality of digital counters is configured to accumulate output signals from the plurality of memory cells after the plurality of memory cells are sequentially activated. Claim 19 A method for in-memory computation, comprising: providing activation signals to different word lines of memory through an activation circuit; accumulating output signals on each of a plurality of columns of memory through each of a plurality of digital counters, wherein a plurality of memory cells are on each of the plurality of columns, the plurality of memory cells store a plurality of bits representing weights of a neural network, and the plurality of memory cells of each of the plurality of columns correspond to the different word lines of memory; adding the output signals of the plurality of digital counters through an adder circuit; accumulating the output signals of the adder circuit through an accumulator; and generating a first clock signal for the activation circuit based on a second clock signal for the adder circuit through a frequency converter. Claim 20 In claim 19, the method for memory-in-computation further comprises the step of generating one or more pulses based on the output signals of the plurality of memory cells of the column through a pulse generator coupled to each of the plurality of columns, and the step of accumulating the output signals through the digital counter for each of the columns is based on the one or more pulses. Claim 21 A method for in-memory computation according to claim 19, further comprising the step of counting the amount of specific logic values ​​of the output signals on each of the plurality of columns through the digital counter for each of the columns. Claim 22 A method for in-memory computation according to claim 19, further comprising the step of generating the output signals on each of the plurality of columns during a plurality of computation cycles through the plurality of memory cells of each of the plurality of columns, wherein the digital counter for each of the column is configured to count the amount of logic highs of the output signals. Claim 23 A method for in-memory computation according to claim 19, wherein each of the plurality of digital counters comprises a 1 counter circuit. Claim 24 A method for in-memory computation according to claim 19, wherein each of the plurality of digital counters comprises a set of flip-flops, the clock input of a first flip-flop of the set of flip-flops is coupled to the column, and the output of the first flip-flop is coupled to the clock input of a second flip-flop of the set of flip-flops. Claim 25 A method for in-memory computation according to claim 24, further comprising the step of generating a bit of the output signal generated by each of the digital counters of the plurality of digital counters through each of the flip-flops of the set. Claim 26 A method for memory-in-computation according to claim 24, further comprising the step of latching the output signal of each of the respective flip-flops of the set of flip-flops through a half-latch circuit. Claim 27 A method for in-memory computation according to claim 19, further comprising: a step of applying a first delay to the output signals generated by the plurality of memory cells in the first column of the plurality of columns; and a step of applying a second delay to the output signals generated by the plurality of memory cells in the second column of the plurality of columns. Claim 28 In paragraph 27, the above first delay is different from the above second delay, a method for in-memory computation. Claim 29 A method for in-memory computation according to claim 19, further comprising: a step of accumulating output signals on each of a plurality of different columns of the memory through each of a plurality of other digital counters, wherein a plurality of other memory cells are on each of the plurality of different columns, and the plurality of other memory cells store a plurality of bits representing weights of the neural network, and each of the plurality of other columns corresponds to different word lines of the memory; a step of selecting output signals from the plurality of digital counters through a multiplexer during a first weighting cycle, wherein the output signals are added through an adder circuit based on the selection of the output signals; a step of selecting other output signals from the plurality of other digital counters through the multiplexer during a second weighting cycle; and a step of adding the other output signals through an adder circuit based on the selection of the other output signals. Claim 30 An apparatus for in-memory computation, comprising: means for providing activation signals to different word lines of memory; means for counting the amount of specific logic values ​​of output signals on each of a plurality of columns of memory, wherein a plurality of memory cells are on each of the plurality of columns, and the plurality of memory cells are configured to store a plurality of bits representing weights of a neural network, and the plurality of memory cells of each of the plurality of columns correspond to different word lines of memory; means for adding output signals of the means for counting; means for accumulating output signals of the means for adding; and means for generating a first clock signal for the means for providing based on a second clock signal for the means for adding.