Systems and methods for input reference generation technique for in-memory computing array

The reference voltage circuit in the in-memory computing array dynamically selects activation voltages based on input analog voltages, addressing efficiency challenges and enhancing computational speed and energy efficiency.

WO2025122550A1PCT designated stage expired Publication Date: 2025-06-12ENCHARGE AI INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/058364
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-04
Filing Date
2024-12-04
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

In-memory computing arrays face challenges in providing efficient reference voltages to computing cells, which affects the speed and energy consumption of neural network computations.

Method used

The proposed solution involves a reference voltage circuit that selects appropriate activation voltages for computing cells based on input analog voltages, using a multiplexer and controller system that adjusts voltages during Reset and Evaluate phases, potentially utilizing a look-up table for optimal voltage selection.

Benefits of technology

This approach enhances the efficiency of in-memory computing by reducing the number of distinct voltages required, improving energy efficiency and computational speed while maintaining accurate analog voltage levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024058364_12062025_PF_FP_ABST
    Figure US2024058364_12062025_PF_FP_ABST
Patent Text Reader

Abstract

A device can include a compute-in-memory (CIM) array of computing cells, including a plurality of rows and a plurality of columns, each computing cell including a memory cell, a computing circuitry, and a capacitor for storing a result of computation, wherein the computing cell receives a pair of activation voltages that are representative of an input analog voltage to be multiplied with data stored in the memory cell. A reference voltage circuit receiving a first set of voltages and a second set of voltages, the reference voltage circuit configured to: during a Reset duration, select a voltage of a first value from the first set of voltages as a first of the pair of activation voltages and select a voltage of a second value from the second set of voltages as a second of the pair of activation voltages.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR INPUT REFERENCE GENERATION TECHNIQUE FOR IN-MEMORY COMPUTING ARRAYCROSS-REFERENCE TO RELATED APPLICATIONThis Application claims priority to and the benefit of United States Provisional Patent Application Number 63 / 606,019, filed December 4, 2023.TECHNICAL FIELD

[0001] This disclosure relates to in-memory computing arrays, and in particular to providing reference voltages to computing cells in in-memory computing arrays.DESCRIPTION OF THE RELATED TECHNOLOGY

[0002] Using in-memory computing for neural network acceleration is an emerging and innovative approach that leverages the unique properties of memory devices to enhance the speed and efficiency of neural network computations. Traditional neural network training and inference processes involve moving data back and forth between memory (RAM) and processing units (CPUs or GPUs), which can be a significant bottleneck in terms of speed and energy consumption. In-memory computing seeks to overcome these limitations by processing data directly within the memory itself.SUMMARY

[0003] In some aspects, the techniques described herein relate to an in-memory computing architecture, including: a compute-in-memory (CIM) array of computing cells, the CIM array including a plurality of rows and a plurality of columns, each computing cell including a memory cell, a computing circuitry, and a capacitor for processing and storing a result of computation, wherein the computing cell receives a pair of activation voltages that are representative of an input analog voltage to be multiplied with data stored in the memory cell; a reference voltage circuit receiving a first set of voltages and a second set of voltages, the reference voltage circuit configured to: during a Reset duration, select a voltage of a first value from the first set of voltages as a first of the pair of activation voltages and select a voltage of a second value from the second set of voltages as a second of the pair of activation voltages, during an Evaluate duration subsequent to the Reset duration, select a voltage of a third value from the first set of voltages as the first of the pair of activation voltages and select a voltage of a fourth value from the second set ofvoltages as the second of the pair of activation voltages.

[0004] In some aspects, the techniques described herein relate to an in-memory computing architecture, wherein the first set of voltages is same as the second set of voltages.

[0005] In some aspects, the techniques described herein relate to an in-memory computing architecture, wherein the reference voltage circuit is configured to select voltages for the first of the pair of activation voltages and the second of the pair of activation voltages based on data or signal indicating a value of the input analog voltage.

[0006] In some aspects, the techniques described herein relate to an in-memory computing architecture, wherein the reference voltage circuit is configured to select voltages for the first of the pair of activation voltages and the second of the pair of activation voltages based on data or signal indicating a value of the input analog voltage and a look-up table including a list of input analog voltages and corresponding values of pairs of voltages.

[0007] In some aspects, the techniques described herein relate to an in-memory computing architecture, wherein the reference voltage circuit includes: a first multiplexer including a plurality of first input ports coupled with the first set of voltages, a first control input port receiving a first control input, and a first output port outputting one of the first set of voltages as the first of the pair of activation voltages based on the first control signal, a second multiplexer including a plurality of second input ports coupled with the second set of voltages, a second control input port receiving a second control input, and a second output port outputting one of the second set of voltages as the second of the pair of activation voltages based on the second control input, and a controller coupled with the first multiplexer and the second multiplexer, the controller configured to provide the first multiplexer and the second multiplexer with the first control input and the second control input, respectively, based on a multi -bit digital signal.

[0008] In some aspects, the techniques described herein relate to an in-memory computing architecture, wherein the controller uses a multi-bit digital signal and a lookup table.

[0009] In some aspects, the techniques described herein relate to an in-memory computing architecture, wherein a total number of distinct voltages among the first set of voltages and the second set of voltages is less than the total number of possible digital codes, which is based on a multi-bit digital representation of the input analog voltage.

[0010] In some aspects, the techniques described herein relate to an in-memorycomputing architecture, wherein any two adjacent input analog voltages of all possible input analog voltages are separated by a same value.

[0011] In some aspects, the techniques described herein relate to a method for providing analog voltages to in-memory computing architecture including a compute in-memory (CIM) array of computing cells, the CIM array including a plurality of rows and a plurality of columns, each computing cell including a memory cell, a computing circuitry, and a capacitor for processing and storing a result of computation, wherein the computing cell receives a pair of activation voltages that are representative of an input analog voltage to be multiplied with data stored in the memory cell, and a reference voltage circuit receiving a first set of voltages and a second set of voltages, the method including: during a Reset duration, selecting a voltage of a first value from the first set of voltages as a first of a pair of activation voltages and select a voltage of a second value from the second set of voltages as a second of the pair of activation voltages; and during an Evaluate duration subsequent to the Reset duration, selecting a voltage of a third value from the first set of voltages as the first of the pair of activation voltages and select a voltage of a fourth value from the second set of voltages as the second of the pair of activation voltages.

[0012] In some aspects, the techniques described herein relate to a method, wherein the first set of voltages are the same as the second set of voltages.

[0013] In some aspects, the techniques described herein relate to a method, including: selecting voltages for the first of the pair of the pair of activation voltages and the second of the pair of activation voltages based on data or signal indicating a value of the input analog voltage.

[0014] In some aspects, the techniques described herein relate to a method, including: selecting voltages for the first of the pair of activation voltages and the second of the pair of activation voltages based on data or signal indicating a value of the input analog voltage and a look-up table including a list of input analog voltages and corresponding values of pairs of voltages.

[0015] In some aspects, the techniques described herein relate to a method, wherein the reference voltage circuit includes: a first multiplexer including a plurality of first input ports coupled with the first set of voltages, a first control input port receiving a first control input, and a first output port outputting one of the first set of voltages as the first of the pair of activation voltages based on the first control input, a second multiplexer including a plurality of second input ports coupled with the second set of voltages, a second control input port receiving a second control input, and a second output portoutputting one of the second set of voltages as the second of the pair of activation voltages based on the second control input, and a controller coupled with the first multiplexer and the second multiplexer, the controller configured to provide the first multiplexer and the second multiplexer with the first control input and the second control input, respectively, based on a multi -bit digital signal.

[0016] In some aspects, the techniques described herein relate to a method, including: the controller using the multi -bit digital signal and a look-up table to provide first multiplexer and the second multiplexer with the first control input and the second control input, respectively.

[0017] In some aspects, the techniques described herein relate to a method, wherein a total number of distinct voltages among the first set of voltages and the second set of voltages is less than a total number of possible digital codes, which is based on a multibit digital representation of the input analog voltage.

[0018] In some aspects, the techniques described herein relate to a method, wherein any two adjacent input analog voltages of all of the possible input analog voltages are separated by a same value.

[0019] In some aspects, the techniques described herein relate to a system, including a compute-in-memory (CIM) array of computing cells organized as a plurality of rows and a plurality of columns, each computing cell including a memory cell for storing data, computing circuitry for processing a computation, and a capacitor for storing a result of the computation, wherein the computing cell receives a pair of activation voltages that represent an input analog voltage to be multiplied with the data; and a reference voltage circuit configured to: receive a first set of voltages and a second set of voltages; during a Reset duration, select a voltage of a first value from the first set of voltages as a first of the pair of activation voltages and select a voltage of a second value from the second set of voltages as a second of the pair of activation voltages, and during an Evaluate duration subsequent to the Reset duration, select a voltage of a third value from the first set of voltages as the first of the pair of activation voltages and select a voltage of a fourth value from the second set of voltages as the second of the pair of activation voltages.

[0020] In some aspects, the techniques described herein relate to a system, wherein the first set of voltages is same as the second set of voltages.

[0021] In some aspects, the techniques described herein relate to a system, wherein the reference voltage circuit is configured to select voltages for the first of the pair ofactivation voltages and the second of the pair of activation voltages based on data or signal indicating a value of the input analog voltage.

[0022] In some aspects, the techniques described herein relate to a system, wherein the reference voltage circuit is configured to select voltages for the first of the pair of activation voltages and the second of the pair of activation voltages based on data or signal indicating a value of the input analog voltage and a look-up table including a list of input analog voltages and corresponding values of pairs of voltages.

[0023] In some aspects, the techniques described herein relate to a system, wherein the reference voltage circuit includes a first multiplexer including a plurality of first input ports coupled with the first set of voltages, a first control input port receiving a first control input, and a first output port outputting one of the first set of voltages as the first of the pair of activation voltages based on the first control signal, a second multiplexer including a plurality of second input ports coupled with the second set of voltages, a second control input port receiving a second control input, and a second output port outputting one of the second set of voltages as the second of the pair of activation voltages based on the second control input, and a controller coupled with the first multiplexer and the second multiplexer, the controller configured to provide the first multiplexer and the second multiplexer with the first control input and the second control input, respectively, based on a multi -bit digital signal.

[0024] In some aspects, the techniques described herein relate to a system, wherein the controller uses the multi-bit digital signal and a look-up table.

[0025] In some aspects, the techniques described herein relate to a system, wherein a total number of distinct voltages among the first set of voltages and the second set of voltages is less than the total number of possible digital codes, which is based on a multi-bit digital representation of the input analog voltage.

[0026] In some aspects, the techniques described herein relate to a system, wherein any two adjacent input analog voltages of all possible input analog voltages are separated by a same value.BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 depicts a block diagram of an example in-memory computing architecture.

[0028] Figure 2 shows a block diagram of a compute in-memory array.

[0029] Figure 3 shows an example circuit diagram of the computing cells discussed above in relation to Figure 2.

[0030] Figure 4 shows a Reset phase and an Evaluate phase of a voltage circuit performing dynamic-range doubling.

[0031] Figure 5 shows a first example reference voltage circuit.

[0032] Figure 6 A and Figure 6B show a second example reference voltage circuit.

[0033] Figure 7 shows a first table of a first example set of activation voltages for a desired set of input analog voltages.

[0034] Figure 8 shows a second table of a second example set of activation voltages for a desired set of input analog voltages.

[0035] Figure 9 shows a third table of a third example set of activation voltages for a desired set of input analog voltages.

[0036] Figure 10 shows a fourth table of a fourth example set of activation voltages for a desired set of input analog voltages.

[0037] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION

[0038] The various concepts introduced above and discussed in greater detail below can be implemented in any of numerous ways, as the described concepts are not limited to any particular manner of implementation. Examples of specific implementations and applications are provided primarily for illustrative purposes.

[0039] As will be apparent to those of skill in the art upon reading this disclosure, each of the individual aspects described and illustrated herein has discrete components and features which can be readily separated from or combined with the features of any of the other several aspects without departing from the scope or spirit of the present disclosure.

[0040] Any recited method can be carried out in the order of events recited or in any other order that is logically possible. That is, unless otherwise expressly stated, it is in no way intended that any method or aspect set forth herein be construed as requiring that its steps be performed in a specific order. Accordingly, where a method claim does not specifically state in the claims or descriptions that the steps are to be limited to a specific order, it is no way intended that an order be inferred, in any respect. This holds for any possible nonexpress basis for interpretation, including matters of logic with respect to arrangement of steps or operational flow, plain meaning derived from grammatical organization or punctuation, or the number or type of aspects described in the specification.

[0041] All publications mentioned herein are incorporated herein by reference to discloseand describe the methods and / or materials in connection with which the publications are cited. All such publications and patents are herein incorporated by references as if each individual publication or patent were specifically and individually indicated to be incorporated by reference. Such incorporation by reference is expressly limited to the methods and / or materials described in the cited publications and patents and does not extend to any lexicographical definitions from the cited publications and patents. Any lexicographical definition in the publications and patents cited that is not also expressly repeated in the instant specification should not be treated as such and should not be read as defining any terms appearing in the accompanying claims. The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided herein can be different from the actual publication dates, which can require independent confirmation.

[0042] While aspects of the present disclosure can be described and claimed in a particular statutory class, such as the system statutory class, this is for convenience only and one of skill in the art will understand that each aspect of the present disclosure can be described and claimed in any statutory class.

[0043] It is also to be understood that the terminology used herein is for the purpose of describing particular aspects only and is not intended to be limiting. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the disclosed compositions and methods belong. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the specification and relevant art and should not be interpreted in an idealized or overly formal sense unless expressly defined herein.

[0044] It should be noted that ratios, concentrations, amounts, and other numerical data can be expressed herein in a range format. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint. It is also understood that there are a number of values disclosed herein, and that each value is also herein disclosed as “about” that particular value in addition to the value itself. For example, if the value “10” is disclosed, then “about 10” is also disclosed. Ranges can be expressed herein as from “about” oneparticular value, and / or to “about” another particular value. Similarly, when values are expressed as approximations, by use of the antecedent “about,” it will be understood that the particular value forms a further aspect. For example, if the value “about 10” is disclosed, then “10” is also disclosed.

[0045] When a range is expressed, a further aspect includes from the one particular value and / or to the other particular value. For example, where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure, e.g. the phrase “x to y” includes the range from ‘x’ to ‘y’ as well as the range greater than ‘x’ and less than ‘y’. The range can also be expressed as an upper limit, e.g. ‘about x, y, z, or less’ and should be interpreted to include the specific ranges of ‘about x’, ‘about y’, and ‘about z’ as well as the ranges of Tess than x’, less than y’, and Tess than z’. Likewise, the phrase ‘about x, y, z, or greater’ should be interpreted to include the specific ranges of ‘about x’, ‘about y’, and ‘about z’ as well as the ranges of ‘greater than x’, greater than y’, and ‘greater than z’. In addition, the phrase “about ‘x’ to ‘y’”, where ‘x’ and ‘y’ are numerical values, includes “about ‘x’ to about ‘y’”.

[0046] It is to be understood that such a range format is used for convenience and brevity, and thus, should be interpreted in a flexible manner to include not only the numerical values explicitly recited as the limits of the range, but also to include all the individual numerical values or sub-ranges encompassed within that range as if each numerical value and sub-range is explicitly recited. To illustrate, a numerical range of “about 0.1% to 5%” should be interpreted to include not only the explicitly recited values of about 0.1% to about 5%, but also include individual values (e.g., about 1%, about 2%, about 3%, and about 4%) and the sub-ranges (e.g., about 0.5% to about 1.1%; about 5% to about 2.4%; about 0.5% to about 3.2%, and about 0.5% to about 4.4%, and other possible sub-ranges) within the indicated range.

[0047] As used herein, the terms “about,” “approximate,” “at or about,” and “substantially” mean that the amount or value in question can be the exact value or a value that provides equivalent results or effects as recited in the claims or taught herein. That is, it is understood that amounts, sizes, formulations, parameters, and other quantities and characteristics are not and need not be exact, but can be approximate and / or larger or smaller, as desired, reflecting tolerances, conversion factors, rounding off, measurement error and the like, and other factors known to those of skill in the art such that equivalent results or effects are obtained. In some circumstances, the value that provides equivalentresults or effects cannot be reasonably determined. In such cases, it is generally understood, as used herein, that “about” and “at or about” mean the nominal value indicated ±10% variation unless otherwise indicated or inferred. In general, an amount, size, formulation, parameter or other quantity or characteristic is “about,” “approximate,” or “at or about” whether or not expressly stated to be such. It is understood that where “about,” “approximate,” or “at or about” is used before a quantitative value, the parameter also includes the specific quantitative value itself, unless specifically stated otherwise.

[0048] Prior to describing the various aspects of the present disclosure, the following definitions are provided and should be used unless otherwise indicated. Additional terms can be defined elsewhere in the present disclosure.

[0049] includingincludingincludingincludingincludesincludingincludingAs used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items. Expressions such as “at least one of,” when preceding a list of elements, modify the entire list of elements and do not modify the individual elements of the list.

[0050] As used in the specification and the appended claims, the singular forms “a,” “an” and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to “a proton beam degrader,” “a degrader foil,” or “a conduit,” includes, but is not limited to, two or more such proton beam degraders, degrader foils, or conduits, and the like.

[0051] The terms “configured for” or “configured to,” as used herein with respect to a specified operation or function, refer to a device, component, circuit, structure, machine, signal, etc. that is physically constructed, programmed, formatted and / or arranged to perform the specified operation or function.

[0052] The various concepts introduced above and discussed in greater detail below can be implemented in any of numerous ways, as the described concepts are not limited to any particular manner of implementation. Examples of specific implementations and applications are provided primarily for illustrative purposes.

[0053] As used herein, the terms “optional” or “optionally” means that the subsequently described event or circumstance can or cannot occur, and that the description includes instances where said event or circumstance occurs and instances where it does not.

[0054] Unless otherwise specified, temperatures referred to herein are based on atmospheric pressure (i.e., one atmosphere).

[0055] In-memory Computing Architecture

[0056] Figure 1 depicts a block diagram of an example in-memory computingarchitecture 100. The in-memory computing architecture 100 can be adapted, for example, to a scalable neural network accelerator architecture based on in-memory computing (IMC). However, the in-memory computing architecture 100 is not limited to neural network applications, and can be employed in numerous applications where high data throughput with low power consumption is desired. The in-memory computing architecture 100 includes a plurality of Compute In-Memory unit (CIMU) tiles 102. The plurality of CIMU tiles 102 are arranged in an array within the architecture. The plurality of CIMU tiles 102 can be individually enabled / disabled based on the computations to be carried out by the in-memory computing architecture 100. In examples where the inmemory computing architecture 100 can be used to implement neural networks, the neural networks can be mapped to one or more CIMU tiles of the plurality of CIMU tiles 102. The remainder of the CIMU tiles of the plurality of CIMU tiles 102 can be disabled to reduce power consumption.

[0057] The in-memory computing architecture 100 can include, in part, activation buffers 104, segmented weight buffers 106, and one or more phase-locked loops (PLLs) 108. The activation buffers 104 can provide signals representative of activations from previous stages of computation, for instance previous layers in a neural network. The segmented weight buffers 106 can provide data required for computation together with the activations / data from previous stages, for instance these weight buffers could store the weights of neural network layers. The one or more PLLs 108 can provide reference clock signals to various portions of the in-memory computing architecture 100. The in-memory computing architecture 100 can also include off-chip interfaces 110 (referenced in Figure 1 as “off-chip control” element 110) for communication with off-chip processors or software to send and receive control or data signals. The off-chip interface 110 can, by itself or in concert with other elements, provide circuits and protocols for high-speed interfaces for wired or wireless connections involving data, control signals, or both, to other processors or other arrays of CIMU tiles, for example, enabling the in-memory computing architecture 100 to scale upward as desired.

[0058] Each of the plurality of CIMU tiles 102 can include a plurality of CIMUs 112, an on-chip network 114, and a weight network 116. While Figure 1 shows each of the plurality of CIMU tiles 102 including four CIMUs 112, this is only an example, and the CIMU tiles 102 can include fewer or more CIMUs 112. One or more of the CIMUs 112 can include a compute in-memory (CIM) array 118, compute dataflow buffers 120, programmable digital single instruction multiple data (SIMD) module 122, and aprogramming and control module 124. The CIM array 118 can be an array of computing cells, discussed further below. The CIM array 118 can carry out computations based on data stored in the computing cells and data provided by the activation buffers 104. The computing cells can be used to perform computational operations between inputs and data stored in a memory cell within the computing cells. The operations can include logical operations (AND, NOR, etc.) or multiplication operations carried out between inputs. The CIM array 118 can carry out matrix operations between multi -bit operands, which is particularly useful in neural network computations where activations are multiplied with weights. In some such applications, the weights can be stored in the memory cells of the CIM array 118 and activations can be provided as input vectors. Each computing cell in the CIM array 118 can perform the multiplication operation between a 1 -bit weight and a portion of the input activation, which can be represented in digital or analog signal form. Some example computing cells can generate a result that is in the form of an electrical signal. For example, the computing cell can output an analog voltage that is representative of the computation result. In some other examples, the computing cell can output an electrical current that is representative of the computation result. The electrical signals of various computing cells can be accumulated and processed to generate the overall matrix multiplication result. For example, electrical signals representative of computation from all computing cells in a single column of the CIM array 118 can be accumulated to represent a portion of the computation. Accumulated electrical signals from multiple columns of computing cells of the CIM array 118 can be combined and processed to generate an overall matrix multiplication result. For instances where the electrical signal generated by the computing cells is an electrical current, the currents from various computing cells within a column can be summed to generate a representative accumulated electrical current. In instances where the electrical signal generated by the computing cell is an analog voltage, the analog voltage generated by each computing cell can be stored in capacitors within the computing cell and then accumulated as a voltage that is representative of a portion of the overall matrix multiplication result. The accumulated result, whether an electrical current or an analog voltage, can be converted into digital form using analog to digital converters (ADCs) and further processed, stored, or passed on to other CIM arrays 118 for further computations.

[0059] The programmable digital SIMD 122 can have an instruction set for flexible element-wise operation and the compute dataflow buffers 120 can support wide range of neural network dataflows. Each CIMU 112 can provide a high-level of configurabilityand can be abstracted into a software library of instructions for interfacing with a compiler (for allocating / mapping an application, neural network and the like to the architecture), and where instructions can thus also be added prospectively. That is, the library can include single / fused instructions such as element mult / add, h(») activation, (N-step convolutional stride + matrix-vector-multiplication (MVM) + batch norm. +h(») activation + max. pool), (dense + MVM) and the like. In various nonlimiting examples, h(») can indicate an activation function, including without limitation the rectified linear unit ReLU(x) function, the sigmoid function (o(x)), and other such functions. Max pooling, a downsampling technique for reducing spatial dimensions to maintain computational efficiency while retaining other important features of the CIMU array or the network, can also be a subject of the computation. The N-step convolutional stride can refer to the number of pixels or other information bits that a kernel or convolutional filter moves or glides across the input image during convolution to effect operations like feature detection, pattern recognition, blurring, image sharpening, image recognition, and the like.

[0060] The on-chip network 114 (OCN) can include routing channels within Network In / Out Blocks, and a Switch Block, which provides flexibility via a disjoint architecture as shown, for example, by the disjoint buffer switch 133 in the enlarged view of OCN 114. This flexibility, among other benefits, enables modules that are independent of one another to work in parallel. The OCN 114 works with configurable CIMU input / output ports to optimize data structuring to / from an in-memory computing engine, to maximize data locality across MVM dimensionalities and tensor depth / pixel indices. The OCN 114 routing channels can include bidirectional wire pairs as shown by the duo-directional pipelined routing structure 131 in the expanded view if the OCN 114, so as to ease repeater / pipeline-FF insertion, while providing sufficient density.

[0061] The in-memory computing architecture 100 can be used to implement a neural network (NN) accelerator, wherein a plurality of compute in memory units (CIMUs 112) are arrayed and interconnected using a very flexible on-chip network (OCN 114) wherein the outputs of one CIMU can be connected to or flow to the inputs of another CIMU or to multiple other CIMUs, the outputs of many CIMUs can be connected to the inputs of one CIMU, the outputs of one CIMU can be connected to the inputs of another CIMU and so on. The OCN 114 can be implemented as a single on-chip network, as a plurality of on-chip network portions, or as a combination of on-chip and off-chip network portions.

[0062] The CIMUs 112 can be surrounded by an on-chip network 114 for moving activations between CIMUs 112 (activation network) as well as moving weights from embedded L2 memory to CIMUs 112 (weight-loading interface). This has similarities with architectures used for coarse-grained reconfigurable arrays (CGRAs), but with cores providing high-efficiency MVM and element-wise computations targeted for neural network acceleration. Various options exist for implementing the on-chip network. The approach in Figure 1 enables routing segments along a CIMU 112 to take outputs from that CIMU 112 and / or to provide inputs to that CIMU 112. In this manner data originating from any CIMU 112 can be routed to any CIMU 112, and any number of CIMUs 112.

[0063] Each CIMU 112 is associated with an input buffer (not shown) for receiving computational data from the on-chip network and composing the received computational data into an input vector for matrix vector multiplication (MVM) processing by the CIMU to generate thereby computed data includingincludingincluding an output vector.

[0064] Each CIMU 112 is associated with a shortcut buffer (not shown), for receiving computational data from the on-chip network 114, imparting a temporal delay to the received computational data, and forwarding delayed computation data toward a next CIMU 112 or an output in accordance with a dataflow map such that dataflow alignment across multiple CIMUs 112 is maintained. At least some of the input buffers can be configured to impart a temporal delay to computational data received from the on-chip network 114 or from a shortcut buffer. The dataflow map can support pixel-level pipelining to provide pipeline latency matching.

[0065] The temporal delay imparted by a shortcut or input buffers includes at least one of an absolute temporal delay, a predetermined temporal delay, a temporal delay determined with respect to a size of input computational data, a temporal delay determined with respect to an expected computational time of the CIMU 112, a control signal received from a dataflow controller, a control signal received from another CIMU 112, and a control signal generated by the CIMU 112 in response to the occurrence of an event within the CIMU. In some aspects, at least one of the input buffer and shortcut buffers of each of the plurality of CIMUs 112 in the array of CIMUs 112 can be configured in accordance with a dataflow map supporting pixel-level pipelining to provide pipeline latency matching. The array of CIMUs 112 can also include parallelized computation hardware configured for processing input data received from at least one of respective input and shortcut buffers.

[0066] A least a subset of the CIMUs 112 can be associated with on-chip network 114portions including operand loading network portions configured in accordance with a dataflow of an application mapped onto the IMC. The application mapped onto the IMC includes a neural network (NN) mapped onto the IMC such that parallel output computed data of configured CIMUs executing at a given layer are provided to configured CIMUs 112 executing at a next layer, said parallel output computed data forming respective NN feature-map pixels.

[0067] The input buffer can be configured for transferring input NN feature-map data to parallelized computation hardware within the CIMU in accordance with a selected stride step, such as discussed above. The NN can include a convolution neural network (CNN), and the input buffer can be used to buffer a number of rows of an input feature map corresponding to a size or height of the CNN kernel.

[0068] The CIM array 118 in each CIMU 112 can perform matrix vector multiplication (MVM) in accordance with a bit-parallel, bit-serial (BPB S) computing process in which single bit computations are performed using an iterative barrel shifting with column weighting process, followed by a results accumulation process.

[0069] Figure 2 shows additional details of a portion of the in-memory computing architecture 100 shown in Figure 1, and in particular, details of an example compute inmemory (CIM) array 200 and associated components. The CIM array 200 can be used, for example, to implement, in part, the CIM array 118 discussed above in relation to the in-memory computing architecture 100 shown in Figure 1. In one example implementation, the CIM array 200 can include a fully row / column-parallel (1152 rowX 256 column) array of computing cells 202 of an in-memory-computing (IMC) macro enabling N-bit (5-bit) input processing. The number of rows (1152), the number or columns (256), and the number of bits (5-bit) of input shown in Figure 2 are only examples, and any or all of these quantities and circuit configurations can be varied based on desired implementations. The computing cells can be used to perform computational operations between inputs and data stored in a memory cell within the computing cells. The operations can include logical operations (AND, NOR, etc.) or multiplication operations carried out between inputs. In some examples, the operands of the computation can be 1-bit each. In some other examples, one of the operands can be an analog signal (voltage or current) while the other operand can be a 1-bit operand stored in the memory cell.

[0070] The in-memory computing architecture 100, in some examples, can be utilized for matrix vector multiplication (MVM) operations, which dominate compute-intensive anddata-intensive Al workloads, in a manner that reduces compute energy and data movement by orders of magnitude. This is achieved through efficient analog compute in the computing cells 202, and by thus accessing a compute result (e.g., inner product), rather than individual bits, from memory. But, doing so fundamentally instates an energy / throughput-vs.-SNR tradeoff, where going to analog introduces compute noise and accessing a compute result increases dynamic range (i.e., reducing SNR for given readout architecture). The computing cells 202, which store computational results in the form of a voltage in capacitors within the computing cells 202, can employ metal-fringing capacitors, which can achieve very low noise from analog nonidealities, and thus have the potential for extremely high dynamic range.

[0071] Figure 2 shows a block diagram of the CIM array 200 including a 1152 (row) X 256 (col.) array of 10T (“ten transistor”) SRAM computing cells 202, which in this example are multiplying bit-cells (M-BCs) (such as, for example, a 10T M-BC 202, although the number of transistors of SRAM interface 204 and M-BCs 202 is purely implementation-dependent and the circuit can use different numbers of transistors or other circuit elements without departing from the principles of the disclosure); peripheral circuits for standard writing / reading thereto (e.g., a bit line (BL) decoder 204 and 256 BL drivers 206-1 through 206-256 (collectively referred to as BL drivers 206), a word line (WL) or address decoder 208 and 1152 WL drivers 210-1 through 210-1152 (collectively referred to as WL drivers 210), and control block 212 for controlling the BL decoder such as SRAM interface 204 and the WL decoder 208); peripheral circuitry for providing 5-bit input-vector elements thereto (e.g., 1152 Dynamic-Range Doubling (DRD) DACs 214-1 through 214-1152 (collectively referred to as DRD DACs 214), and a corresponding inmemory computing (IMC) controller (“IMC control block”) 216); peripheral circuitry for digitizing the compute result from each column (e.g., 256 8-bit successive approximation register (“SAR”) ADCs 218-1 through 218-256 (collectively referred to as SAR ADCs 218), and column reset mechanisms 220-1 through 220-256 (collectively referred to as column reset mechanisms 220) (e.g., CMOS switches configured to pull the output voltage levels of column compute lines CLs to a reset voltage VRST during a Reset phase of operation, and allow the voltage levels of column compute lines CLs to reflect their respective compute results during an Evaluate phase of operation). For example, the RST switches corresponding to CL1-CL256, or a subset thereof, can close to produce the desired reset voltage VRST during the reset phase. The RST switches can then open during an ensuing evaluation phase, thereby enabling the voltage values at the CLs to reflect thecomputed product.

[0072] In addition, the lower right portion of Figure 2 depicts an example enlarged view of a representative one of the 256 8-bit ADCs, which includes various switch mechanisms ADCRST (Analog-to-Digital Converter Reset), ADCSMP (Analog-to-Digital Converter Sample) and voltage designations VADCRST and VCMPR, the latter voltage designation connected in this example to a positive terminal of the comparator CMPR and the former voltage designation selectively applied via the ADCRST and ADCSMP switch to reset the comparator. The negative terminal of comparator CMPR receives a CL value when the circuit is activated. An output of comparator CMPR is coupled to SAR logic for outputting an 8-bit digital result. It will be appreciated, however, that the implementation details of the above circuits are representative in nature and that variations to the circuits are possible without departing from the scope or spirit of the present disclosure.

[0073] While writing / reading is typically performed row-by-row, MVM operations are typically performed by applying input-vector elements corresponding to neural -network input activations to all rows at once. That is, each DRD DAC 214j, in response to a respective 5-bit input-vector element Xj[4:0], generates a respective differential output signal (lAj / IAbj) which is subjected to a 1-bit multiplication with the stored weights (Aij / Abij) at each computing cell 202j in the corresponding row of computing cells 202, and accumulation through charge-redistribution across computing cells 202 capacitors CM-BC on the compute line (CL) to yield an inner product in each column, which is then digitized via the respective SAR ADCs 218 of each column as noted above.

[0074] Figure 3 shows an example circuit diagram of the computing cells 202 discussed above in relation to Figure 2. The computing cells 202 can include a highly dense structure for achieving weight storage and multiplication, thereby minimizing data-broad- cast distance and control signals within the context of i-row, j -column arrays implemented using such computing cells, such as the 1152 (row) X 256 (col.) CIM array 200 of 10T SRAM multiplying bit cells (M-BCs).

[0075] The exemplary computing cells 202 includes a six-transistor bit cell portion 222 (here, NMOS transistors 226a, 226b, 226e, 226f and PMOS transistors 226c and 226d), a first switch SW1, a second switch SW2, a capacitor C, a word line (WL) 224, a first bit line (BLj) 227, a second bit line (BLbj) 228, and a compute line (CL) 230.

[0076] The six-transistor bit cell portion 222 is depicted as being located in a middle portion of the computing cells 202, and includes six transistors 226a-226f in this example. The 6-transistor bit cell portion 222 can be used for storage, and to read and write data.In one example, the 6-transistor bit cell portion 222 stores the filter weight. In some examples, data is written to the computing cells 202 through the word line (WL) 224, the first bit line (BL) 227, and the second bit line (BLb) 228.

[0077] The computing cells 202 can include a first CMOS switch SW1 and a second CMOS switch SW2. The first switch SW 1 is depicted as being controlled by a first stored signal Aij such that, when closed, the first switch SW1 couples one of the received differential output signals (lA / IAb) provided by the DRD DACs 214, illustratively IA, to a first terminal of the capacitor C. The second switch SW2 is depicted as being controlled by a second stored signal Abij such that, when closed, the second switch SW2 couples the other one of the received differential output signals of the corresponding DRD DACs 214, illustratively lAb, to the first terminal of the capacitor C. The second terminal of the capacitor C is connected to a compute line (CL) 230 via an output port 232 that provides a result of the computation of the computing cell 202. It is noted that in various other examples, the input signals provided to the first and second switches SW 1 and SW2 can include a fixed voltage (e.g., Vaa), ground, or some other voltage level.

[0078] The computing cells 202, including the first SW1 and second SW2 switches, can implement computation on the data stored in the six-transistor bit cell portion 222. The result of a computation is sampled as charge on the capacitor C. According to various implementations, the capacitor C can be positioned above the computing cell 202 and utilize no additional area on the circuit. In some implementations, a logic value of either Vdd or ground is stored on the capacitor C. In other implementations, the voltage stored on the capacitor C can include a positive or negative voltage in accordance with the operation of the first and the second switches SW 1 and SW2, and the output voltage level generated by the corresponding DRD DACs 214 shown in Figure 2.

[0079] Thus, with continued reference to Figure 3, the value that is stored on the capacitor C is highly stable, since the capacitor C value is either driven up to a fixed analog voltage or down to ground. In some examples, the capacitor C is a metal-oxide-metal (MOM) finger capacitor, and in some examples, the capacitor C can be about 0.1 femto-Farhads (fF) to about 10 fF or can be about 1.2 fF. MOM capacitors have very good matching temperature and process characteristics, and thus have highly linear and stable compute operations. Note that other types of logic functions can be implemented using the computing cells 202 by changing the way the transistors 226a-226f and / or the first and the second switches SW1 and SW2 are connected and / or operated during the Reset and Evaluate phases of operation. The six-transistor bit cell portion 222 can be implementedusing different numbers of transistors and can have different architectures. In some examples, the six-transistor bit cell portion 222 can be a SRAM, DRAM, MRAM, or an RRAM.

[0080] Input Reference Voltage Generation

[0081] Figure 4 shows a Reset phase and an Evaluate phase of a voltage circuit performing dynamic-range doubling.

[0082] As described previously with respect to Figure 2, the CIM array 200 can include an array of 10T SRAM computing cells 202, which in this example are multiplying bitcells (M-BCs). The CIM array 200 can include peripheral circuitry for providing 5-bit input-vector elements thereto (e.g., 1152 Dynamic-Range Doubling (“DRD”) DACs 214- 1 through 214-1152 (collectively referred to as DRD DACs 214), and a corresponding controller (“control block”) 216.

[0083] In the example of Figure 2, each DRD DAC 214j, in response to a respective 5- bit input-vector element Xj[4:0], generates a respective differential output signal (lAj / IAbj) which is subjected to a 1 -bit multiplication with the stored weights (Aij / Abij) at each computing cell 202j in the corresponding row of computing cells 202, and accumulation through charge-redistribution across computing cells 202 capacitors on the compute line (CL) to yield an inner product in each column, which is then digitized via the respective SAR ADCs 218 of each column.

[0084] Multiply and accumulate (MAC) operations are performed on an input signal, provided on lA / IAb, and the bit stored in the M-BC by using input activation data to determine how the complementary signals on lA / IAb transition between the Reset and Evaluate phases. The M-BC stored data is used to determine which of lA / IAb signals is selected to drive the capacitor bottom plate, to implement multiplication logic, with the output product represented by the voltage change on its capacitor’s bottom plate. Accumulation is then performed by the Evaluate phase charge redistribution across connected M-BC capacitors in the column.

[0085] lA / IAb correspond to multi-level analog voltage signals that are to be transmitted through the switches within the M-BC. This provides an option of providing input data through digital -to-analog converters to enable multi -bit input data multiplication with 1- bit stored data. To support up to 4-bit activation inputs, voltage transitions corresponding to 24(16) analog voltage levels are typically required. Generation of accurate analog voltage levels can typically be expensive in terms of power usage and silicon area required. Further, the need for many levels can lead to additional error sources andmismatches.

[0086] As illustrated in Figure 4, to reduce the number of levels, the in-memory computing architecture 100 takes advantage of an approach known as Dynamic-Range Doubling (DRD). DRD takes advantage of the fact that MBC computation depends on the voltage transition on the capacitor bottom plate between the Reset and Evaluate phases 402 and 404, respectively, and that 16 level transitions between these two phases can be achieved from a total of just 9 analog reference levels (VACT C, VACTi-VACTs). This is performed by using eight positive-going transitions (achieved by starting at the lowest voltage, VACT_C, during Reset 402 and transitioning to a higher voltage, VACTi- VACTs, during Evaluate 404), and eight negative-going transitions (achieved by starting at a higher voltage VACTi-VACTs, during Reset 402 and transitioning to the lowest voltage, VACT C, during Evaluate 404) to make up the 16 levels. This example technique improves DAC energy and area by reducing the number of levels required to 16 transitions for five-bit input data. DRD utilizes a Reset duration 402 to apply 16 transition level input data on IAi or IAbi if, for example the data’s sign bit (which can, but need not, be a most significant bit (MSB)) is positive or negative, respectively, or if the data transitions from positive-to-negative, or negative-to-positive, such as shown in Figure 4. While generally the MSB need not be a sign bit, in some examples, it can be formatted as such by applying a fixed offset to the IMC inputs and / or outputs. Other techniques can be possible for reducing the number of levels otherwise needed for a specified number of bits.

[0087] Accordingly, in the example of Figure 4, the Reset phase 402 can include a first positive-input duration 406 in which the switch corresponding to IAi is closed, the switch corresponding to IAbi is open and the lowest voltage VACT C on IAi is coupled to the bottom of the capacitor. If IAi is VACT C =0, for example, the voltage at the bottom of the capacitor relative to IAi is 0. The Reset phase 402 can also include a negative-input duration 408, which is achieved by starting at a higher voltage (IAN = VACTTXX=I..8) in which the switch corresponding to IAb\ = VCT C is open and the switch corresponding to IAN is closed.

[0088] During example Evaluate duration 410, the input voltages IAi and IAbi are swapped. The voltage on IAi transitions from the starting lowest positive voltage IAi = VACT_C at Reset phase 402 to higher positive input voltages IAi = VACTXX=I..8. Similarly, during example Evaluate duration 412 to negative transitioning voltages, e.g., from IA the voltage IAN transitions from higher voltages IAN = VACTXX=I..8 at Resetduration 408 to the lowest voltage IAN = VACT C at Evaluate duration 412.

[0089] The above use of input drivers with DRD DACs has several advantages. The swapping of the voltages of L and IAbi between the reset and evaluation phases 402 and 404, respectively, can cause compute line (CLj) charge redistribution by driving a voltage change of up to twice the power (VDD) voltage across the M-BC capacitors in a column. As noted in Figure 4, the maximum voltage change is twice the voltage from the DRD DAC. In terms of dynamic range, this phenomenon enables use of the five-bit input activations Xj [4,0] (Figure 2) with half the required DAC levels, and it increases the compute line (CLj) voltage swing to overcome attenuation due to CLj parasitic capacitance. Further, the above configuration has advantages relating to energy and throughput. For example, the M-BC capacitors are driven by lower voltages resulting from lower number of needed DAC levels and the reduction in the maximum DAC level enabled by dynamic range doubling.

[0090] Figure 5 shows a first example reference voltage circuit 500. The first example reference voltage circuit 500 can be utilized to generate the pair of activation voltages lA / IAb provided to the computing cells 202 of the CIM array 200. The reference voltage circuit 500 can be utilized for each row of computing cells 202 of the CIM array 200. For example, referring to Figure 2, the reference voltage circuit 500 can be used to generate each of the lA / IAb voltage pairs for a single row of computing cells 202. Similar reference voltage circuits 500 can be used for generating the lA / IAb voltage pairs for each of the other rows of computing cells 202.

[0091] The first example reference voltage circuit 500 shown in Figure 5 includes two multiplexers, muxl 501 (also referred to as “a first multiplexer”) and mux2 502 (also referred to as “a second multiplexer”). The muxl 501 can include a plurality of first input ports 504 that are coupled with a first set of voltages Vn - VIN. The muxl 501 also can include a first control input port 506 receiving a first control input 508 and a first output port 510 that can output one of the first set of voltages as the first of the pair of activation voltages (lA / IAb) based on the first control signal 508. The mux2 502 can include a plurality of second input ports 505 coupled with a second set of voltages V21 - V2M. The mux2 502 can also include a second control input port 507 that can receive a second control input 509 and a second output port 511 that can output one of the second set of voltages as the second of the pair of activation voltages (lA / IAb) based on the second control signal 509. The muxl 501 and mux2 502 are controlled by a mux controller 503, which is coupled with the muxl 501 and the mux2 502. The mux controller 503 can beconfigured to provide the muxl 501 and the mux2 502 with the first control signal 508 and the second control signal 509, respectively. The mux controller 503 is used to control which input voltages (Vn - VIN) and (V21 - V?M) is selected as IA and lAb, respectively, during Reset and then during Evaluate phase. For example, the mux controller 503 can receive a multi-bit digital signal ACT[K:0] to determine which of input voltages is selected during the Reset and the Evaluate phases to achieve the desired bit transitions in the computing cells 202. In some examples, the first set of voltages (Vn - VIN) can be the same in number and values to the second set of voltages (V21 - V2M). In some other examples, the total number of voltages in the first set of voltages can be different from the number of voltages in the second set of voltages. Further, at least one value of the voltage in the first set of voltages can be different from the values of the second set of voltages. In some instances, each of the first set of voltages and the second set of voltages can be further divided into additional sets of voltages. For example, the first set of voltages or the second set of voltages can include two sets of voltages each. In some such examples, the two sets of voltages can be distinct. The control inputs to the two multiplexers can select input voltages from the multiple sets of voltages within the first and second sets of voltages. Having multiple sets of voltages to select voltages from provides flexibility in selecting the desired IA and lAb during the Reset and Evaluate phases.

[0092] ACT[K:0] determines the total number of lA / IAb transitions required (indicates 2K+1transitions). For N number of input voltages in muxl 501 and M number of input voltages in mux2 502, there can be up N2and M2number of output voltage transitions that can be generated from N+M input voltages on the IA and lAb respectively.

[0093] The mux controller 503 selects a voltage for each of muxl 501 and mux2 502 to apply as activation voltages IA and lAb, respectively in the Reset and Evaluate phases (also referred to durations). The voltages selected as the activation voltage can be selected in any suitable manner based on the input analog voltage transition requested or required. For example, the input analog voltage transition can be represented by the data signal ACT[K:0] input or some other data signal indicating the value of the input analog voltage in the Reset and Evaluate phases. As an example, the digital signal can encode the value of the input analog voltage transition in n bits, which the first example reference voltage circuit 500 can decode to determine the value of the input analog voltage in the Reset and Evaluate phases. In another example, the input analog voltages in the Reset and Evaluate phases can be obtained from a lookup table. A multi-bit digital signal can be used todetermine an index to a look-up table, which can include the values for the IA and lAb for the Reset and Evaluate phases. The controller 503 can then send the appropriate first control signal 508 and second control signal 509 to the muxl 501 and mux2 502 to output the desired values of IA and I Ab at the outputs of the multiplexers. In some instances, the input analog voltage transition on IA (or lAb) can be denoted by the difference between the voltages at IA (or lAb) in the Reset phase and in the Evaluate phase. That is, the difference between the value of IA for the Reset phase and the Evaluate phase is the input analog voltage transition provided to the computing cells 202. A similar operation is performed to obtain the input analog voltage provided by the mux2 502 in the lAb.

[0094] Because the first example reference voltage circuit 500 of Figure 5 includes two multiplexers, a value of lAb for the Reset phase and a value for the Evaluate phase can be determined concurrently with the values for IA. Thus, two input analog voltages, one for IA and one for lAb, can be provided concurrently. To reduce the total number of input voltages needed, some or all of the input to mux2 502 can be made equal to the inputs to muxl 501.

[0095] Figure 6A and Figure 6B show a second example reference voltage circuit 600. The second example reference voltage circuit 600 can be utilized as an alternative to the first example reference voltage circuit 500 discussed in relation to Figure 5 to provide desired voltage levels for IA and lAb during the Reset and the Evaluate phases. Figure 6A shows the portion of the second example reference voltage circuit 600 for generating the voltages for IA and Figure 6B shows the portion of the second example reference voltage circuit 600 for generating the voltages for lAb. The second example reference voltage circuit 600 can receive a first set of voltages Vn - VIN and V21 - V2N (shown in Figure 6A) and a second set of voltages V31 - V3N and V41 - V4N (shown in Figure 6B). In some examples, the voltages Vn - VIN can be different from the voltages V21 - V2N and can be viewed as distinct set of voltages. Similarly, voltages V31 - VIN can be different from the voltages V41 - V4N and can be viewed as distinct set of voltages. In some other examples, two or more of the sets Vn - VIN, V21 - V2N, V31 - V3N, and V41 - V4N, can have at least one voltage of the same value. The second example reference voltage circuit 600 includes a first input multiplexer 602, a second input multiplexer 604, and a first output multiplexer 606. Outputs of the first input multiplexer 602 and the second input multiplexer 604 are provided as inputs to the first output multiplexer 606. The second example reference voltage circuit 600 also includes a third input multiplexer608, a fourth input multiplexer 610, and a second output multiplexer 612. Outputs of the third input multiplexer 608 and the fourth input multiplexer 610 are provided as inputs to the second output multiplexer 612. The first set of voltages Vn - VIN and V21 - V2N can be provided as inputs to the first input multiplexer 602 and the second input multiplexer 604, respectively. The second set of voltages V31 - V3N and V41 - V4N can be provided as inputs to the third input multiplexer 608 and the fourth input multiplexer 610, respectively. The output of the first output multiplexer 606 can be provided as the input analog voltage IA and the output of the second output multiplexer 612 can be provided as input analog voltage lAb.

[0096] The second example reference voltage circuit 600 also includes control logic for controlling the operation of the multiplexers. For example, the ACT[K-1 :0] digital signal can be provided as a control input to the first input multiplexer 602, the second input multiplexer 604, the third input multiplexer 608, and the fourth input multiplexer 610. In particular, the ACT[K-1 :0] digital signal can indicate the value of the desired input analog voltage and can control which ones of the set of inputs are to be selected as outputs at the input multiplexers and provided to the respective output multiplexers. In some instances, the ACT[K:0] can be a result of a look up operation from a look up table based on a desired input analog voltage transition.

[0097] An XOR gate 614 receives as input a Reset signal and the MSB ACT[K] of the ACT[K:0] digital signal and outputs a control signal to the first output multiplexer 606. For example, referring to Figure 6A, the ACT[K-l :0] digital signal can select a voltage of a first value from the first set of voltages to be output by the first input multiplexer 602 and can select a voltage of a third value from the first set of voltages to be output by the second input multiplexer 604. Thus, during the Reset duration, the voltage of the third value output by the second input multiplexer 604 can be provided as the IA, and during the Evaluate duration, the voltage of the first value output by the first input multiplexer 602 can be provided as the IA. Of course, depending on polarity of the Reset signal, the first value from the first set of voltages provided by the first input multiplexer 602 can be applied during the Reset duration and the third value from the first set of voltages can be provided by the second input multiplexer 604 during the Evaluate duration instead. Similarly, referring to Figure 6B, second and the fourth values of voltages from the second set of voltages can be provided by the third input multiplexer 608 and the fourth input multiplexer 610, where the fourth value is output during the Reset duration and the second value is output during the Evaluate duration.

[0098] In some instances, the first set of voltages Vn - VIN and V21 - V2N can be the same as the second set of voltages V31 - V3N and V41 - V4N. In some examples, the total number of distinct voltages among the first set of voltages and the second set of voltages is less than the total number of possible digital codes, which is based on a multi-bit digital representation of the input analog voltage. Thus, for example, if the total number of discrete values of the input analog voltage is 16, represented by a 4-bit digital representation, the total number of digital codes is equal to 16. The first example reference voltage circuit 500 and the second example reference voltage circuit 600 can allow the total number distinct voltages in the first set of voltages and the second set of voltages to be less than 16. It should be noted that the input analog voltage provided by the first example reference voltage circuit 500 and the second example reference voltage circuit 600 is the difference between the voltages at each of the pair of activation voltages during the Reset duration and the Evaluate duration. Thus, for example, for IA, the input analog voltage is the difference between the IA provided during the Reset duration and the IA provided during the Evaluate duration.

[0099] The first example reference voltage circuit 500 and the second example reference voltage circuit 600 are only examples, and that other circuits can also be designed to provide the activation voltages. It should be noted that the first example reference voltage circuit 500 and the second example reference voltage circuit 600 provide the ability to select values for the first of the pair of activation voltages independently of selecting the values for the second of the pair of activation voltages during both Reset and Evaluate durations. This ability allows for a reduction in the total number of unique voltages in the first set of voltages and the second set of voltages. The ability of selecting the activation voltages independently of each other can allow various multiplication operations at the computing cell.

[0100] Figure 7 shows a first table 700 of a first example set of activation voltages for a desired set of input analog voltages. In particular, the first table 700 depicts a scenario where the pair of activation voltages (IA and lAb) can provide a + / - 1 multiplication operation at the computing cells using nine unique voltages for the first set of voltages and the second set of voltages. The first table 700 includes columns listing values for IA and lAb during Reset and Evaluate durations, as well as columns for the corresponding values of the voltage at the capacitor C terminal coupled with the column line CL based on the stored bit value in the six -transistor bit cell portion 222 of the computing cell. In this specific example, the pair of activation voltages switch between the same two valuesfor IA and lAb. For example, referring to the first row of the first table 700, during Reset duration the IA has a value of 100 mV and the lAb has a value of 500 mV. During the Evaluate duration, which follows the Reset duration, the voltages on IA and lAb are flipped. That is, during the Evaluate phase, the IA has a value of 500 mV and the lAb has the value of 100 mV. As discussed in relation to Figure 3 and Figure 4, the stored bit (weight) in the computing cell is multiplied by the activation voltage based on the value of the stored bit. For example, if the stored bit is “1” then the switch SW1 is closed and the transition of the IA from 100 mV to 500 mV will result in the voltage of 400 mV at the terminal of the capacitor C that is coupled with the column line CL. Similarly, if the stored bit is “0” then the switch SW2 is closed and the transition of the lAb from 500 mV to 100 mV will result in the voltage of -400 mV at the terminal of the capacitor C coupled with the column line CL. As the polarity of the voltages at the capacitor terminal is positive for when “1” is stored in the computing cell and negative for when “0” is stored in the computing cell, the computing cell operation provides a + / - 1 multiplication.

[0101] Figure 8 shows a second table 800 of a second example set of activation voltages for a desired set of input analog voltages. In particular, the second table 800 is similar to the first table 700, except that the second table 800 generates the values at the terminal of the capacitor based on only 6 unique first and second set of input voltage values (as opposed to 9 unique first and second set of input voltage values in the first table 700). As discussed herein, reducing the total number of unique voltages needed to operate the computing cell can reduce total power consumed by the in-memory computing architecture 100. The approach in table 800 can be extended, to enable various other values at the terminal of the capacitor, based on other combinations of activation voltages.

[0102] Figure 9 shows a third table 900 of a third example set of activation voltages for a desired set of input analog voltages. In particular, the third table 900 depicts a 1 / 0 multiplication by the computing cell corresponding to the stored bit. For example, if a “1” is stored in the computing cell, the capacitor terminal voltage is the difference between the IA in the Reset and the Evaluate durations (which is non-zero). If, for example, a “0” is stored in the computing cell, the capacitor terminal voltage is the difference between the lAb in the Reset and the Evaluate durations (which is always zero because the lAb selected during the Reset and Evaluate durations is the same).

[0103] Figure 10 shows a fourth table 1000 of a fourth example set of activation voltages for a desired set of input analog voltages. In particular, the fourth table 1000 depicts a +1 / -2 multiplication by the computing cell corresponding to the stored bit value of“l’7“0”. For example, referring to the first row, when a “1” is stored in the computing cells, the result is a positive voltage of 400 mV, which is the difference between the voltages 500 mV and 100 mV corresponding to the IA values during Evaluate and Reset durations. If the stored value is “0”, the capacitor terminal voltage is the difference (-800 mV) in the lAb values of 900 mV and 100 mV during the Reset and Evaluate durations, respectively. As the magnitude of the capacitor voltage when the stored value is “0” is twice the magnitude of the capacitor terminal voltage when the stored value is “1”, the computing cell provides a +1 / -2 multiplication.

[0104] The described tables provide only a few examples of the various levels of multiplication factors that can be implemented using the first example reference voltage circuit 500 and the second example reference voltage circuit 600 discussed herein in relation to Figure 5 and Figure 6A and Figure 6B. The ability of select the voltages for IA and lAb independently of each other, enables the implementation of various possible levels of multiplication factors with the same multiplying bit cell and / or a reduced number of unique voltage levels for the first set and the second set of input voltages.

[0105] In each of the tables discussed herein in relation to FIGS. 7-10, the capacitor voltage columns include values that are a difference between the IA or lAb in the Reset and the Evaluate durations. 16 different values corresponding to a 4 bit digital code representing an input analog voltage are possible. The series of values in each capacitor voltage column has a constant step change. For example, referring to Figure 10, the capacitor voltage values under “1” stored in memory have constant step change of 50 mV from one row to the next. The constant step change ensures a linear transition between input analog voltage values over the entire range. Further, there can be a threshold value below which the values of IA and lAb cannot be set. For example, referring to fourth table 1000 in Figure 10, the threshold value can be a value above 100 mV. The voltage generators and drivers providing the first set of voltages and the second set of voltages can have difficulties generating lower voltages accurately and reliably as the analog circuits used to generate the reference voltages typically need some amount of headroom themselves. In such instances, the values utilized for IA and lAb during the Reset and the Evaluate durations are ensured to be above the threshold value that the voltage generators or drivers can reliably provide.ASPECTS OF THE DISCLOSURE

[0106] The present disclosure will be better understood upon reading the following numbered aspects, which should not be confused with the claims. Each of the numberedaspects described below can, in some instances, be combined with aspects described elsewhere in the disclosure. The following listing of example aspects is supported by the disclosure provided herein.

[0107] Aspect 1. An in-memory computing architecture, including: a compute-inmemory (CIM) array of computing cells, the CIM array including a plurality of rows and a plurality of columns, each computing cell including a memory cell, a computing circuitry, and a capacitor for processing and storing a result of computation, wherein the computing cell receives a pair of activation voltages that are representative of an input analog voltage to be multiplied with data stored in the memory cell; and a reference voltage circuit receiving a first set of voltages and a second set of voltages, the reference voltage circuit configured to: during a Reset duration, select a voltage of a first value from the first set of voltages as a first of the pair of activation voltages and select a voltage of a second value from the second set of voltages as a second of the pair of activation voltages, and during an Evaluate duration subsequent to the Reset duration, select a voltage of a third value from the first set of voltages as the first of the pair of activation voltages and select a voltage of a fourth value from the second set of voltages as the second of the pair of activation voltages.

[0108] Aspect 2. The in-memory computing architecture of any one of Aspects 1-8, wherein the first set of voltages is same as the second set of voltages.

[0109] Aspect 3. The in-memory computing architecture of any one of Aspects 1-8, wherein the reference voltage circuit is configured to select voltages for the first of the pair of activation voltages and the second of the pair of activation voltages based on data or signal indicating a value of the input analog voltage.

[0110] Aspect 4. The in-memory computing architecture of any one of Aspects 1-8, wherein the reference voltage circuit is configured to select voltages for the first of the pair of activation voltages and the second of the pair of activation voltages based on data or signal indicating a value of the input analog voltage and a look-up table including a list of input analog voltages and corresponding values of pairs of voltages.

[0111] Aspect 5. The in-memory computing architecture of any one of Aspects 1-8, wherein the reference voltage circuit includes: a first multiplexer including a plurality of first input ports coupled with the first set of voltages, a first control input port receiving a first control input, and a first output port outputting one of the first set of voltages as the first of the pair of activation voltages based on the first control input, a second multiplexer including a plurality of second input ports coupled with the second set of voltages, asecond control input port receiving a second control input, and a second output port outputting one of the second set of voltages as the second of the pair of activation voltages based on the second control input, and a controller coupled with the first multiplexer and the second multiplexer, the controller configured to provide the first multiplexer and the second multiplexer with the first control input and the second control input, respectively, based on a multi -bit digital signal.

[0112] Aspect 6. The in-memory computing architecture of any one of Aspects 1-8, wherein the controller uses the multi-bit digital signal and a look-up table.

[0113] Aspect 7. The in-memory computing architecture of any one of Aspects 1-8, wherein a total number of distinct voltages among the first set of voltages and the second set of voltages is less than a total number of possible digital codes, which is based on a multi-bit digital representation of the input analog voltage.

[0114] Aspect 8. The in-memory computing architecture of any one of Aspects 1-7, wherein any two adjacent input analog voltages of all possible input analog voltages are separated by a same value.

[0115] Aspect 9. A method for providing analog voltages to in-memory computing architecture including a compute in-memory (CIM) array of computing cells, the CIM array including a plurality of rows and a plurality of columns, each computing cell including a memory cell, a computing circuitry, and a capacitor for processing and storing a result of computation, wherein the computing cell receives a pair of activation voltages that are representative of an input analog voltage to be multiplied with data stored in the memory cell, and a reference voltage circuit receiving a first set of voltages and a second set of voltages, the method including: during a Reset duration, selecting a voltage of a first value from the first set of voltages as a first of a pair of activation voltages and select a voltage of a second value from the second set of voltages as a second of the pair of activation voltages; and during an Evaluate duration subsequent to the Reset duration, selecting a voltage of a third value from the first set of voltages as the first of the pair of activation voltages and select a voltage of a fourth value from the second set of voltages as the second of the pair of activation voltages.

[0116] Aspect 10. The method of any one of Aspects 9-16, wherein the first set of voltages are the same as the second set of voltages.

[0117] Aspect 11. The method of any one of Aspects 9-16, including: selecting voltages for the first of the pair of the pair of activation voltages and the second of the pair of activation voltages based on data or signal indicating a value of the input analogvoltage.

[0118] Aspect 12. The method of any one of Aspects 9-16, including: selecting voltages for the first of the pair of activation voltages and the second of the pair of activation voltages based on data or signal indicating a value of the input analog voltage and a look-up table including a list of input analog voltages and corresponding values of pairs of voltages.

[0119] Aspect 13. The method of any one of Aspects 9-16, wherein the reference voltage circuit includes: a first multiplexer including a plurality of first input ports coupled with the first set of voltages, a first control input port receiving a first control input, and a first output port outputting one of the first set of voltages as the first of the pair of activation voltages based on the first control input, a second multiplexer including a plurality of second input ports coupled with the second set of voltages, a second control input port receiving a second control input, and a second output port outputting one of the second set of voltages as the second of the pair of activation voltages based on the second control input, and a controller coupled with the first multiplexer and the second multiplexer, the controller configured to provide the first multiplexer and the second multiplexer with the first control input and the second control input, respectively, based on a multi -bit digital signal.

[0120] Aspect 14. The method of any one of Aspects 9-16, including: the controller using the multi-bit digital signal and a look-up table to provide first multiplexer and the second multiplexer with the first control input and the second control input, respectively.

[0121] Aspect 15. The method of any one of Aspects 9-16, wherein a total number of distinct voltages among the first set of voltages and the second set of voltages is less than a total number of possible digital codes, which is based on a multi-bit digital representation of the input analog voltage.

[0122] Aspect 16. The method of any one of Aspects 9-15, wherein any two adjacent input analog voltages of all possible input analog voltages are separated by a same value.

[0123] Aspect 17. A system, including: a compute-in-memory (CIM) array of computing cells organized as a plurality of rows and a plurality of columns, each computing cell including a memory cell for storing data, computing circuitry for processing a computation, and a capacitor for storing a result of the computation, wherein the computing cell receives a pair of activation voltages that represent an input analog voltage to be multiplied with the data; and a reference voltage circuit configured to: receive a first set of voltages and a second set of voltages; during a Reset duration, selecta voltage of a first value from the first set of voltages as a first of the pair of activation voltages and select a voltage of a second value from the second set of voltages as a second of the pair of activation voltages, and during an Evaluate duration subsequent to the Reset duration, select a voltage of a third value from the first set of voltages as the first of the pair of activation voltages and select a voltage of a fourth value from the second set of voltages as the second of the pair of activation voltages.

[0124] Aspect 18. The system of any one of Aspects 17-24, wherein the first set of voltages is same as the second set of voltages.

[0125] Aspect 19. The system of any one of Aspects 17-24, wherein the reference voltage circuit is configured to select voltages for the first of the pair of activation voltages and the second of the pair of activation voltages based on data or signal indicating a value of the input analog voltage.

[0126] Aspect 20. The system of any one of Aspects 17-24, wherein the reference voltage circuit is configured to select voltages for the first of the pair of activation voltages and the second of the pair of activation voltages based on data or signal indicating a value of the input analog voltage and a look-up table including a list of input analog voltages and corresponding values of pairs of voltages.

[0127] Aspect 21. The system of any one of Aspects 17-24, wherein the reference voltage circuit includes: a first multiplexer including a plurality of first input ports coupled with the first set of voltages, a first control input port receiving a first control input, and a first output port outputting one of the first set of voltages as the first of the pair of activation voltages based on the first control input, a second multiplexer including a plurality of second input ports coupled with the second set of voltages, a second control input port receiving a second control input, and a second output port outputting one of the second set of voltages as the second of the pair of activation voltages based on the second control input, and a controller coupled with the first multiplexer and the second multiplexer, the controller configured to provide the first multiplexer and the second multiplexer with the first control input and the second control input, respectively, based on a multi -bit digital signal.

[0128] Aspect 22. The system of any one of Aspects 17-24, wherein the controller uses the multi-bit digital signal and a look-up table.

[0129] Aspect 23. The system of any one of Aspects 17-24, wherein a total number of distinct voltages among the first set of voltages and the second set of voltages is less than a total number of possible digital codes, which is based on a multi-bit digitalrepresentation of the input analog voltage.

[0130] Aspect 24. The system of any one of Aspects 17-23, wherein any two adjacent input analog voltages of all possible input analog voltages are separated by a same value.

[0131] Details disclosed with respect to the methods described herein included in one example or aspect can be applied to other examples and aspects. Any aspect of the present disclosure that has been described herein can be disclaimed, i.e., exclude from the claimed subject matter whether by proviso or otherwise.

[0132] Various modifications to the implementations described in this disclosure may be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other implementations without departing from the spirit or scope of this disclosure. Thus, the claims are not intended to be limited to the implementations shown herein, but are to be accorded the widest scope consistent with this disclosure, the principles and the novel features disclosed herein.

Claims

CLAIMSWhat is claimed is:

1. An in-memory computing architecture, comprising: a compute-in-memory (CIM) array of computing cells, the CIM array comprising a plurality of rows and a plurality of columns, each computing cell including a memory cell, a computing circuitry, and a capacitor for processing and storing a result of computation, wherein the computing cell receives a pair of activation voltages that are representative of an input analog voltage to be multiplied with data stored in the memory cell; and a reference voltage circuit receiving a first set of voltages and a second set of voltages, the reference voltage circuit configured to: during a Reset duration, select a voltage of a first value from the first set of voltages as a first of the pair of activation voltages and select a voltage of a second value from the second set of voltages as a second of the pair of activation voltages, and during an Evaluate duration subsequent to the Reset duration, select a voltage of a third value from the first set of voltages as the first of the pair of activation voltages and select a voltage of a fourth value from the second set of voltages as the second of the pair of activation voltages.

2. The in-memory computing architecture of claim 1 , wherein the first set of voltages is same as the second set of voltages.

3. The in-memory computing architecture of claim 1, wherein the reference voltage circuit is configured to select voltages for the first of the pair of activation voltages and the second of the pair of activation voltages based on data or signal indicating a value of the input analog voltage.

4. The in-memory computing architecture of claim 1, wherein the reference voltage circuit is configured to select voltages for the first of the pair of activation voltages and the second of the pair of activation voltages based on data or signal indicating a value of the input analog voltage and a look-up table including a list of input analog voltages and corresponding values of pairs of voltages.

5. The in-memory computing architecture of claim 1, wherein the reference voltage circuit comprises: a first multiplexer including a plurality of first input ports coupled with the first set of voltages, a first control input port receiving a first control input, and a first output port outputting one of the first set of voltages as the first of the pair of activation voltages based on the first control input, a second multiplexer including a plurality of second input ports coupled with the second set of voltages, a second control input port receiving a second control input, and a second output port outputting one of the second set of voltages as the second of the pair of activation voltages based on the second control input, and a controller coupled with the first multiplexer and the second multiplexer, the controller configured to provide the first multiplexer and the second multiplexer with the first control input and the second control input, respectively, based on a multi-bit digital signal.

6. The in-memory computing architecture of claim 5, wherein the controller uses the multi-bit digital signal and a look-up table.

7. The in-memory computing architecture of claim 1, wherein a total number of distinct voltages among the first set of voltages and the second set of voltages is less than a total number of possible digital codes, which is based on a multi-bit digital representation of the input analog voltage.

8. The in-memory computing architecture of claim 7, wherein any two adj acent input analog voltages of all possible input analog voltages are separated by a same value.

9. A method for providing analog voltages to in-memory computing architecture comprising a compute in-memory (CIM) array of computing cells, the CIM array comprising a plurality of rows and a plurality of columns, each computing cell including a memory cell, a computing circuitry, and a capacitor for processing and storing a result of computation, wherein the computing cell receives a pair of activation voltages that are representative of an input analog voltage to be multiplied with data stored in the memorycell, and a reference voltage circuit receiving a first set of voltages and a second set of voltages, the method comprising: during a Reset duration, selecting a voltage of a first value from the first set of voltages as a first of a pair of activation voltages and select a voltage of a second value from the second set of voltages as a second of the pair of activation voltages; and during an Evaluate duration subsequent to the Reset duration, selecting a voltage of a third value from the first set of voltages as the first of the pair of activation voltages and select a voltage of a fourth value from the second set of voltages as the second of the pair of activation voltages.

10. The method of claim 9, wherein the first set of voltages are the same as the second set of voltages.

11. The method of claim 9, comprising: selecting voltages for the first of the pair of the pair of activation voltages and the second of the pair of activation voltages based on data or signal indicating a value of the input analog voltage.

12. The method of claim 9, comprising: selecting voltages for the first of the pair of activation voltages and the second of the pair of activation voltages based on data or signal indicating a value of the input analog voltage and a look-up table including a list of input analog voltages and corresponding values of pairs of voltages.

13. The method of claim 9, wherein the reference voltage circuit includes: a first multiplexer including a plurality of first input ports coupled with the first set of voltages, a first control input port receiving a first control input, and a first output port outputting one of the first set of voltages as the first of the pair of activation voltages based on the first control input, a second multiplexer including a plurality of second input ports coupled with the second set of voltages, a second control input port receiving a second control input, and a second output port outputting one of the second set of voltages as the second of the pair of activation voltages based on the second control input, anda controller coupled with the first multiplexer and the second multiplexer, the controller configured to provide the first multiplexer and the second multiplexer with the first control input and the second control input, respectively, based on a multi-bit digital signal.

14. The method of claim 13, comprising: the controller using the multi-bit digital signal and a look-up table to provide first multiplexer and the second multiplexer with the first control input and the second control input, respectively.

15. The method of claim 9, wherein a total number of distinct voltages among the first set of voltages and the second set of voltages is less than a total number of possible digital codes, which is based on a multi-bit digital representation of the input analog voltage.

16. The method of claim 15, wherein any two adjacent input analog voltages of all possible input analog voltages are separated by a same value.

17. A system, comprising: a compute-in-memory (CIM) array of computing cells organized as a plurality of rows and a plurality of columns, each computing cell including a memory cell for storing data, computing circuitry for processing a computation, and a capacitor for storing a result of the computation, wherein the computing cell receives a pair of activation voltages that represent an input analog voltage to be multiplied with the data; and a reference voltage circuit configured to: receive a first set of voltages and a second set of voltages; during a Reset duration, select a voltage of a first value from the first set of voltages as a first of the pair of activation voltages and select a voltage of a second value from the second set of voltages as a second of the pair of activation voltages, and during an Evaluate duration subsequent to the Reset duration, select a voltage of a third value from the first set of voltages as the first of the pair of activation voltages and select a voltage of a fourth value from the second set of voltages as the second of the pair of activation voltages.

18. The system of claim 17, wherein the first set of voltages is same as the second set of voltages.

19. The system of claim 17, wherein the reference voltage circuit is configured to select voltages for the first of the pair of activation voltages and the second of the pair of activation voltages based on data or signal indicating a value of the input analog voltage.

20. The system of claim 17, wherein the reference voltage circuit is configured to select voltages for the first of the pair of activation voltages and the second of the pair of activation voltages based on data or signal indicating a value of the input analog voltage and a look-up table including a list of input analog voltages and corresponding values of pairs of voltages.

21. The system of claim 17, wherein the reference voltage circuit comprises: a first multiplexer including a plurality of first input ports coupled with the first set of voltages, a first control input port receiving a first control input, and a first output port outputting one of the first set of voltages as the first of the pair of activation voltages based on the first control input, a second multiplexer including a plurality of second input ports coupled with the second set of voltages, a second control input port receiving a second control input, and a second output port outputting one of the second set of voltages as the second of the pair of activation voltages based on the second control input, and a controller coupled with the first multiplexer and the second multiplexer, the controller configured to provide the first multiplexer and the second multiplexer with the first control input and the second control input, respectively, based on a multi-bit digital signal.

22. The system of claim 21 , wherein the controller uses the multi -bit digital signal and a look-up table.

23. The system of claim 17, wherein a total number of distinct voltages among the first set of voltages and the second set of voltages is less than a total number of possible digital codes, which is based on a multi-bit digital representation of the input analog voltage.

24. The system of claim 23, wherein any two adjacent input analog voltages of all possible input analog voltages are separated by a same value.

Citation Information

Patent Citations

  • Digital-to-Analogue Converter and Neuromorphic Circuit Using Such a Converter

    US20130144821A1

  • Input circuitry for analog neural memory in a deep learning artificial neural network

    US20230048411A1

  • Scalable array architecture for in-memory computing

    US20230074229A1