In-memory computing based on bitcell architecture
Patent Information
- Application Number
- CN202211231850.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-27
- Filing Date
- 2022-09-30
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2042-09-30
Smart Images

Figure CN115936081B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates in its entirety to memory arrays, such as those used in learning machines / inference machines (e.g., artificial neural networks (ANN)). Background Technology
[0002] Known applications in computer vision, speech recognition, and signal processing benefit from the use of learning machines / inference machines such as deep convolutional neural networks (DCNNs). DCNNs are computer-based tools that process large amounts of data and adaptively "learn" by incorporating relevant features within neighboring data, making broad predictions, and improving predictions based on reliable conclusions and novel incorporations. DCNNs are arranged in multiple "layers," with different types of predictions performed at each layer.
[0003] For example, if multiple 2D images of a face are provided as input to a DCNN, the DCNN will learn various facial features, such as edges, curves, angles, points, color contrast, bright spots, dark spots, etc. These features are learned in one or more first layers of the DCNN. Then, in one or more second layers, the DCNN will learn various recognizable facial features, such as eyes, eyebrows, forehead, hair, nose, mouth, cheeks, etc.; each feature is distinguishable from all other features. That is, the DCNN learns to recognize and distinguish eyes from eyebrows or any other facial features. Summary of the Invention
[0004] In one embodiment, the in-memory compute memory cell includes a first cell having a latch, a write bit line, and a complementary write bit line, and a second cell having a latch, a write bit line, and a complementary write bit line. The write bit line of the first cell is coupled to the complementary write bit line of the second cell, and the complementary write bit line of the first cell is coupled to the write bit line of the second cell. In one embodiment, the first cell and the second cell are foundry cells.
[0005] In one embodiment, the memory array includes a plurality of bit cells arranged as a set of bit cell rows intersecting with a plurality of columns of the memory array. The memory array also has a plurality of in-memory computing (IMC) cells arranged as a set of IMC cell rows of the memory array intersecting with a plurality of columns of the memory array. Each IMC cell in the memory array includes a first bit cell having a latch, a write bit line, and a complementary write bit line, and a second bit cell having a latch, a write bit line, and a complementary write bit line. In each IMC cell, the write bit line of the first bit cell is coupled to the complementary write bit line of the second bit cell, and the complementary write bit line of the first bit cell is coupled to the write bit line of the second bit cell.
[0006] In one embodiment, the system includes multiple in-memory computing (IMC) memory arrays. Each IMC memory array includes: a plurality of bit cells arranged as a set of bit cell rows intersecting with a plurality of columns of the IMC memory array; and a plurality of in-memory computing (IMC) cells arranged as a set of IMC cell rows intersecting with a plurality of columns of the IMC memory array. An IMC cell has: a first bit cell having a latch, a write bit line, and a complementary write bit line; and a second bit cell having a latch, a write bit line, and a complementary write bit line. In the IMC cell, the write bit line of the first bit cell is coupled to the complementary write bit line of the second bit cell, and the complementary write bit line of the first bit cell is coupled to the write bit line of the second bit cell. The system has an accumulation circuitry system coupled to the columns of the multiple IMC memory arrays.
[0007] In one embodiment, the method includes: storing weight data in multiple rows of an in-memory computing (IMC) memory array, the IMC memory array being arranged as multiple cell rows intersecting with multiple cell columns, the IMC memory array including a set of bit cell rows and a set of IMC cell rows. Each IMC cell in the IMC memory array includes a first cell having a latch, a write bit line, and a complementary write bit line; and a second bit cell having a latch, a write bit line, and a complementary write bit line, wherein the write bit line of the first cell is coupled to the complementary write bit line of the second bit cell, and the complementary write bit line of the first cell is coupled to the write bit line of the second bit cell. Feature data is stored in one or more rows in the set of IMC cell rows. IMC cells in columns of the IMC memory array multiply the feature data stored in the IMC cells with weight data stored in the columns of the IMC cells. In one embodiment, a non-transitory computer-readable medium has content that, in operation, configures a computing system to perform a method. Attached Figure Description
[0008] Non-limiting and non-exhaustive embodiments are described with reference to the following accompanying drawings, wherein the same reference numerals refer to the same parts in the various views unless the context otherwise indicates. The dimensions and relative positions of elements in the drawings are not necessarily drawn to scale. For example, the shapes of various elements are selected, enlarged, and positioned to improve drawing clarity. Specific shapes of the drawn elements have been selected for ease of identification in the drawings. Furthermore, for ease of illustration, some elements known to those skilled in the art are not illustrated in the drawings. One or more embodiments are described below with reference to the accompanying drawings, wherein:
[0009] Figure 1 It is a functional block diagram of an embodiment of an electronic device or system having a processing core and a memory, according to one embodiment;
[0010] Figure 2 The illustration shows a conventional system including a memory array and dedicated computing circuitry, which can be used, for example, to perform computations in a neural network;
[0011] Figure 3 A more detailed illustration shows an example of a typical eight-transistor bit cell;
[0012] Figure 4 An example embodiment of an in-memory computing (IMC) unit is illustrated;
[0013] Figure 5 The diagram shows... Figure 4 The logical equivalent of the IMC unit;
[0014] Figure 6 The illustration shows an embodiment of a memory array that can be used in a neural network to provide IMC tiles for kernel storage and feature buffers and to perform multiplication in an IMC mode.
[0015] Figure 7 The illustration shows what can be used for, for example Figure 6 An embodiment of the shielding circuit system in an embodiment of a memory array;
[0016] Figure 8 , Figure 9 and Figure 10 The diagram illustrates example signals used to control the operation of the IMC memory array in various operating modes;
[0017] Figure 11 An embodiment of a memory array implementing an IMC tile with multiple feature tiles is illustrated;
[0018] Figure 12 The illustration shows an example of a system that uses multiple IMC tiles to perform multiplication-accumulation operations;
[0019] Figure 13A and Figure 13B An additional embodiment of a system employing multiple IMC tiles to perform multiplication-accumulation operations is illustrated; and
[0020] Figure 14 An embodiment of a method for performing IMC operations using an IMC memory array is illustrated. Detailed Implementation
[0021] To provide a thorough understanding of the various disclosed embodiments, certain specific details are illustrated in the following description, together with the accompanying drawings. However, those skilled in the art will recognize that the disclosed embodiments can be practiced without one or more of these specific details, or using other methods, components, devices, materials, etc., in various combinations. In other instances, to avoid unnecessarily obscuring the description of the embodiments, well-known structures or components associated with the environment of this disclosure, including but not limited to interfaces, power supplies, physical component layouts, etc., in a computer memory environment, have not been shown or described. Additionally, the various embodiments may be methods, systems, or devices.
[0022] Throughout the specification, claims, and drawings, unless the context clearly specifies otherwise, the following terms shall have the meaning explicitly associated herein. The term “this document” refers to the specification, claims, and drawings associated with this application. Unless the context clearly specifies otherwise, the phrases “in one embodiment,” “in another embodiment,” “in various embodiments,” “in some embodiments,” “in other embodiments,” and their other variations refer to one or more features, structures, functions, limitations, or characteristics of this disclosure, and are not limited to the same or different embodiments. As used herein, the term “or” is an inclusive “or” operator and is equivalent to the phrases “A or B, or both” or “A or B or C, or any combination thereof,” and lists with additional elements are treated similarly. Unless the context clearly indicates otherwise, the term “based on” is not exclusive and allows for reliance on additional features, functions, aspects, or limitations not described. Additionally, throughout the specification, the meanings of “a,” “an,” and “described” include both singular and plural references.
[0023] Computations performed by DCNNs or other neural networks often involve repetitive computations on large amounts of data. For example, many learning / inference machines compare known information or kernels with unknown data or feature vectors, such as comparing a known group of pixels with a portion of an image. A common comparison is the dot product between the kernel and the feature vector. However, kernel size, feature size, and depth tend to vary across different layers of the neural network. In some cases, dedicated computational circuitry can be used to implement these operations on varying datasets.
[0024] Figure 1 This is a functional block diagram of an embodiment of an electronic device or system 100 to which the described embodiments can be applied. System 100 includes one or more processing cores or circuits 102. Processing core 102 may include, for example, one or more processors, state machines, microprocessors, programmable logic circuits, discrete circuit systems, logic gates, registers, etc., and various combinations thereof. The processing core can control the overall operation of system 100, the execution of application programs by system 100, etc.
[0025] System 100 includes one or more memories, such as one or more volatile and / or non-volatile memories, which can store all or part of the instructions and data related to, for example, control of system 100, applications executed by system 100, and operations. As shown, system 100 includes one or more cache memories 104, one or more main memories 106, and one or more auxiliary memories 108. One or more of the memories 104, 106, and 108 may include a memory array, which can be shared by one or more processes executed by system 100 during operation.
[0026] System 100 may include one or more sensors 120 (e.g., image sensors, audio sensors, accelerometers, pressure sensors, temperature sensors, etc.), one or more interfaces 130 (e.g., wireless communication interfaces, wired communication interfaces, etc.), and other circuitry 150, including antennas, power supplies, etc., as well as a main bus system 170. The main bus system 170 may include one or more data, address, power, and / or control buses coupled to the various components of system 100. System 100 may also include additional bus systems, such as bus system 162 communicatively coupling cache memory 104 and processing core 102, bus system 164 communicatively coupling cache memory 104 and main memory 106, bus system 166 communicatively coupling main memory 106 and processing core 102, and bus system 168 communicatively coupling main memory 106 and auxiliary memory 108.
[0027] System 100 also includes a neural network circuit system 140, as shown in the figure. The neural network circuit system 140 includes one or more in-memory computing (IMC) memory arrays 142, which are referenced below. Figures 4 to 14 The discussion includes multiple IMC memory cells.
[0028] Figure 2This is a conceptual diagram illustrating a conventional system 200 including a memory array 210 and dedicated computing circuitry 230, which can be used, for example, as custom computing tiles to perform computations in a neural network. The memory array 210 includes a plurality of cells 212 arranged in a column-row configuration, where multiple cell rows intersect with multiple cell columns. Each cell can be addressed via a specific column and a specific row (e.g., via read bits and word lines). Details of the functions and components used to access specific memory cells are known to those skilled in the art and are not described herein for the sake of simplicity. The number of cells 212 illustrated in the memory array 210 is for illustrative purposes only, and systems employing embodiments described herein may include more or fewer cells in more or fewer columns and more or fewer rows. Each cell 212 of the memory array 210 stores a single data bit.
[0029] As shown, the output of memory array 210 (e.g., weights for the illustrated convolution operation) is provided to dedicated computing circuitry 230. This dedicated computing circuitry 230 includes, for example, multiplication and accumulation circuitry 232 and a trigger library 234 to store feature data for activation during computation. Such dedicated computing circuitry 230 is bulky, thus requiring significant chip area and potentially consuming substantial power in addition to increased memory utilization issues.
[0030] Figure 3 A more detailed illustration is provided. Figure 2 Example bit cell 212 of memory array 210. The illustrated bit cell 212 includes a first write bit line WBL and a second write bit line WBLB, a write word line WWL, a read word line RWL, and a latch 214.
[0031] For example, when applied to matrix-vector multiplication operations such as those employed in modern deep neural network (DNN) architectures, leveraging dedicated memory structures to improve the energy efficiency and computational density of such structures can facilitate low-cost ANN devices and the introduction of in-memory computation (IMC) in non-von Neumann architectures. Neural network operations may require extreme levels of parallel access, potentially involving access to multiple rows within the same memory instance, which can pose significant challenges to the reliability of bit cell contents. These reliability issues can lead to information loss and reduced accuracy, which, when employing high levels of parallelism, can have a significant impact on the statistical accuracy of neural network inference tasks.
[0032] The inventors have developed a novel IMC cell architecture that can be used in memory arrays instead of conventional bit cells to facilitate in-memory computation. Memory arrays utilizing such an IMC memory cell architecture facilitate IMC, for example, by enabling high levels of access (such as access to multiple columns in a memory instance), while maintaining high levels of reliability and increasing computational density. Such memory arrays can be used as IMC tiles for neural network computation to provide storage and multiplier logic, as well as as general-purpose memory in other operational modes. The novel IMC cell can be based on a specific configuration of foundry bit cells (such as 6, 8, 10, or 12 transistor foundry bit cells), which can provide significant density gains.
[0033] Figure 4 An example architecture of one embodiment of the IMC unit 402 is illustrated, and Figure 5 The diagram shows... Figure 4 The logical equivalent of IMC cell 402 is 402'. IMC cell 402 includes a first cell 404 with a first cell latch 406 and a second cell 408 with a second cell latch 410. The first cell 404 and the second cell 408 are coupled to a first write bit line and a second write bit line of an intersecting bit cell, the second write bit line of which complements the first write bit line. The first cell 404 and the second cell 408 can be foundry bit cells, such as 6, 8, 10, or 12 transistor foundry bit cells. As shown, the first write bit line WBL of the first cell 404 is coupled to the second write bit line WBLB of the second cell 408, and the second write bit line WBLB of the first cell 404 is coupled to the first write bit line WBL of the second cell 408. The feature word line (FWL) is used as a feature write line when the IMC cell operates in IMC mode (e.g., to provide XOR multiplication functionality for binary neural networks), and as a word write line when the IMC cell operates as a standard SRAM cell. The XOR enable line enables the IMC cell to operate in XOR mode. Latches 406 and 410 store the feature data and a supplement (X, Xb) to the feature data, eliminating the need for... Figure 2 The trigger library 234 is required. Weight data and supplementary weight data (w, wb) are provided on the read word lines, which serve as read bit lines for the weight array in IMC mode. The IMC cell provides the feature data and weight data XOR on the line XOR (W, X), eliminating the need for a separate XOR multiplier.
[0034] like Figure 4As shown, the precharge circuit system 411 receives the precharge signal PCHXOR to precharge the XOR (W, X) line and enables latch 412 to store the result of the XOR operation, which can be provided to an adder to perform a multiplication-accumulation operation. Alternatively, the output of the XOR (W, X) line can be directly provided to, for example, an adder. The precharge signal PCHXOR can be managed by a memory array management circuit system (such as...) Figure 1 The memory array management circuit system 160 is generated.
[0035] Some embodiments of the IMC cell can be customized. For example, one embodiment of the IMC cell may employ two 12-transistor bit cells to generate a matched signal (XOR) or a mismatched signal (XNOR).
[0036] Figure 6 An example embodiment of a memory array 610 is illustrated, which can be used to provide kernel storage and feature buffers for neural networks. For example, the memory array 610 can be used as in-memory computation tiles to perform computations in a neural network, such as a 1-bit binary neural network XOR implementation. The memory array 610 includes multiple cells arranged in a column-row configuration, where multiple cell rows intersect with multiple cell columns 630. Figure 2 Unlike embodiments of memory array 210, memory array 610 includes a first set 642 of one or more rows 644 of bit cells 212 and a second set 646 of one or more rows 648 of IMC cells 402. Figure 6 As shown, the logically equivalent IMC unit 402' is shown on the left, and an embodiment of implementing IMC unit 402 is shown in more detail on the right. Figure 6 The arrangement shown facilitates IMCXOR operations using foundry 8T-bit cells based on push rules, which can provide high-density IMC memory cells that can be easily integrated into conventional arrays of SRAMs.
[0037] In the IMC operation mode, bit cells 212 in a first set 642 of one or more rows 644 of bit cells 212 can be configured to store kernel data (e.g., weights), and a second set 646 of one or more rows 648 of IMC cells 402 can be configured as a feature buffer, wherein each IMC cell 402 can be configured as a trigger to store feature data and as a one-bit binary neural network XOR multiplier to XOR the stored feature data with the weights stored in another row of the IMC cell, and can be obtained on the read bit line to provide another XOR input.
[0038] In SRAM operation mode, each cell 212, 402 can be addressed via specific columns and specific rows (e.g., via read bit lines and word lines). Details of the functions and components used to access specific memory cells (e.g., address decoders) are known to those skilled in the art and are not described herein for the sake of simplicity. The examples of row 644 of bit cell 212 and row 648 of IMC cell 402 illustrated in memory array 610 are for illustrative purposes only, and systems employing the embodiments described herein may include more or fewer cells in more or fewer columns and more or fewer rows. For example, embodiments may have two or more rows 648 of IMC cells in array 610 (e.g., two rows of IMC cells (see reference '...')). Figure 11 (e.g., the four rows of the IMC cell). Bit cells 212 in one or more rows 644 of the bit cell can be, for example, 6, 8, 10, or 12 transistor bit cells; for example, one or more rows 644 can be implemented using a conventional SRAM array. The first bit cell 404 and the second bit cell 408 of the IMC cell 402 can be, for example, 6, 8, 10, or 12 transistor bit cells. Bit cells 212 and bit cells 404, 408 can employ different bit cell implementations (e.g., bit cell 212 can be 6 transistor bit cells, while bit cells 404, 408 can be 8 transistor bit cells, etc.).
[0039] As mentioned above, Figure 6 The embodiments facilitate feature data storage and parallel computing using a push-rule-based IMC cell arrangement. This allows for the full integration of IMC cells 402 into the memory array 610 to provide IMC tiles with high cell density. The memory array 610 can be accessed using a streaming interface in a first-in-first-out (FIFO) manner, rather than using the memory-mapped addressing typically used in SRAM. The streaming data interface to the IMC tiles further simplifies the integration of multiple tiles and allows for the addition of efficient data relocation and data mover engines before or after the IMC tiles.
[0040] Figure 7 The illustration shows that, for example, Figure 6 An embodiment of the shielding circuitry system 750 used in the memory array 610 is described. For convenience, refer to... Figure 6 To describe Figure 7Each IMC cell 402 in row 648 may have a corresponding local masking control circuitry system 750. As shown, the local masking control circuitry system 750 has an inverter 752 that receives a local masking control signal Mask and an AND gate 754 that receives a global PCHXOR signal. The AND gate 754 provides a local PCHXOR signal as an output, which is used to precharge the IMC cell 402 to IMC mode. The local masking control circuitry system 750 facilitates masking of specific columns of the memory array by disabling XOR operations for selected columns of IMC cells 402, such as... Figure 6 The memory array 610 is used to mask one or more columns 630. In neural network processing, computationally specific masking can be frequently employed to, for example, provide increased resolution. The local masking control signal Mask can be controlled by the system for a particular application. The local masking control signal can be kernel-specific. In one embodiment, the masking control signal used to mask the inputs of the memory can be reused to mask columns of the memory array at the outputs.
[0041] Figure 8 , Figure 9 and Figure 10 The diagram illustrates example control signals used to control the operation of the IMC memory array in various operating modes, and will be referenced for convenience. Figure 6 The memory array 610 is used for description. PCH is a conventional precharge signal used in SRAMs. The PCH signal can be used to control the operation of the memory array 610 in normal or IMC operation mode. Figure 8 The illustration depicts the force applied to the IMC memory array (such as, during normal memory reads of IMC cell 402, e.g., when the array is not operating in IMC mode). Figure 6 Example control signals for the memory array 610. Figure 9 The illustration shows example control signals applied to the IMC memory array during the writing of feature data to IMC cell 402 or kernel data to bit cell 212 of the array. Figure 10 The illustration shows example control signals applied to the IMC memory array during the reading of the XOR result from IMC cell 402. In compute mode, the PCHOFF pulse can be used to capture or latch the XOR result based on the XOR evaluation delay.
[0042] Figure 11 The diagram illustrates a memory array 1110 for implementing IMC tiles, which have multiple feature tiles implemented using multiple rows of IMC cells. For convenience, reference will be made to... Figure 6 The memory array 610 is used to describe Figure 11The memory array 1110 includes a set 642 of one or more rows 644 of bit cells 212, and, as shown, a set 646 of two rows 644 of IMC cells 402 (for illustration purposes). Figure 11 The logical equivalent 402' is shown in the diagram. IMC cell row 648 can be used as feature tiles, where the selection circuitry or feature data selection bus 1170 is used to select a feature tile from the feature tiles used for a specific IMC calculation. This facilitates the reuse of kernel data (weights) with different feature data. In the case of convolutional layer operations, feature data is provided in a streaming manner to maximize feature reuse. Support for cross-streaming feature data can further improve feature data reuse in convolutional layers. Additional rows 644 of the IMC cell can be used for kernel storage in other IMC and SRAM operation modes (e.g., the four rows 644 of the IMC cell can be configured to provide 0–4 feature tiles, where rows 644 not used as feature tiles can be used as kernel storage rows). This provides a flexible geometry with additional available outputs. In some operation configurations, adder-based accumulation can be used, while in others, passive element-based accumulation (e.g., capacitive accumulation) and various combinations thereof can be used.
[0043] Figure 12 An embodiment of system 1200 is illustrated, which employs multiple n-memory arrays configured as IMC tiles to implement multiplication-accumulation operations in adder-based accumulation. For convenience, reference will be made to... Figure 6 To describe Figure 12 System 1200 includes multiple memory arrays 1210 having N columns and configured as IMC tiles, each memory array being coupled to a corresponding N-bit adder 1280. Each memory array 1210 includes: rows of bit cells (participating...) Figure 6 The set of 642 (in 644) Figure 12 The core is typically implemented using 6 or 8 transistor bit cells; and one or more rows of IMC cells (see [reference]). Figure 6 The set 646 (of which 648 can typically be implemented using 8, 10, or 12 transistor bit cell pairs) is shown in the figure. The n×Log2N adder 1290 provides the final accumulated value and can optionally compensate for system bias. When masking is used to provide sparsity in the accumulation of the output XOR result, Figure 12 The embodiments described may be particularly useful.
[0044] Figure 13A An embodiment of system 1300 is illustrated, which employs a plurality of n memory arrays configured as IMC tiles to implement multiplication-accumulation operations using capacitor-based accumulation. For convenience, reference will be made to... Figure 1 and Figure 6 To describe Figure 13A System 1300 includes multiple IMC memory arrays 610, each IMC memory array 610 being coupled to a corresponding capacitor element 1380 to accumulate the result of an XOR calculation performed by the IMC memory array 610. The capacitor element 1380 is coupled to a matching line 1382 to generate a matching signal Match, which can be provided as an input to an analog-to-digital converter ADC 1384. Match line 1382 is also selectively coupled to bias capacitor elements CbiasP 1386 and CbiasN 1388 via a switch 1387. The switches can be controlled by system 100 to provide a programmable bias. The bias capacitor CbiasP can store a positive bias charge, for example, based on a PCHOFF signal or a delayed PCHOFF signal, while the bias capacitor CbiasN can store a negative bias charge, for example, based on an inverted PCHOFF signal or a delayed version of an inverted PCHOFF signal. Applying a programmable bias to the accumulated value facilitates batch normalization in neural network applications.
[0045] Capacitor elements 1380, 1386, and 1388 may include device-based capacitors (e.g., Nmos, Pmos), metal capacitors, trench capacitors, and various combinations thereof.
[0046] The ADC 1384 also receives a reference voltage Vref, which may correspond to, for example, an n / 2 matching line 1392 bump equivalent. The output of the ADC 1384 indicates the count of the XOR accumulation. The output can be provided to a multi-level analog-to-digital converter 1396 to provide a multi-bit classification output.
[0047] Figure 13B Another embodiment of system 1300' is illustrated, which employs a plurality of n memory arrays configured as IMC tiles to implement multiplication-accumulation operations using capacitor-based accumulation. Figure 13B System 1300' and Figure 13A The difference in system 1300 is that each of the multiple IMC memory arrays 610 also generates complementary XORB results, which are provided to corresponding capacitor elements 1392 to accumulate the results of the XORB calculations performed by the IMC memory arrays 610, thereby generating a mismatch signal Matchb on mismatch line 1394. In addition to the match signal on line 1382, the Matchb signal on line 1394 is provided to successive approximation (SA) circuit 1398. The output of SA circuit 1398 indicates whether the accumulated match exceeds the accumulated mismatch and can be used as a classification signal. The capacitor biasing element can also be similar to the reference... Figure 13A The discussion is coupled to the mismatch line 1394.
[0048] Figure 14 An embodiment of a method 1400 for performing IMC operations using an IMC memory array is illustrated, and for convenience, references are provided. Figure 1 , Figure 6 , Figure 12 The method is described in conjunction with Figure 13. It can be, for example, under the control of the memory management circuitry system 160 of claim 1 and using… Figure 6 The IMC memory array 610 is used to perform this.
[0049] At 1402, method 1400 stores weight data in multiple rows of an in-memory computation (IMC) memory array. For example, when such rows of IMC cells are configured to operate in bit-cell operation mode, the weight data may be stored in one or more rows 644 of the set of bit-cell rows 642, or may be stored in one or more rows 648 of the set of IMC cell rows 646, or various combinations thereof. The method proceeds from 1402 to 1404.
[0050] At 1404, method 1400 stores feature data in one or more rows of the IMC memory array. For example, feature data may be stored in one or more rows 648 of a set of rows 646 of IMC cells, the set of rows 646 being configured to operate in IMC operation mode. Method 1400 proceeds from 1404 to 1406.
[0051] At 1406, method 1400 multiplies the feature data stored in the IMC cells of one or more columns of the IMC memory array with the weight data stored in the corresponding column. For example, IMC cell 402 of column 630 can XOR the feature data stored in the latch of column IMC cell 402 with the weight data stored in other cells of column 630. The multiplication can be repeated for addition columns of IMC array 610, or for different IMC cells of column 630. Method 1400 proceeds from 1406 to 1408.
[0052] At 1408, method 1400 accumulates the results of the multiplication operation. For example, adder 1280 or capacitor 1380 can be used to accumulate the results.
[0053] Figure 14Embodiments of method 1400 may exclude all the actions shown, may include additional actions, may combine actions, and may perform actions in various orders. For example, in some embodiments, the accumulation at 1408 may be omitted; in some embodiments, storing weight data at 1402 may occur after or in parallel with storing feature data at 1404; loops may be employed (e.g., a loop that loads the weight dataset, followed by a loop that loads the feature data and the accumulation result); actions to compensate for bias may be performed; additional actions to generate classification signals may be performed; and various combinations thereof.
[0054] Some embodiments may take the form of or include a computer program product. For example, according to one embodiment, a computer-readable medium is provided, which includes a computer program adapted to perform one or more of the methods or functions described above. The medium may be a physical storage medium, such as a read-only memory (ROM) chip, or a disc such as a digital multifunction disc (DVD-ROM), optical disc (CD-ROM), hard disk, memory, network, or a portable media article that will be read by a suitable drive or via a suitable connection (including one or more barcodes or other related codes stored on one or more such computer-readable media and readable by a suitable reader device).
[0055] Furthermore, in some embodiments, some or all of these methods and / or functions may be implemented or provided in other ways, such as at least in part in firmware and / or hardware, including but not limited to: one or more application-specific integrated circuits (ASICs), digital signal processors, discrete circuit systems, logic gates, standard integrated circuits, controllers (e.g., by executing appropriate instructions, and including microcontrollers and / or embedded controllers), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), and devices employing RFID technology and various combinations thereof.
[0056] The various embodiments described above can be combined to provide further embodiments. These and other changes can be made to the embodiments based on the detailed description above. Generally, the terminology used in the appended claims should not be construed as limiting the claims to the specific embodiments disclosed in the specification and claims, but should be interpreted to include all possible embodiments and the full scope of the authorized equivalents of these claims. Therefore, the claims are not limited by this disclosure.
Claims
1. An in-memory computing memory unit, comprising: The first bit unit has a latch, a write bit line, and a complementary write bit line; as well as The second bit unit has a latch, a write bit line, and a complementary write bit line, wherein the write bit line of the first bit unit is coupled to the complementary write bit line of the second bit unit, and the complementary write bit line of the first bit unit is coupled to the write bit line of the second bit unit, wherein the first bit unit includes a read word line, and the second bit unit includes a read word line, and in the in-memory computation mode, the in-memory computation unit performs an XOR operation on the data stored in the latch and the data provided on the read word line.
2. The in-memory computing memory unit according to claim 1, wherein the first bit unit and the second bit unit are spoofing units.
3. The in-memory computing memory unit according to claim 2, wherein the first bit unit and the second bit unit are eight transistor bit units.
4. The in-memory computing memory unit according to claim 1, wherein the first bit unit includes a read bit line, and the second bit unit includes a read bit line, and in the in-memory computing operation mode, the in-memory computing unit performs an XOR operation on the feature data stored in the latch and the weight data provided on the read word line.
5. A memory array, comprising: The memory array has a plurality of bit cells arranged as a set of bit cell rows of the memory array that intersect with a plurality of columns of the memory array. as well as The memory array has a plurality of in-memory compute (IMC) cells, which are arranged as a set of IMC cell rows of the memory array intersecting with the plurality of columns of the memory array, wherein each IMC cell in the memory array comprises: The first bit unit has a latch, a write bit line, and a complementary write bit line; as well as The second bit unit has a latch, a write bit line, and a complementary write bit line, wherein the write bit line of the first bit unit is coupled to the complementary write bit line of the second bit unit, and the complementary write bit line of the first bit unit is coupled to the write bit line of the second bit unit, wherein the first bit unit of the IMC unit includes a read word line, and the second bit unit of the IMC unit includes a read word line, and in the in-memory computation mode of the array, the IMC unit selectively performs an XOR operation on the data stored in the latch and the data provided on the read word line.
6. The memory array according to claim 5, wherein the plurality of bit cells, the first bit cell of the IMC unit, and the second bit cell of the IMC unit are foundry bit cells.
7. The memory array according to claim 6, wherein, The plurality of bit units are six transistor bit units; and The first bit unit and the second bit unit of the IMC unit are eight transistor bit units.
8. The memory array according to claim 5, wherein the first bit unit of the IMC unit includes a read bit line, and the second bit unit of the IMC unit includes a read bit line, and in the in-memory computation mode of the array, the IMC unit selectively performs an XOR operation on the feature data stored in the latch and the weight data provided on the read word line.
9. The memory array of claim 5, wherein the array includes a precharge circuit system coupled to the plurality of IMC cells.
10. The memory array of claim 9, wherein the precharge circuit system includes a shielding circuit system that selectively shields the outputs of columns of the array during operation.
11. The memory array of claim 5, wherein the set of bit cell rows comprises a plurality of bit cell rows, and the set of IMC cell rows comprises a plurality of IMC cell rows.
12. The memory array of claim 11, comprising a selection circuit system coupled to a set of IMC cell rows, wherein the selection circuit system selects a row among the plurality of IMC cell rows in an operation.
13. The memory array of claim 11, wherein the set of IMC cell rows comprises four IMC cell rows.
14. The memory array of claim 12, wherein, in operation, a single row in the set of IMC cell rows is configured to operate in IMC operation mode or bit cell operation mode.
15. A system comprising: Multiple in-memory compute (IMC) memory arrays, each IMC memory array comprising: A plurality of bit cells, the plurality of bit cells being arranged as a set of bit cell rows intersecting with a plurality of columns of the IMC memory array; and The IMC memory array comprises multiple in-memory compute IMC cells, which are arranged as a set of IMC cell rows intersecting with the multiple columns of the IMC memory array. Each IMC cell in the IMC memory array has: The first bit unit, the first bit unit having a latch, a write bit line, and a complementary write bit line; and A second bit unit, the second bit unit having a latch, a write bit line, and a complementary write bit line, wherein the write bit line of the first bit unit is coupled to the complementary write bit line of the second bit unit, and the complementary write bit line of the first bit unit is coupled to the write bit line of the second bit unit; and An accumulation circuit system coupled to the columns of the plurality of IMC memory arrays. The first bit unit of the IMC unit includes a read word line, and the second bit unit of the IMC unit includes a read word line. In the in-memory computing operation mode of the system, the IMC unit selectively performs an XOR operation on the data stored in the latch and the data provided on the read word line.
16. The system of claim 15, wherein the plurality of bit cells of the plurality of IMC memory arrays, the first bit cell of the IMC cell, and the second bit cell of the IMC cell are foundry bit cells.
17. The system of claim 15, wherein the first bit unit of the IMC unit includes a read bit line, and the second bit unit of the IMC unit includes a read bit line, and in the in-memory computation mode of the system, the IMC unit selectively performs an XOR operation on the feature data stored in the latch and the weight data provided on the read word line.
18. The system of claim 15, wherein the plurality of IMC memory arrays includes a pre-charge circuit system coupled to the plurality of IMC cells.
19. The system of claim 18, wherein the pre-charge circuit system includes a shielding circuit system that selectively shields the outputs of the columns of the array during operation.
20. The system of claim 15, wherein the accumulator circuitry comprises a plurality of adders.
21. The system of claim 15, wherein the accumulation circuit system comprises one or more capacitors.
22. The system of claim 21, comprising one or more bias capacitors capable of being selectively coupled to the accumulator circuitry.
23. The system of claim 21, comprising a readout circuitry system coupled to the one or more capacitors.
24. The system of claim 23, wherein the readout circuitry system includes an analog-to-digital converter.
25. The system of claim 23, wherein the readout circuit system includes a successive approximation circuit.
26. A method comprising: Weight data is stored in multiple rows of an in-memory computation IMC memory array, the IMC memory array being arranged as multiple cell rows intersecting with multiple cell columns, the IMC memory array comprising a set of bit cell rows and a set of IMC cell rows, wherein each IMC cell in the IMC memory array comprises: The first bit unit, the first bit unit having a latch, a write bit line, and a complementary write bit line; and The second bit unit has a latch, a write bit line, and a complementary write bit line, wherein the write bit line of the first bit unit is coupled to the complementary write bit line of the second bit unit, and the complementary write bit line of the first bit unit is coupled to the write bit line of the second bit unit. The feature data is stored in one or more rows of the set of IMC cell rows; and Using the IMC cells of the columns of the IMC memory array, the feature data stored in the IMC cells is multiplied by the weight data stored in the columns of the IMC cells.
27. The method of claim 26, further comprising controlling the operation mode of an individual row in the set of IMC cell rows.
28. The method of claim 26, wherein the multiplication is performed on a set of the plurality of columns, and the method includes accumulating the result of the multiplication operation on the set of columns.
29. The method of claim 28, comprising accumulating the results of multiplication operations of multiple IMC memory arrays.
30. The method of claim 29, further comprising applying a bias to the accumulated result of the multiplication operation.
31. The method of claim 29, further comprising generating a classification signal for a neural network based on the accumulated multiplication results of the plurality of IMC memory arrays.
32. A non-transitory computer-readable medium having content, said content, in operation, configuring a computing system to execute a method, said method comprising: Weight data is stored in multiple rows of an in-memory computation IMC memory array, the IMC memory array being arranged as multiple cell rows intersecting with multiple cell columns, the IMC memory array comprising a set of bit cell rows and a set of IMC cell rows, wherein each IMC cell in the IMC memory array comprises: The first bit unit, the first bit unit having a latch, a write bit line, and a complementary write bit line; and The second bit unit has a latch, a write bit line, and a complementary write bit line, wherein the write bit line of the first bit unit is coupled to the complementary write bit line of the second bit unit, and the complementary write bit line of the first bit unit is coupled to the write bit line of the second bit unit. The feature data is stored in one or more rows of the set of IMC cell rows; and Using the IMC cells of the columns of the IMC memory array, the feature data stored in the IMC cells is multiplied by the weight data stored in the columns of the IMC cells.
33. The non-transitory computer-readable medium of claim 32, wherein the content includes instructions executed by the computing system.
Citation Information
Patent Citations
Multiport memory with twisted bitlines
US6909663B1