Process-in-memory device

KR103025605B1Active Publication Date: 2026-09-29KYUNGPOOK NAT UNIV IND ACADEMIC COOP FOUND
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
KR1020250207167
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-09-29
Estimated Expiration
2045-12-23

Smart Images

  • Figure 112025145551721-PAT00003_ABST
    Figure 112025145551721-PAT00003_ABST
Patent Text Reader

Abstract

A process-in-memory device according to one embodiment disclosed herein comprises: a memory array including a plurality of memory cells; each of the memory cells comprising: a first inverter and a second inverter mutually cross-coupled; a first access transistor connected between the first inverter and a first bit line; a second access transistor connected between the second inverter and a second bit line; a third PMOS transistor connected between a first PMOS transistor and a first NMOS transistor constituting the first inverter; and a fourth PMOS transistor connected between a second PMOS transistor and a second NMOS transistor constituting the second inverter; wherein the gates of the first access transistor and the second access transistor are connected to a word line, and the gates of the third PMOS transistor and the fourth PMOS transistor are connected to a column-assist line, and a word line booster that supplies a boosting voltage higher than a supply voltage to the word line; and an XNOR sensing amplifier that generates an XNOR operation result based on the voltage levels of the first bit line and the second bit line. and a population counter that receives a plurality of the above XNOR operation results and generates an accumulation count; wherein the process-in-memory device is reconfigurable into a memory mode and a computing mode, and in the computing mode, the column-assist line is set to a first voltage so that the third PMOS transistor and the fourth PMOS transistor are turned off, and a first word line storing first data and a second word line storing second data are simultaneously activated to the boosting voltage, and the XNOR sensing amplifier performs a bitwise XNOR operation between the first data and the second data based on the voltage levels of the first bit line and the second bit line, and the population counter can perform an XNOR-and-accumulate operation by accumulating the bitwise XNOR operation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a process-in-memory device. Background Technology

[0002] In traditional Von Neumann computer architectures, the frequent movement of large amounts of data between memory and processing units results in significant energy consumption and latency. Process-In-Memory (PIM) is an approach that resolves this memory-wall problem by performing operations within memory units, thereby meeting the throughput and energy efficiency requirements of data-intensive computing applications such as speech recognition, image recognition, autonomous driving, and artificial intelligence. PIM designs must be compatible with logic CMOS processes, achieve low latency and high energy efficiency, and provide computations robust to process variations.

[0003] Recently, SRAM-based PIM macros utilize analog circuit approaches, such as bitline discharge operations or charge domain operations using capacitors, to perform Multiply-Accumulate (MAC) or Multiply-Average (MAV) operations essential for Convolutional Neural Network (CNN) processing. While analog PIM macros offer advantages such as high energy efficiency, high throughput, and low latency, they suffer from significant design overhead due to data conversion and analog-specific non-ideality issues caused by process, voltage, and temperature (PVT) variations. Additionally, analog PIM macros are susceptible to noise, which leads to a degradation in computational accuracy.

[0004] Digital-based Computational SRAM (C-SRAM) designs perform accurate bit-by-bit operations on bitlines through multi-wordline activation. Multi-wordline activation is an essential operation for digital C-SRAM designs to achieve high throughput and high energy efficiency. However, when the PIM is implemented with standard 6T SRAM cells, significant read disturbance occurs. Since 6T SRAM shares read and write bitlines, bit cells in the same column flip when the bitline voltage drops to a relatively low value, and this read disturbance significantly reduces the reliability of the C-SRAM design.

[0005] Therefore, a new process-in-memory architecture is required that maintains high accuracy of digital computation while eliminating read interference, and efficiently performs XNOR-and-accumulate operations suitable for AI applications such as binary neural networks through multi-wordline activation. (Patent Document 1) US 10699778 B2 (Patent Document 2) US 11176991 B1 The problem to be solved

[0006] One objective of the embodiments disclosed in this document is to provide a process-in-memory device that eliminates read interference and improves the reliability of memory operation through a digital process-in-memory architecture based on differential crosspoint 8T SRAM cells.

[0007] Another objective of the embodiments disclosed in this document is to provide a process-in-memory device capable of performing both general cache SRAM operations and XNOR-and-accumulate operations for binary neural networks by providing a dual operation mode reconfigurable to memory mode and computing mode.

[0008] Another objective of the embodiments disclosed in this document is to provide a process-in-memory device that solves the problem of loss of computational accuracy due to data conversion overhead and PVT variation in analog designs through a fully digital XNOR-and-accumulate operation.

[0009] Another objective of the embodiments disclosed in this document is to provide a process-in-memory device that provides high energy efficiency and computational accuracy optimized for binary neural network processing by performing bit-by-bit XNOR operations and population count operations through multiple wordline simultaneous activation.

[0010] The technical problems of the embodiments disclosed in this document are not limited to those mentioned above, and other unmentioned technical problems will be clearly understood by those skilled in the art from the description below. means of solving the problem

[0011] A process-in-memory device according to one embodiment disclosed herein comprises: a memory array including a plurality of memory cells; each of the memory cells comprising: a first inverter and a second inverter mutually cross-coupled; a first access transistor connected between the first inverter and a first bit line; a second access transistor connected between the second inverter and a second bit line; a third PMOS transistor connected between a first PMOS transistor and a first NMOS transistor constituting the first inverter; and a fourth PMOS transistor connected between a second PMOS transistor and a second NMOS transistor constituting the second inverter; wherein the gates of the first access transistor and the second access transistor are connected to a word line, and the gates of the third PMOS transistor and the fourth PMOS transistor are connected to a column-assist line, and a word line booster that supplies a boosting voltage higher than a supply voltage to the word line; and an XNOR sensing amplifier that generates an XNOR operation result based on the voltage levels of the first bit line and the second bit line. and a population counter that receives a plurality of the above XNOR operation results and generates an accumulation count; wherein the process-in-memory device is reconfigurable into a memory mode and a computing mode, and in the computing mode, the column-assist line is set to a first voltage so that the third PMOS transistor and the fourth PMOS transistor are turned off, and a first word line storing first data and a second word line storing second data are simultaneously activated to the boosting voltage, and the XNOR sensing amplifier performs a bitwise XNOR operation between the first data and the second data based on the voltage levels of the first bit line and the second bit line, and the population counter can perform an XNOR-and-accumulate operation by accumulating the bitwise XNOR operation results.

[0012] In one embodiment, the third PMOS transistor is connected to a first data node which is one end of the first NMOS transistor, and the fourth PMOS transistor is connected to a second data node which is one end of the second NMOS transistor, and the first data node and the second data node can maintain mutually inverted data.

[0013] In one embodiment, in the memory mode, the column-assist line is raised to the first voltage during a read operation, and the third PMOS transistor and the fourth PMOS transistor are turned off, thereby separating the first inverter and the second inverter from the first bit line and the second bit line, so that read interference can be eliminated.

[0014] In one embodiment, a negative voltage generator that supplies a negative voltage to the column-assist line is further included; and in the memory mode, the column-assist line is lowered to the negative voltage during a write operation, thereby turning on the third PMOS transistor and the fourth PMOS transistor, so that the write capability can be improved.

[0015] In one embodiment, the boosting voltage may be 1.3 times the supply voltage, and the negative voltage may be -0.5 times the supply voltage.

[0016] In one embodiment, the first bit line is charged to one of a ground level, an intermediate level, or a high level according to a combination of the first data and the second data, and the XNOR detection amplifier can generate the XNOR operation result based on the voltage level difference between the first bit line and the second bit line.

[0017] In one embodiment, the population counter has a Wallace-tree structure and includes a plurality of full adders and half adders to compress the plurality of XNOR operation results.

[0018] In one embodiment, the population counter may receive 128 bitwise XNOR operation results to generate an 8-bit cumulative count in the range of 0 to 128.

[0019] In one embodiment, the memory array includes a pair of storage blocks, and in the computing mode, two word lines are simultaneously activated from each of the pair of storage blocks, so that a total of four word lines can be simultaneously activated.

[0020] In one embodiment, the first data is an input activation value of a binary neural network, the second data is a network weight of a binary neural network, and the XNOR-and-accumulate operation may be an operation for binary neural network processing. Effects of the invention

[0021] According to a process-in-memory device according to one embodiment disclosed in this document, the third PMOS transistor and the fourth PMOS transistor are turned off by a column-assist line during a read operation, thereby separating the data node from the bit line and completely eliminating read interference, which allows the data stability of the memory cell to be maintained even when multiple word lines are activated and improves the reliability of the C-SRAM design.

[0022] According to a process-in-memory device according to one embodiment disclosed in this document, reconfiguration between a memory mode and a computing mode is possible by an external signal, allowing for the selective use of a memory mode operating as a general cache SRAM and a computing mode performing XNOR-and-accumulate operations for a binary neural network, thereby supporting both memory operations and neural network operations with a single hardware, which can improve chip area efficiency.

[0023] According to a process-in-memory device according to one embodiment disclosed in this document, an analog-to-digital converter is unnecessary through a fully digital XNOR-and-accumulate operation using an XNOR sensing amplifier and a population counter, thereby eliminating area and power overhead, and providing extremely accurate operation results even with process, voltage, and temperature variations, which can improve the accuracy of binary neural network processing.

[0024] According to a process-in-memory device according to one embodiment disclosed in this document, by simultaneously activating two word lines storing an input active value and a network weight, the bit line is charged to three voltage levels according to the data combination, and a bitwise XNOR operation is performed based on this, thereby improving energy efficiency by more than 13.3 times and improving latency by more than 53 times compared to conventional sequential read and external operation methods.

[0025] According to a process-in-memory device according to one embodiment disclosed in this document, a Wallace-tree structured population counter can complete two 128-input XNOR-and-accumulate operations within one cycle by compressing the results of 128 bitwise XNOR operations to generate an 8-bit accumulation count, thereby achieving an energy efficiency of 16.66 TOPS / W as a next-generation AI hardware accelerator for binary neural network processing.

[0026] The effects according to the present disclosure are not limited to those described above, and other unmentioned effects will be clearly understood by a person skilled in the art from the description below. Brief explanation of the drawing

[0027] FIG. 1 is a circuit diagram of a differential crosspoint 8T SRAM cell (100) used as a memory cell of a process-in-memory device according to one embodiment of the present invention. FIG. 2 is a diagram showing the bias conditions for each operating mode of a differential crosspoint 8T SRAM cell (100) according to one embodiment of the present invention. FIG. 3 is a block diagram of a reconfigurable process-in-memory device according to one embodiment of the present invention. FIG. 4 is a physical and logical configuration diagram of a process-in-memory device according to one embodiment of the present invention. FIG. 5a is a waveform of a read operation in memory mode according to one embodiment of the present invention. FIG. 5b is a waveform of a write operation in memory mode according to one embodiment of the present invention. FIG. 6 is a diagram showing a bitline charging mechanism for XNOR operation according to one embodiment of the present invention. FIG. 7 is a diagram of the XNOR-and-accumulate operation configuration according to one embodiment of the present invention. FIG. 8 is a diagram showing a Wallace-tree structure population counter according to one embodiment of the present invention. Specific details for implementing the invention

[0028] Hereinafter, exemplary embodiments according to the present invention will be described in detail with reference to the contents described in the attached drawings. However, the present invention is not limited or restricted by exemplary embodiments. Unless otherwise defined, all terms used in this specification (including technical and scientific terms) shall be used in a meaning that is commonly understood by those skilled in the art to which this disclosure belongs, but this may vary depending on the intent of those skilled in the art, case law, the emergence of new technology, etc.

[0029] Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise. In certain cases, terms have been selected at the applicant's discretion, and in such cases, their meanings will be described in detail in the relevant explanatory sections. Accordingly, terms used in this disclosure should be defined not merely by their names, but based on their meanings and the content throughout this disclosure.

[0030] Throughout this specification, when a part is described as "comprising" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Furthermore, the singular form used in this specification includes the plural form unless specifically stated otherwise. Additionally, the expression "at least one of a, b, and / or c" as used throughout this specification may encompass 'a alone', 'b alone', 'c alone', 'a and b', 'a and c', 'b and c', or 'a, b, and c all'.

[0031] Meanwhile, terms such as "first and / or second" used in this specification may be used to describe various components, but they are used solely for the purpose of distinguishing one component from another and are not intended to limit the scope to the components referred to by such terms. For example, without departing from the scope of the present invention, the first component may be named the second component, and the second component may also be named the first component.

[0032] Additionally, terms such as “…part,” “…module,” etc., as described in this specification refer to a unit that processes at least one function or operation, which may be implemented in hardware or software, or a combination of hardware and software. Furthermore, embodiments of this disclosure may be represented in this specification by functional block configurations and various processing steps. These functional blocks may be implemented by various numbers of hardware and / or software configurations that execute specific functions. For example, embodiments of this disclosure may employ integrated circuit configurations such as memory, processing, logic, look-up tables, etc., which can execute various functions under the control of one or more microprocessors or other control devices.

[0033] Similar to how the components disclosed herein may be executed as software programs or software elements, embodiments of the present disclosure may be implemented in programming or scripting languages ​​such as C, C++, Java, assembler, etc., including various algorithms implemented as combinations of data structures, processes, routines, or other programming configurations. Functional aspects may be implemented as algorithms executed on one or more processors. Additionally, the present embodiments may employ prior art for at least one of electronic configuration, signal processing, and data processing. Terms such as “mechanism,” “element,” “means,” and “configuration” may be used broadly and are not limited to mechanical and physical configurations. The above terms may include the meaning of a series of software processes (routines) in conjunction with a processor, etc.

[0034] Each block of the process flow diagrams attached to this specification and combinations of the flow diagrams may be executed by computer program instructions. Since these computer program instructions may be loaded into the processor of a general-purpose computer, a computer for special purposes, or other programmable data processing equipment, the instructions executed through the processor of the computer or other programmable data processing equipment create means for performing the functions described in the flow diagram block(s).

[0035] These computer program instructions may be stored in computer-available or computer-readable memory that can be directed toward a computer or other programmable data processing equipment to implement a function in a specific way, and the instructions stored in said computer-available or computer-readable memory may also produce a manufactured item containing instruction means that performs the function described in the flowchart block(s).

[0036] Since computer program instructions can be loaded onto a computer or other programmable data processing equipment, instructions that perform a series of operation steps on the computer or other programmable data processing equipment to create a process executed by the computer can also provide steps for executing the functions described in the flowchart block(s).

[0037] Additionally, each block may represent a module, segment, or part of code containing one or more executable instructions for executing a specified logical function(s). Furthermore, in some alternative execution examples, the functions mentioned in the blocks may occur out of order. For instance, two blocks described in succession may actually be executed substantially simultaneously, or the blocks may be executed in reverse order according to their corresponding functions.

[0038] The "electronic device" or "terminal" mentioned in this specification may be implemented as a computer or portable terminal capable of connecting to a terminal via a network. Here, the computer includes, for example, a notebook, desktop, or laptop equipped with a web browser, and the portable terminal may include, for example, any type of handheld wireless communication device that ensures portability and mobility, such as a communication-based terminal like IMT (International Mobile Telecommunication), CDMA (Code Division Multiple Access), W-CDMA (W-Code Division Multiple Access), or LTE (Long Term Evolution), or a smartphone or tablet PC. Additionally, the "electronic device" or "terminal" mentioned in this specification may also include a processor, memory for storing and executing program data, permanent storage such as a disk drive, a communication port for communicating with an external device, and user interface devices such as a touch panel, a key, or a button.

[0039] In the present disclosure, methods implemented as software modules or algorithms may be stored on a computer-readable recording medium as computer-readable code or program instructions executable on a processor. The computer-readable recording medium may include magnetic storage media (e.g., ROM (read-only memory), RAM (random-access memory), floppy disks, hard disks, etc.) and optical reading media (e.g., CD-ROM, DVD (Digital Versatile Disc)). The computer-readable recording medium may be distributed and executed across networked computer systems.

[0040] Hereinafter, various embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In describing the embodiments, technical details that are well known in the art to which the present invention pertains and are not directly related to the present invention will be omitted. This is to ensure that the essence of the present invention is conveyed more clearly without obscuring it by omitting unnecessary explanations. For the same reason, some components in the accompanying drawings may be exaggerated, omitted, or schematically depicted. Furthermore, the size of each component does not entirely reflect its actual size. Throughout this specification, the same reference numerals may refer to the same or corresponding components.

[0041] FIG. 1 is a circuit diagram of a differential crosspoint 8T SRAM cell (100) used as a memory cell of a process-in-memory device according to one embodiment of the present invention. In one embodiment, the differential crosspoint 8T SRAM cell (100) may include a first inverter (101) and a second inverter (102) that are cross-coupled. The first inverter (101) may include a first PMOS transistor (P1) and a first NMOS transistor (N1), and the second inverter (102) may include a second PMOS transistor (P2) and a second NMOS transistor (N2). For example, the first inverter (101) and the second inverter (102) are cross-coupled to form a bidirectional feedback loop and can maintain bi-stable data when the cell is not accessed.

[0042] In one embodiment, the differential crosspoint 8T SRAM cell (100) may include a third PMOS transistor (P3) connected between a first PMOS transistor (P1) and a first NMOS transistor (N1) that constitute a first inverter (101). According to the embodiment, one end of the third PMOS transistor (P3) may be connected to a first data node (DN) which is one end of the first NMOS transistor (N1), and the other end may be connected to the first PMOS transistor (P1). Similarly, a fourth PMOS transistor (P4) may be connected between a second PMOS transistor (P2) and a second NMOS transistor (N2) that constitute a second inverter (102). For example, one end of the fourth PMOS transistor (P4) may be connected to a second data node ( / DN) which is one end of the second NMOS transistor (N2), and the other end may be connected to the second PMOS transistor (P2). In one embodiment, the first data node (DN) and the second data node ( / DN) can maintain mutually inverted data.

[0043] In one embodiment, the gate terminal of the first access transistor (N3) is connected to the word line (WL) and can be connected between the first bit line (BL) and the first inverter (101). According to an embodiment, the gate terminal of the second access transistor (N4) is connected to the word line (WL) and can be connected between the second bit line ( / BL) and the second inverter (102). For example, the first access transistor (N3) and the second access transistor (N4) can be turned on when the word line (WL) is activated to connect the bit line pair (BL, / BL) and the inside of the cell.

[0044] In one embodiment, the gate terminals of the third PMOS transistor (P3) and the fourth PMOS transistor (P4) may each be connected to a column assist line (CAL, 110). According to the embodiment, the column assist line (CAL, 110) may be adaptively controlled according to the operating mode to determine the turn-on / off state of the third PMOS transistor (P3) and the fourth PMOS transistor (P4). For example, the third PMOS transistor (P3) and the fourth PMOS transistor (P4) may perform the role of ensuring the stability of the data node and eliminating read interference in each operating mode.

[0045] FIG. 2 is a diagram showing bias conditions for different operating modes of a differential crosspoint 8T SRAM cell (100) according to an embodiment of the present invention. In one embodiment, the differential crosspoint 8T SRAM cell (100) may have different bias conditions in standby, read, and write operations.

[0046] According to an embodiment, in standby mode, the bitline pair (BL, / BL) is connected to the supply voltage (VDD), and the wordline (WL) and column-assist line (CAL, 110) can be connected to ground. For example, when the column-assist line (CAL, 110) is in a ground state, the third PMOS transistor (P3) and the fourth PMOS transistor (P4) remain in a turned-on state, and the cross-coupled first inverter (101) and second inverter (102) can maintain bistable data.

[0047] In one embodiment, during a read operation, the column-assist line (CAL, 110) may be raised to a first voltage so that the third PMOS transistor (P3) and the fourth PMOS transistor (P4) are turned off. According to the embodiment, the first voltage may be the supply voltage (VDD). For example, when the column-assist line (CAL, 110) is raised to the supply voltage (VDD), the third PMOS transistor (P3) and the fourth PMOS transistor (P4) are initially turned off, and the bitline pair (BL, / BL) may be discharged to ground. In one embodiment, when the wordline (WL) is raised to the boosting voltage (VPP), one of the bitlines may be charged according to the stored data state. According to the embodiment, the boosting voltage (VPP) may be 1.3 times the supply voltage (VDD).

[0048] For example, when the first data node (DN) stores a logic '0' and the second data node ( / DN) stores a logic '1', the second PMOS transistor (P2) conducts, allowing the read cell current (ICELL) to flow through the second PMOS transistor (P2) and the second access transistor (N4) to the second bit line ( / BL). In one embodiment, since the first data node (DN) and the second data node ( / DN) are separated from the bit line pair (BL, / BL) while reading the cell, the differential crosspoint 8T SRAM cell (100) can eliminate read interference by providing a read mechanism that does not interfere with the data stored internally. According to the embodiment, when the word line (WL) returns to ground and the column-assist line (CAL, 110) returns to ground, positive feedback of the cross-coupled first inverter (101) and second inverter (102) can restore their respective data states.

[0049] In one embodiment, during a write operation, the column-assist line (CAL, 110) may be lowered to a negative voltage (NVGG). According to the embodiment, the negative voltage (NVGG) may be -0.5 times the supply voltage (VDD). For example, to write a logic '0' to the first data node (DN) when the first data node (DN) is initially at the supply voltage (VDD) and the second data node ( / DN) is ground, the first bit line (BL) may be set to ground and the second bit line ( / BL) may be set to the supply voltage (VDD). In one embodiment, when the word line (WL) rises to the boosting voltage (VPP), the first access transistor (N3) may be turned on so that the passing node (PN) transitions from the supply voltage (VDD) to ground. According to the embodiment, since the gates of the third PMOS transistor (P3) and the fourth PMOS transistor (P4) are biased to a negative voltage, the first data node (DN) can be easily discharged from the supply voltage (VDD) to ground, thereby improving the write capability. For example, the drop of the first data node (DN) triggers the second inverter (102), and positive feedback within the cell can change the contents of the memory bit.

[0050] FIG. 3 is a block diagram of a reconfigurable process-in-memory device (200) according to one embodiment of the present invention. In one embodiment, the process-in-memory device (200) may include a memory array (280). According to the embodiment, the memory array (280) may include a plurality of differential crosspoint 8T SRAM cells and may include an XNOR sense amplifier, a bitline precharger, a memory cell, a bitline sense amplifier, a bitline discharger, and a column gate. For example, the memory array (280) may include a pair of 16-kilobit storage blocks, each storage block consisting of 128 columns and 128 rows to provide a total capacity of 32-kilobit.

[0051] In one embodiment, the process-in-memory device (200) may include a wordline booster (230). According to the embodiment, the wordline booster (230) may supply a boosting voltage higher than the supply voltage to the wordline. For example, the wordline booster (230) may increase the channel conductance of the access transistor in read and write operations and enable simultaneous activation of multiple wordlines in computing mode by generating a boosting voltage (VPP) corresponding to 1.3 times the supply voltage (VDD) and supplying it to the wordline selected through the row decoder.

[0052] In one embodiment, the process-in-memory device (200) may include a negative voltage generator (250). According to the embodiment, the negative voltage generator (250) may supply a negative voltage to the column-assist line. For example, the negative voltage generator (250) may generate a negative voltage (NVGG) corresponding to -0.5 times the supply voltage (VDD) and supply it to the column-assist line through a column signal driver, thereby strongly turning on the third PMOS transistor and the fourth PMOS transistor during a write operation to improve write capability.

[0053] In one embodiment, the process-in-memory device (200) may include a population counter (270). According to the embodiment, the population counter (270) may receive a plurality of XNOR operation results to generate an accumulation count. For example, the population counter (270) may receive bit-by-bit XNOR operation results from an XNOR detection amplifier in a memory array (280), perform a population count operation, and provide the result as an output signal (P0 to PK-1) through a population output buffer.

[0054] In one embodiment, the process-in-memory device (200) may be reconfigurable into a memory mode and a computing mode. According to the embodiment, a mode selection buffer receives an external signal (FM) and transmits it to control logic, and the control logic may select one of the two operating modes accordingly. For example, when the external signal (FM) is in a low state, the process-in-memory device (200) may operate in memory mode to perform read and write access as a general cache SRAM. In one embodiment, when the external signal (FM) is in a high state, the process-in-memory device (200) may enter computing mode to perform XNOR-and-accumulate operations for a binary neural network.

[0055] In one embodiment, the process-in-memory device (200) may further include a row address buffer, a column address buffer, a predecoder, control logic, a block signal driver, a row decoder, a column signal driver, a block detection amplifier, a write driver, and a data input / output buffer. According to the embodiment, the row address buffer and the column address buffer receive external address signals (AROW, ACOL), and the predecoder decodes them and transmits them to the row decoder and the column signal driver. For example, the control logic receives external control signals such as write enable ( / WE), chip select ( / CS), and clock (CLK) to generate timing signals according to each operation mode, and the block signal driver transmits them to the memory array (280) and peripheral circuits. In one embodiment, the block detection amplifier and the write driver are connected to selected bit lines through column gates and can exchange 8-bit wide data (I0 ~ IN-1, O0 ~ ON-1) with the data input / output buffer.

[0056] In one embodiment, the process-in-memory device (200) may operate on a single power supply. According to the embodiment, the process-in-memory device (200) receives only the supply voltage (VDD) from an external source and can generate the necessary boosting voltage (VPP) and negative voltage (NVGG) itself through an internal wordline booster (230) and negative voltage generator (250). For example, each operation cycle may automatically start from the positive edge of the clock (CLK).

[0057] FIG. 4 is a physical and logical configuration diagram of a process-in-memory device (200) according to an embodiment of the present invention. In one embodiment, the memory array may include a first storage block (210) and a second storage block (220) arranged in a bidirectional symmetrical arrangement. According to the embodiment, the first storage block (210) and the second storage block (220) each have a capacity of 16-kilobit, and each storage block may be composed of 128 columns and 128 rows. For example, a row decoder may be positioned in the center between the first storage block (210) and the second storage block (220) to control the word lines of both storage blocks.

[0058] In one embodiment, the basic configuration of the process-in-memory device (200) may be 4K-word × 8-bit. According to the embodiment, within each storage block (210, 220), a bitline sensing amplifier, a bitline discharger, a thermal gate, and a thermal signal driver may be placed at the bottom of the array, and an XNOR sensing amplifier and a bitline precharger may be placed at the top of the array. For example, a population counter (270) may be placed above the first storage block (210) and the second storage block (220).

[0059] In one embodiment, a pair of data lines (DL, / DL), which are bidirectional data buses of an 8-bit word, can transmit differential signals between a selected column and a block detection amplifier or write driver. According to the embodiment, the main data lines of the 8-bit word can be shared between adjacent storage blocks in a 2:1 column multiplexing manner. For example, the first storage block (210) and the second storage block (220) can share the main data lines to efficiently utilize the chip area.

[0060] In one embodiment, the process-in-memory device (200) may allow switching of functional modes between memory operation and XNOR-and-accumulate operation. According to an embodiment, when the external signal (FM) is in a low state, the process-in-memory device (200) may operate in memory mode to perform read and write access as a general cache SRAM. For example, in memory mode, each storage block (210, 220) operates independently, and a block detection amplifier may amplify the voltage difference of a bitline pair (BL, / BL) and transmit it to a dataline pair (DL, / DL).

[0061] In one embodiment, when the external signal (FM) is in a high state, the process-in-memory device (200) may enter a computing mode. According to the embodiment, in computing mode, the XNOR detection amplifier (290) may perform bitwise XNOR operations, and the population counter (270) may accumulate the results of the operations. For example, in each cycle, two word lines may be simultaneously activated in the first storage block (210) and the second storage block (220), respectively, so that a total of four rows may be turned on simultaneously. In one embodiment, the XNOR detection amplifier (290) may perform a bitwise XNOR operation between the network weight and the input activation value, and transmit the result to the population counter (270) to perform an accumulation count operation for the binary neural network.

[0062] In one embodiment, each population counter (270) can receive 128 XNOR operation results and generate an 8-bit count as output. According to the embodiment, the process-in-memory device (200) can complete two 128-input XNOR-and-accumulate operations within one cycle. For example, compared to an analog XNOR-and-accumulate implementation, throughput is limited compared to a method that activates all rows at once, but the massive area and power overhead caused by analog-to-digital converters is unnecessary, and it can provide extremely accurate operation results even with process, voltage, and temperature variations.

[0063] FIG. 5a is a waveform diagram of a read operation in memory mode according to an embodiment of the present invention. In one embodiment, the supply voltage (VDD) is 1.0 V and the internal signal waveform may be shown under room temperature conditions. According to the embodiment, in the standby state, the bitline precharger is in the ON state, and the bitline sense amplifier, bitline discharger, and thermal gate may be in the OFF state. For example, the wordline (WL) and column-assist line (CAL, 110) are connected to ground, and the bitline pair (BL, / BL) may be precharged with the supply voltage (VDD).

[0064] In one embodiment, a read operation may be initiated by raising the column-assist line (CAL, 110) to the supply voltage (VDD) after the bitline precharger is deactivated. According to the embodiment, when the column-assist line (CAL, 110) is raised to the first voltage, the third PMOS transistor (P3) and the fourth PMOS transistor (P4) may be turned off. For example, the bitline pair (BL, / BL) becomes floating after being discharged to ground, which causes the first inverter and the second inverter to be disconnected from the bitline pair (BL, / BL), thereby eliminating read interference.

[0065] In one embodiment, when the word line (WL) rises to a boosting voltage (VPP), a read cell current may flow depending on the stored data state. According to the embodiment, when the first data node (DN) is low and the second data node ( / DN) is high, the read cell current (ICELL) may flow to the second bit line ( / BL). For example, the read cell current (ICELL) flows to the second bit line ( / BL) through the second PMOS transistor (P2) and the fourth NMOS transistor (N4), which can raise the voltage level of the second bit line ( / BL) while the first bit line (BL) is kept ground. In one embodiment, conversely, when the second data node ( / DN) is low and the first data node (DN) is high, the read cell current (ICELL) flows to the first bit line (BL), which can raise the voltage level of the first bit line (BL) while the second bit line ( / BL) is kept ground.

[0066] In one embodiment, the bitline sense amplifier may operate by activating the sense amplifier positive power supply (SAP) and sense amplifier negative power supply (SAN) signals. According to the embodiment, the bitline sense amplifier may amplify the voltage difference between the first bitline (BL) and the second bitline ( / BL). For example, after the wordline (WL) returns to ground, the full swing read signal may be transmitted to the dataline pair (DL, / DL) by switching the thermal gate. In one embodiment, the delay time from the wordline (WL) to the block sense amplifier output (BSAOUT) may be 1.45 nanoseconds at a supply voltage of 1.0V.

[0067] In one embodiment, when the word line (WL) and column assist line (CAL, 110) return to ground, positive feedback of the cross-coupled first inverter and second inverter can restore their respective data states. According to the embodiment, the third PMOS transistor (P3) and the fourth PMOS transistor (P4) are turned on again so that the differential crosspoint 8T SRAM cell can return to a stable standby state.

[0068] FIG. 5b is a waveform diagram of a write operation in memory mode according to an embodiment of the present invention. In one embodiment, the supply voltage (VDD) is 1.0V and the internal signal waveform may be illustrated under room temperature conditions. According to the embodiment, the write operation may be initiated after the bitline precharger is deactivated, with the write driver transmitting external data to the data line pair (DL, / DL). For example, when the write driver enable signal (WDEN) is activated, the write driver may set the voltage of the data line pair (DL, / DL) according to the input data.

[0069] In one embodiment, the column assist line (CAL, 110) may be lowered to a negative voltage (NVGG). According to the embodiment, the negative voltage (NVGG) may be -0.5 times the supply voltage (VDD). For example, when the column assist line (CAL, 110) is lowered to a negative voltage (NVGG), the third PMOS transistor (P3) and the fourth PMOS transistor (P4) are strongly turned on, which can improve write capability. In one embodiment, external data may be transferred from the data line pair (DL, / DL) to the bit line pair (BL, / BL) by switching the thermal gate.

[0070] In one embodiment, to write a logic '0' to the first data node (DN) when the first data node (DN) is initially at the supply voltage (VDD) and the second data node ( / DN) is at ground, the first bit line (BL) can be set to ground and the second bit line ( / BL) can be set to the supply voltage (VDD). According to the embodiment, when the word line (WL) is activated at the boosting voltage (VPP), the first access transistor (N3) is turned on so that the passing node (PN) can be transitioned from the supply voltage (VDD) to ground. For example, since the gate of the third PMOS transistor (P3) is biased to a negative voltage, the first data node (DN) can be easily discharged from the supply voltage (VDD) to ground.

[0071] In one embodiment, the falling of the first data node (DN) can trigger the second inverter. According to the embodiment, positive feedback within the differential crosspoint 8T SRAM cell can change the contents of the memory bits. For example, when the first data node (DN) falls to ground, the second PMOS transistor (P2) is turned on and the second NMOS transistor (N2) is turned off, so that the second data node ( / DN) can rise to the supply voltage (VDD). In one embodiment, through this positive feedback mechanism, the states of the first data node (DN) and the second data node ( / DN) can be completely reversed.

[0072] In one embodiment, after the word line (WL) returns to ground and the column gate is deactivated, the column-assist line (CAL, 110) may return to ground. According to the embodiment, the bit line pair (BL, / BL) may be precharged again to the supply voltage (VDD). For example, as the third PMOS transistor (P3) and the fourth PMOS transistor (P4) return to a normal turn-on state, the differential crosspoint 8T SRAM cell may be switched to a standby state that stably maintains the newly written data.

[0073] FIG. 6 is a diagram illustrating a bitline charging mechanism for XNOR operations according to an embodiment of the present invention. In one embodiment, convolutional neural networks and deep learning are widely used for modern AI tasks such as traffic prediction, speech recognition, or image classification, but can consume extensive computing and memory resources. According to the embodiment, in a binary neural network, network weights and input activation values ​​are limited to a single bit (+1 or -1), so that the multiplication-accumulate operation, which is the most dominant operation in deep neural networks, can be replaced by a simple XNOR-and-accumulate operation that performs a population count after bitwise XNOR operations. For example, binary neural networks may be suitable candidates for deep learning hardware implementations due to their extreme energy efficiency.

[0074] In one embodiment, a process-in-memory device based on a differential crosspoint 8T SRAM cell can perform XNOR-and-accumulate operations for a binary neural network. According to the embodiment, computing mode operation can be initiated by raising the column-assist line (CAL, 110) to the supply voltage (VDD) after the bitline precharger is deactivated. For example, the bitline pair (BL, / BL) becomes floating after being discharged to ground, and the third PMOS transistor (P3) and the fourth PMOS transistor (P4) are turned off so that the data node can be disconnected from the bitline.

[0075] In one embodiment, a first word line storing an input active value and a second word line storing a network weight may be activated simultaneously. According to the embodiment, depending on the combination of the first data and the second data, the first bit line (BL) may be charged to a ground level, an intermediate level, or a high level. For example, if the bitwise product of the input active value (Ai) and the network weight (Wi) is 00, the first bit line (BL) may not be charged and may maintain a ground level (VBL0). In one embodiment, if the bitwise product of the input active value (Ai) and the network weight (Wi) is 01 or 10, only one bit may charge the first bit line (BL) to form a bit line voltage of an intermediate level (VBL1). According to the embodiment, if the bitwise product of the input active value (Ai) and the network weight (Wi) is 11, the first bit line (BL) may be charged faster to form a bit line voltage of a high level (VBL2).

[0076] In one embodiment, the second bit line ( / BL) can also be charged by the same principle. According to the embodiment, when the bitwise product of the input active value (Ai) and the network weight (Wi) is 11, the bitwise product of the inverted input active value ( / Ai) and the inverted network weight ( / Wi) becomes 00, so the second bit line ( / BL) is not charged and can maintain a ground level. For example, when the bitwise product of the input active value (Ai) and the network weight (Wi) is 01 or 10, only one bit charges the second bit line ( / BL), and a bit line voltage of an intermediate level (VBL1) can be formed. In one embodiment, when the bitwise product of the input active value (Ai) and the network weight (Wi) is 00, the second bit line ( / BL) is charged faster, and a bit line voltage of a high level (VBL2) can be formed.

[0077] In one embodiment, the XNOR sensing amplifier (290) can generate an XNOR operation result based on the voltage level difference between the first bit line (BL) and the second bit line ( / BL). According to the embodiment, the XNOR sensing amplifier (290) may have a low-skewed dynamic inverter-based structure. For example, if the input active value (Ai) and the network weight (Wi) have the same value (00 or 11), the XNOR operation result may be a logic '1', and if they have different values ​​(01 or 10), the XNOR operation result may be a logic '0'. In one embodiment, the XNOR sensing amplifier (290) can detect the 3-level bit line voltage to perform an accurate bit-by-bit XNOR operation.

[0078] FIG. 7 is a diagram of an XNOR-and-accumulate operation configuration according to an embodiment of the present invention. In one embodiment, a process-in-memory device (200) can perform bit-by-bit XNOR operations and population counts by simultaneously activating two word lines for each operation cycle. According to the embodiment, two word lines are simultaneously activated from each of a pair of storage blocks, so that a total of four word lines can be simultaneously activated. For example, in the first storage block, a first word line storing an input activation value and a second word line storing a network weight are simultaneously activated to a boosting voltage (VPP), and in the second storage block, two word lines can be simultaneously activated in the same way.

[0079] In one embodiment, when the XNOR detection amplifier activation signal (XNEN) is enabled, the XNOR detection amplifier (290) can perform an XNOR operation between the input activation values ​​stored row by row and the network weights. According to the embodiment, the XNOR detection amplifier (290) can generate bitwise XNOR operation results based on the voltage levels of the first bit line (BL) and the second bit line ( / BL). For example, since each storage block has 128 columns, 128 bitwise XNOR operation results can be generated for each storage block.

[0080] In one embodiment, the bitwise XNOR operation result generated from the XNOR sensing amplifier (290) may be transmitted to the population counter (270). According to the embodiment, each population counter (270) may receive 128 XNOR operation results as inputs to generate an accumulation count. For example, the population counter (270) may count the number of logic '1's among the 128 inputs to generate an 8-bit count in the range of 0 to 128 as an output. In one embodiment, the process-in-memory device (200) may complete two 128-input XNOR-and-accumulate operations within one cycle.

[0081] In one embodiment, compared to an analog-based XNOR-and-accumulate implementation, it may have limited throughput rather than a method that activates all rows at once. According to the embodiment, the digital architecture eliminates the massive area and power overhead caused by analog-to-digital converters and can provide accurate computation results even with process, voltage, and temperature variations. For example, a fully digital XNOR-and-accumulate operation can prevent loss of accuracy due to non-ideality inherent to analog and guarantee the high computational accuracy required for binary neural network processing.

[0082] In one embodiment, the entire signal flow within the computing mode operation may proceed sequentially from wordline (WL) activation to bitline charging, XNOR sense amplifier (290) operation, population counter (270) operation, and final result output through the population output buffer. According to the embodiment, the delay time from the wordline to the population buffer output may be 2.98 nanoseconds at a supply voltage of 1.0V. For example, with such short delay time and high energy efficiency, the process-in-memory device (200) can provide performance suitable for real-time AI inference applications.

[0083] FIG. 8 is a diagram showing a population counter (270) of a Wallace-tree structure according to an embodiment of the present invention. In one embodiment, the population counter (270) may be implemented as a digital bit counter that receives a plurality of XNOR operation results and generates an accumulation count. According to the embodiment, the population counter (270) has a Wallace-tree structure and includes a plurality of full adders and half adders to compress a plurality of XNOR operation results. For example, the Wallace-tree structure can generate a population count with minimal delay time by efficiently summing input bits in parallel.

[0084] In one embodiment, the population counter (270) may receive 128 bitwise XNOR operation results to generate an 8-bit cumulative count in the range of 0 to 128. According to the embodiment, the population counter (270) may have a 12-layer tree structure that compresses a 128-bit digital input (XNOR<0:127>) into an 8-bit digital output (PCOUT<0:7>). For example, the population counter (270) may include 120 full adders and 15 half adders to achieve the maximum possible bit reduction.

[0085] In one embodiment, a full adder can receive three input bits and generate a sum and a carry output. According to an embodiment, a half adder can receive two input bits and generate a sum and a carry output. For example, each layer of a Wallace-tree structure can use full adders and half adders to progressively reduce the number of input bits, and the final layer can output an 8-bit binary count value. In one embodiment, two types of full adders can be alternately arranged to balance circuit delay and circuit area. According to an embodiment, a 22-transistor full adder including a driving inverter and a 16-transistor full adder without a driving inverter can be alternately arranged to optimize the performance of the entire circuit.

[0086] In one embodiment, the total power consumption and critical path delay of the population counter (270) may be 0.84 milliwatts and 1.05 nanoseconds, respectively, when performing a 128-bit population count under supply voltage of 1.0 V and room temperature conditions. According to the embodiment, the Wallace-tree structure provides fast computation speeds through parallel processing and can ensure digitally accurate count results. For example, such low latency and power consumption can contribute to the process-in-memory device achieving high throughput and energy efficiency.

[0087] In one embodiment, the internal wiring connections of the Wallace-tree structure may be omitted from the drawing. According to the embodiment, 128 input bits are compressed stepwise through 12 layers, and full adders and half adders are appropriately placed in each layer to reduce the number of bits. For example, in the first layer, multiple full adders reduce the 128 inputs to about 85, and in subsequent layers, the number of bits is gradually reduced to finally produce an 8-bit output.

[0088] The above descriptions are specific embodiments for carrying out the present disclosure. The present disclosure will include not only the embodiments described above, but also embodiments that are simply modified or can be easily modified. Furthermore, the present disclosure will include technologies that can be easily modified and implemented using the embodiments described above. Accordingly, the scope of the present disclosure should not be limited to the embodiments described above, but should be defined by the claims set forth below as well as equivalents to the claims of the present disclosure. Explanation of the symbols

[0089] 100: Differential CrossPoint 8T SRAM Cell 101: First Inverter 102: Second Inverter 110: Column-Assist Line (CAL) 200: Process-in-Memory device 210: 1st storage block 220: Second storage block 230: Wordline Booster 250: Negative voltage generator 270: Population Counter 290: XNOR Detect Amplifier

Claims

Claim 1 A process-in-memory device comprises: a memory array including a plurality of memory cells; each of the memory cells comprising: a first inverter and a second inverter mutually cross-coupled; a first access transistor connected between the first inverter and a first bit line; a second access transistor connected between the second inverter and a second bit line; a third PMOS transistor connected between a first PMOS transistor and a first NMOS transistor constituting the first inverter; and a fourth PMOS transistor connected between a second PMOS transistor and a second NMOS transistor constituting the second inverter; wherein the gates of the first access transistor and the second access transistor are connected to a word line, and the gates of the third PMOS transistor and the fourth PMOS transistor are connected to a column-assist line, and a word line booster that supplies a boosting voltage higher than a supply voltage to the word line; an XNOR sensing amplifier that generates an XNOR operation result based on the voltage levels of the first bit line and the second bit line; and a population counter that receives a plurality of the XNOR operation results and generates an accumulation count. A process-in-memory device comprising, wherein the process-in-memory device is reconfigurable into a memory mode and a computing mode, and in the computing mode, the column-assist line is set to a first voltage so that the third PMOS transistor and the fourth PMOS transistor are turned off, and a first word line storing first data and a second word line storing second data are simultaneously activated to the boosting voltage, and the XNOR sense amplifier performs a bitwise XNOR operation between the first data and the second data based on the voltage levels of the first bit line and the second bit line, and the population counter performs an XNOR-and-accumulate operation by accumulating the results of the bitwise XNOR operation. Claim 2 A process-in-memory device according to claim 1, wherein the third PMOS transistor is connected to a first data node which is one end of the first NMOS transistor, and the fourth PMOS transistor is connected to a second data node which is one end of the second NMOS transistor, and the first data node and the second data node maintain mutually inverted data. Claim 3 A process-in-memory device according to claim 1, wherein in the memory mode, the column-assist line is raised to the first voltage during a read operation so that the third PMOS transistor and the fourth PMOS transistor are turned off, thereby separating the first inverter and the second inverter from the first bit line and the second bit line and eliminating read interference. Claim 4 A process-in-memory device according to claim 1, further comprising a negative voltage generator that supplies a negative voltage to the column-assist line; wherein, in the memory mode, the column-assist line is lowered to the negative voltage during a write operation, thereby turning on the third PMOS transistor and the fourth PMOS transistor, so as to improve write capability. Claim 5 A process-in-memory device according to claim 4, wherein the boosting voltage is 1.3 times the supply voltage and the negative voltage is -0.5 times the supply voltage. Claim 6 A process-in-memory device according to claim 1, wherein the first bit line is charged to one of a ground level, an intermediate level, or a high level according to a combination of the first data and the second data, and the XNOR detection amplifier generates the XNOR operation result based on the voltage level difference between the first bit line and the second bit line. Claim 7 A process-in-memory device according to claim 1, wherein the population counter has a Wallace-tree structure and includes a plurality of full adders and half adders to compress the plurality of XNOR operation results. Claim 8 In claim 7, the population counter receives 128 bitwise XNOR operation results and generates an 8-bit cumulative count in the range of 0 to 128, a process-in-memory device. Claim 9 A process-in-memory device according to claim 1, wherein the memory array comprises a pair of storage blocks, and in the computing mode, two word lines are simultaneously activated from each of the pair of storage blocks, so that a total of four word lines are simultaneously activated. Claim 10 A process-in-memory device according to claim 1, wherein the first data is an input activation value of a binary neural network, the second data is a network weight of a binary neural network, and the XNOR-and-accumulate operation is an operation for processing a binary neural network.

Citation Information

Patent Citations

  • Semiconductor Memory Device

    KR1020160093456A

  • Memory structure for artificial intelligence (AI) applications

    US20200251157A1

  • Realization of neural networks with ternary inputs and binary weights in NAND memory arrays

    US20200311523A1

  • Word line booster cell and memory array

    US20240395317A1