Memory addressed using an addressing scheme with minimal addressable unit
Patent Information
- Application Number
- CN202210400904.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-02-08
- Filing Date
- 2022-04-15
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-04-15
AI Technical Summary
快取中的列数可以是最小可寻址单元大小的整数倍,这会导致对页面缓冲器和快取布局的设计限制
[0017]本公开所述的技术利用“非二进制”快取配置来支持高效布局和操作高密度存储器。如本发明所述实施的数据路径配置可克服两个不同快取行存储的一个数据字元(wordof data)的存取模式问题。此外,如本发明所述实施的数据路径配置可克服快取支持以及位于不同列的快取和页面缓冲器的不同字元中的相同位元选择的互连设置的问题。此外,如本发明所述实施的数据路径配置提供了如何设置支持存储器阵列的单元阵列布局的问题解决方式,特别是对于非常大的页面大小。
Smart Images

Figure CN116612799B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of integrated circuit memory technology, to a memory device having a data path configuration between the memory and an interface, and more specifically to a memory that uses an addressing mode having a minimum addressable unit for addressing. Background Technology
[0002] Some types of high-density memory (including NAND flash memory) are configured for page-mode operation, which involves moving relatively large data pages in parallel from the memory array to the page buffer, while the input / output interface may include a smaller number of pins. In some contexts, a page may comprise a sequence of 16 kilobytes (KB) or more, and the input / output interface may include eight pins for data transfer. To facilitate higher throughput and other operations, cache memory may be placed between the page buffer and the input / output interface. Using cache memory, data pages can be quickly transferred from the page buffer to the cache. During the time interval between transferring the next page between the page buffer and the memory array, the input / output interface can utilize cache access for operation.
[0003] In implementations of page buffers and caches with page operations, the layout on the integrated circuit can be a factor. Page buffers and caches are expected to conform to the pitch of multiple data lines that provide data page connections for the memory array. Page buffer and cache units may include clocked storage elements such as latches or flip-flops formed using multiple transistors, and therefore may require a layout pitch larger than the data line pitch. In one approach to address this layout issue, the page buffer and cache are arranged in multiple columns and rows, where each row conforms to an orthogonal pitch equal to the number of columns of multiple data lines. For example, in a layout with 16 columns, rows connect to 16 data lines, store 16 bits, and conform to an orthogonal pitch equal to the 16 data lines.
[0004] The number of columns and cells per row in the cache are typically powers of 2 (e.g., 8 or 16) to facilitate data transfers that are integer multiples of the smallest addressable memory (MIN) size. The base 2 of the address identifies the boundaries of MINs that are integer multiples of the MIN size. The MIN has a size corresponding to the smallest addressable size supported by the memory. The size of the MIN is typically 1 byte or 8 bits. In some devices, the MIN can be 1 word or 16 bits. If the MIN is 8 bits, it will be represented as 1000 in binary. Therefore, the word (8-bit) boundary in a base-2 address will always have three least significant bits equal to 0 (i.e., XXXXXX000). If the MIN is 16 bits, it will be represented as 10000 in binary. Therefore, word (16-bit) boundaries in a radix-2 address will always have four least significant bits equal to 0 (i.e., XXXXX0000). Other sizes of minimum addressable units are also possible. The number of columns in the cache can be an integer multiple of the minimum addressable unit size, which imposes design constraints on the page buffer and cache layout.
[0005] However, the vertical spacing of the page buffer and cache can be important. Therefore, it is desirable to provide a configuration that reduces the area required to implement the page buffer and cache in an integrated circuit device with memory. Summary of the Invention
[0006] This disclosure describes a technique that utilizes an addressing scheme based on a minimum addressable unit (e.g., a byte) to provide a data path for memory addressing, where the minimum addressable unit has a power of 2 size (e.g., 8). The data path is used to transfer data between a memory array and a data interface using multiple transfer storage units having N bits (e.g., 12), where N is a non-integer multiple of the minimum addressable unit size (e.g., 1.5 multiplied by 8, or 0.75 multiplied by 16). In embodiments described herein, N is not a power of 2. This data path configuration can be supported by effectively implementing page buffers and caching.
[0007] For example, the data path may include a page buffer for connecting to the memory array, and a cache for connecting the page buffer to the data interface. The memory array may include multiple data lines for connecting to the page buffer, which is located in X groups of data lines. The cache may include multiple cache unit arrays, each of which is used for each of the X groups of data lines. Each cache unit array can be used to access an integer number of P bits of the page buffer and may include P cache units arranged in N cache rows and P / N cache columns.
[0008] The data path may include multiple circuits coupled to the cache to connect N cache cells of a selected P / N cache row of a selected cache cell array of each of the X cache cell arrays to a set of M interface lines, where M is N multiplied by X. The data interface is sequentially coupled to the M interface lines and includes multiple circuits for transmitting data in D bit units during each interface clock pulse cycle, where D is an integer multiple of the size of the smallest addressable unit (e.g., 8 or 16) using binary addressing.
[0009] In this data path configuration, M is a common multiple of D and N. For example, for the smallest addressable unit of 1 byte (octet), D can be 16 in 8-pin single data rate port mode, 8-pin double data rate port mode, and 16 in 16-pin single data rate port mode, where a single data rate port means that only one piece of data is transmitted per clock pulse.
[0010] The memory may include multiple address translation circuits that generate multiple selection signals in response to an input address to control multiple data path circuits transmitted by N bit transfer memory cells.
[0011] Using the data path described in this invention, a data page based on the addressing mode may include multiple smallest addressable units (e.g., bytes) with sequential addresses. Data pages stored in the memory array are distributed across multiple N-bit memory cells, corresponding to N-bit rows (bit columns) of cache and page buffers, coupled to data lines that are set non-sequentially across X groups. Sequential transfer storage units are interleaved between different cache cell arrays in X groups.
[0012] This disclosure describes a memory including a memory array having multiple data lines for accessing multiple memory cells of the memory array. The multiple data lines comprise X groups of data lines, where X is an integer greater than 1. The memory includes a data path that includes a page buffer and a cache. The page buffer is used to access the multiple data lines. The page buffer includes multiple page buffer units, each for each data line, and the page buffer can be configured in multiple page buffer unit arrays. The cache is coupled to the page buffer. The cache includes a cache unit array for each of the X groups of data lines. Each cache unit array is used to access an integer P buffer units of the page buffer and includes P cache units arranged in an integer N cache columns and P / N cache rows. N is not a power of 2 and is a non-integer multiple of the size of the smallest addressable unit. Multiple data path circuits are coupled to the cache to connect N cache cells of a selected P / N cache row of a selected cache cell array from each of the X groups of cache cell arrays to M interface lines, where M is N multiplied by X. The interface is coupled to the M interface lines and includes circuitry for data transmission in D-bit units, where D is an integer multiple of the size of the least addressable unit, and M is a common multiple of D and N. Multiple address translation circuits are provided to generate multiple selection signals for these data path circuits in response to an input address identifying the least addressable unit.
[0013] Multiple address translation circuits may include an address divider to divide the input address by a factor F equal to M and the size of the smallest addressable unit. The address divider outputs a quotient, which is used to select the cache row and the cache array of each of the X groups of cache arrays. The address divider outputs a remainder used by the interface to select D bit units from the M interface lines for transmission on the interface.
[0014] Multiple address translation circuits can be used to generate a series of selection signals to move and transmit storage units through data path circuits and interfaces to support data page transfers of multiple D-bit units of data, which have sequential addresses on the interface.
[0015] This disclosure describes a page buffer and cache configuration for a data path, wherein the page buffer and cache consist of multiple cell arrays aligned with corresponding sets of data lines of a memory array. Each cell array may include N columns and Y columns. The set of data lines coupled to the page buffer cell array comprises N data lines connected to N cells of each Y-row page buffer cell array. The set of data lines connecting the page buffer to the memory array has an orthogonal pitch (the center-to-center spacing between adjacent parallel bit lines), which is a function of memory technology and node manufacturing, and can be very small for high-density memory. To fit the available space, the N cells in the page buffer and cache must conform to the orthogonal pitch of the N bit lines. This is accomplished by stacking the cells in the page buffer and cache in rows. Therefore, the orthogonal pitch of the memory elements in a row of the cell array's page buffer and cache must be equal to or less than the pitch of the N sets of data lines connected to the column in the unit array.
[0016] Using a number N less than D allows for the implementation of cache and page buffers with a smaller area, which will be used for cache and page buffers with D columns.
[0017] The technology described in this disclosure utilizes a "non-binary" cache configuration to support efficient layout and operation of high-density memory. The data path configuration implemented as described in this invention overcomes the access pattern problem of a single word of data stored in two different cache rows. Furthermore, the data path configuration implemented as described in this invention overcomes the problems of cache support and interconnection settings for selecting the same bit in different words of cache and page buffers located in different columns. Moreover, the data path configuration implemented as described in this invention provides a solution to the problem of how to configure a cell array layout that supports memory arrays, particularly for very large page sizes.
[0018] Other aspects and advantages of the invention will become apparent from a review of the accompanying drawings, detailed description, and claims.
[0019] To provide a better understanding of the above and other aspects of the present invention, specific embodiments are described below in conjunction with the accompanying drawings. Attached Figure Description
[0020] Figure 1 It is a block diagram of an integrated circuit having the memory array and non-binary cache configuration described in this invention.
[0021] Figure 2 The layout of a non-binary cache cell array and a page buffer cell array interconnected with data transfer capabilities between the cache cell array and the page buffer cell array is shown.
[0022] Figure 3 The diagram illustrates the interconnection between cache rows with cache cell arrays and local cache drives, enabling data transfer functionality. Figure 2 The layout.
[0023] Figure 4 It is an alternative block diagram for a memory with a non-binary cache configuration.
[0024] Figure 5 It's similar to the configuration of drawing examples. Figure 1 A block diagram of the memory.
[0025] Figure 6 It is similar to drawing another example configuration. Figure 1 A block diagram of the memory.
[0026] Explanation of reference numerals in the attached figures
[0027] 100: Integrated circuit memory device
[0028] 105: Input / Output Interface
[0029] 107, 131, 143, 144, 450, 460, 461, 470, 471, 480, 490: Lines
[0030] 108: Command Decoder
[0031] 110: Control Logic
[0032] 120: Block
[0033] 130: Bus
[0034] 140: Decoder
[0035] 142,452: Address divider
[0036] 145: Character Line
[0037] 160: Memory Array
[0038] 165,200: Data cable
[0039] 171, 415: Page buffer
[0040] 181,420: Cache
[0041] 184: Data Bus
[0042] 195: Input / Output Port
[0043] 205: Page Buffer Unit Array
[0044] 210: Fast Cache Cell Array
[0045] 215, 220: Columns
[0046] 225: Transmission and storage unit
[0047] 231, 232, 331, 332: Vertical interconnection
[0048] 241, 242, 333: Horizontal interconnection
[0049] 410: Array
[0050] 425: Cache selection and support circuit
[0051] 430: Local cache drive
[0052] 432, 434, 436, 438: Group Cache Drivers
[0053] 440: Interface circuit
[0054] 442: Interface Port
[0055] 454: Y Address Generator
[0056] 456: Cache control block
[0057] CA: Row Address
[0058] Q: Quotient
[0059] R: Remainder
[0060] GP(0)-GP(3): Group
[0061] CA_Q1: Row address quotient
[0062] CA_R1: Row address remainder
[0063] cCLK: Fast clock pulse signal
[0064] iCLK: Interface clock pulse signal
[0065] RD[N-1:0], RD[2N-1:N], RD[3N-1:2N], RD[4N-1:3N]: Read data
[0066] WD[N-1:0], WD[2N-1:N], WD[3N-1:2N], WD[4N-1:3N]: Write data
[0067] YAL: Addressing signal for the cache row within the cache cell array.
[0068] YAH: Addressing signal for cached cell array Detailed Implementation
[0069] in accordance with Figure 1-6 A detailed description of the embodiments of this technology is provided.
[0070] Figure 1 This is a simplified chip block diagram of an integrated circuit device. On a single integrated circuit substrate, the integrated circuit device has a memory array 160, a page buffer 171, and a second-level buffer (referred to in this invention as cache 181). The memory device described in this invention can also be implemented using multichip modules, stacked chips, and other configurations.
[0071] Similar to Figure 1 The memory array 160 in the device may include a NAND flash array that uses dielectric charge trapping or floating gate charge trapping memory cells. Furthermore, the memory array 160 may include other non-volatile memory types, including NOR flash memory, ferroelectric RAM, phase change memory, transition metal oxide programmable resistance memory, etc. Additionally, the memory array 160 may include volatile memory, such as dynamic random access memory (DRAM) and static random access memory (SRAM).
[0072] On the integrated circuit memory device 100, control logic 110 with a command decoder 108 includes logic circuitry, such as a state machine and an address counter, which performs memory operations in response to received commands and addresses to access the memory array 160 via a cache 181 and a page buffer 171. These memory operations include read and write operations. For NAND flash memory, write operations may include programming and erase operations. In some embodiments, when moving data pages from the memory array 160, random page read operations can be performed using address data with byte addresses identifying the page, rather than using address data on page boundaries.
[0073] Control logic 110 outputs control signals (not shown) to elements of the memory device and addresses on bus 130. For example, addresses provided on bus 130 may include the output of an address counter (e.g., a sequential address) or addresses carried in received commands.
[0074] Decoder 140 includes a row decoder coupled to multiple word lines 145 and arranged along the columns of memory array 160, and a column decoder coupled to page buffer 171 and cache 181. Additionally, in this example configuration, address divider 142 is line-coupled to control logic 110 to receive the row address CA on line 131, which includes the row address portion of the address. Page buffer 171 is coupled to multiple sets of data lines 165 arranged along the rows of memory array 160 to read data from and write data to the corresponding groups of memory cells in memory array 160. Figure 1The diagram illustrates four groups of data lines 165, which are coupled to various groups of memory cells in the memory array 160. The number of groups is a function of the data path configuration, as discussed in more detail below. In the layout on the integrated circuit, multiple groups of data lines are physically grouped to connect to rows of page buffer cells in the page buffer 171. These multiple groups of data lines combine multiple data lines usable in parallel, which can be equal to the page width plus additional data. For example, a page may include 16K bits + 2K additional bits. Each group of data lines may include one-quarter of the number of page data lines.
[0075] The memory array may include bit lines and word lines. As used in this invention, a data line is a data path that connects a bit line to the page buffer 171. In some configurations, there is one data line per bit line. In some configurations, two bit lines may be selectively connected and share a single data line.
[0076] Page buffer 171 may include one or more storage elements for each data line, which serve as page buffer units. Decoder 140 may select specific storage units of memory array 160 and couple those specific storage units to page buffer 171 via corresponding data lines. Figure 1 In this configuration, memory array 160 includes four sets of memory cells. During page operations, data lines 165 are configured as four sets of data lines corresponding to the four sets of memory cells in the array, and these data lines can be used in parallel to transfer pages to page buffer 171. Page buffer 171 can store data written to or read from these specific memory cells in parallel. Page buffer 171 may have a page width of several kilobytes (e.g., 16KB or 32KB) or more additional bits (e.g., 2KB) that can be used for page-related metadata (e.g., associated ECC codes).
[0077] Cache 181 is coupled to the page buffer and may include at least one storage element for each data line, the at least one storage element serving as a cache unit. Data pages can be transferred from the page buffer 171 to cache 181 in a shorter time compared to the time required to transfer data pages from the array's storage units to the page buffer 171.
[0078] The page buffer units in page buffer 171 and the cache units in cache 181 may include clocked storage elements, such as flip-flops and latches using a multiple row by multiple column architecture. Cache 181 can operate in response to a cache clock (cCLK) signal.
[0079] Other embodiments may include a three-level buffer structure, which includes a page buffer 171 and two additional buffer levels. Furthermore, other configurations of the buffer memory structure may be utilized, utilizing the data path circuitry between the page buffer and the interface.
[0080] Data interfaces such as input / output (I / O) interface 105 are connected to page buffer 171 and cache 181 via data bus 184, which includes M interface lines. Data bus 184 may have a bus width smaller than a page, substantially smaller than a page width. In the examples described below, the interface lines in data bus 184 may include N interface lines in each group of cache 181, where N is a non-integer multiple of the minimum addressable unit of I / O interface 105, and data bus 184 is selected according to the data path configuration described below.
[0081] I / O interface 105 may include a buffer for parallel latching of data from all sets of bus lines, the buffer latching N times the number of bit sets (e.g., Nx4) in each cache clock cycle.
[0082] Input / output data, address, and control signals move between the data interface, command decoder 108, and control logic 110, and input / output ports (I / O ports) 195 on the integrated circuit memory device 100 or other data sources inside or outside the integrated circuit memory device 100.
[0083] exist Figure 1In the example shown, the control logic 110 of the bias arrangement state machine controls the application of the bias arrangement supply voltage, which is generated or provided by one or more voltage supplies in block 120, such as read, program, and erase voltages, including page read functions that transfer data from pages in the memory array to page buffers.
[0084] Control logic 110 and command decoder 108 constitute a controller, which can be implemented using special-purpose logic circuitry having a state machine and supporting logic. In an alternative embodiment, the control logic includes a general-purpose processor, which can be implemented on the same integrated circuit that executes a computer program to control device operation. In still other embodiments, a combination of special-purpose logic circuitry and a general-purpose processor can be used to implement the control logic.
[0085] I / O interface 105 or other types of data interfaces, and array addresses on bus 130, can be operated using "binary" addressing, such that boundary addresses defined by powers of 2 address each memory cell in the memory array and each memory cell transferred on I / O interface 105, including 8, 16, 32, etc., corresponding to a specific order of binary numbers as addresses. In the technique described in this invention, page buffer 171 and cache 181 are configured with "non-binary" transfer memory cells, such that some memory cells in the cache and page buffer have boundaries not defined by powers of 2. For example, embodiments of this invention include page buffers and caches, wherein transfer memory cells (one transfer memory cell per cache column in the following example) have N bits, which are non-integer multiples of the smallest addressable memory cell size. In one example, for a minimum addressable memory cell of one byte, the cache is 12 bits, or 1.5 times the size of the minimum addressable memory cell, and the I / O interface 105 is used for external transmissions of an integer multiple of the size of the minimum addressable memory cell for each cycle of the interface clock pulse signal iCLK. For a minimum addressable memory cell of one byte, the interface may transmit 8 bits per clock pulse cycle or 16 bits per clock pulse cycle, as described above.
[0086] like Figure 1As shown, address divider 142 (“divide by F”) receives an address bit from line 131 that specifies a column address CA, which identifies the first bit at the boundary of the minimum addressable unit to be accessed. Address divider 142 divides the column address CA by a factor F equal to the quantity M (row size N multiplied by the number of groups X) divided by the size of the minimum addressable unit (the number of N-bit cache transfer units X comprises F minimum addressable units). For at least some column addresses, division by F produces a quotient and a remainder. Address divider 142 provides a quotient Q on line 143, which is applied to control logic 110 to support address generation for “non-binary” addressing of the cache and page buffer. Address divider 142 provides a remainder R on line 144, which is applied to I / O interface 105 to support address generation for "binary" addressing of I / O interface 105. In one approach, the external address bit of row address CA is converted by address divider 142 into an internally addressable quotient Q and remainder R. The quotient Q is sent to the controller's Y address counter for the start counter address, while the remainder R is sent to I / O interface 105 for the start address. The controller's Y address counter and the I / O interface 105's address counter increment in response to memory operations to address the required data. The counter address is then sent to the Y decoder of decoder 140 to select a row in the cell array within the cache array and increments the counter address in response to memory operations to address the required data.
[0087] I / O interface 105 may include buffers, shift register buffers, or other supporting circuitry with a transceiver to transmit and receive data on the port using an interface clock pulse signal (iCLK). For example, integrated circuit memory device 100 may include input / output ports using 8 or 16 pins for receiving and transmitting bus signals. Another pin may be connected to a clock line carrying the interface clock pulse signal iCLK. Yet another pin may be connected to a control line carrying a chip-enable or chip-select signal CS#. In some embodiments, I / O port 195 may be connected to on-chip host circuitry, such as a general-purpose processor or application-specific circuitry, or a combination of modules providing system-on-a-chip functionality supported by memory array 160.
[0088] The interface clock pulse signal iCLK can be provided by one of the I / O ports 195 or generated internally. The cache clock pulse signal cCLK can be generated in response to the interface clock pulse signal iCLK or generated independently, and the cache clock pulse signal cCLK can have different clock pulse rates. Data can be transferred between the I / O interface 105 and the cache using the cache clock pulse signal cCLK, and data can be transferred through the I / O port 195 of the I / O interface 105 using the interface clock pulse signal iCLK.
[0089] In one embodiment, I / O interface 105 is a byte-wide interface that includes a set of eight I / O ports 195 for data transmission, and is configured to operate in double data rate mode, wherein two bytes are output one byte at a time on the rising and falling edges of an interface clock cycle. In other embodiments, I / O interface 105 is a byte-wide interface with a set of eight I / O ports 195, which may operate in double data rate mode (two bytes per clock cycle) or single data rate mode (one byte per clock cycle) depending on the configuration settings carried by the device or commands used for the interface.
[0090] Other types of data interfaces with parallel and serial interfaces can also be used. The serial interface may be based on or conform to the Serial Peripheral Interface (SPI) bus specification, where the command channel shares I / O pins used for address and data. I / O ports 195 on a specific integrated circuit memory device 100 can be used to provide output data with an I / O data width that is an integer multiple of the minimum addressable memory cell size. For a minimum addressable cell of one byte, the minimum addressable memory cell can be 8, 16, 32, or more bits in parallel per interface clock pulse cycle.
[0091] In some implementations, the data interface may be connected to binary addressable memory, such as SRAM or DRAM, on the same chip or the same multichip module using a bus or other type of interconnect.
[0092] Page buffer 171 and cache 181 consist of multiple cell arrays aligned with corresponding sets of data lines of memory array 160. Each cell array may include N columns and Y rows. The set of data lines comprises N data lines connected to N cells in each of the Y columns of the page buffer. The set of data lines connecting the page buffer to the memory array has an orthogonal pitch (the center-to-center spacing between adjacent parallel bit lines), which is a function of memory technology and node manufacturing; for high-density memory, the set of data lines can be very small. To fit the available space, the N cells in the page buffer and cache must conform to the orthogonal pitch of the N bit lines. This is accomplished by stacking the cells in the page buffer and cache in rows. Therefore, the orthogonal pitch of the memory elements in a row of cell arrays in page buffer 171 and cache 181 must be equal to or less than the pitch of the set of N bit lines connected by the column in the unit array.
[0093] Figure 2 The illustration depicts the layout of a page buffer unit array 205 and a cache unit array 210 in an example configuration according to the technology described in this invention. In this example, assuming the smallest addressable unit is one byte, the interface outputs at most two bytes per interface clock cycle in double data rate mode.
[0094] In this example, page buffer cell array 205 has N = 12 columns × Y = 16 rows of cells coupled to 192 data lines 200, where the cells in each row are labeled b0 to b11. The number N is a non-integer multiple of 8. Therefore, each row of page buffer cell array 205 is configured and conforms to the orthogonal spacing of 12 data lines. Similarly, cache cell array 210 has 12 columns × 16 rows of cells, where the cells in each row are again labeled b0 to b11. Figure 2 The cell array includes columns 215 of Y = 16 cache selection and control circuits, corresponding to cache column numbers CC0-CC15 for each of the 16 cache rows. Furthermore, the cell array includes columns 220 of N = 12 cache drivers, where one of the bits b0-b11 of each selected row in the cache cell array 210 provides a data path for the N = 12-bit transfer memory unit 225.
[0095] The cell array is configured with N columns, for example, 12 columns based on the addressable system's bytes, where N is a non-integer multiple of the smallest addressable cell size, rather than an integer multiple of the binary number of rows, such as sixteen columns. This saves area on the integrated circuit. The reduction in the number of columns (16-12=4) translates into an area saving equal to the width of the cell array multiplied by the vertical pitch of the eliminated rows (4 eliminated rows in this example). In this example, reducing from 16 columns to 12 columns, the area of the page buffer and cache is reduced by 25%. The flexibility provided by the implementation where N is a non-integer multiple of the smallest addressable memory cell size allows for a smaller layout.
[0096] Therefore, in this technique, the cache cell array outputs N-bit transfer memory cells, where N is a non-integer multiple of the smallest addressable cell size. Thus, row selection of a cell array with N-bit transfer cells (e.g., 12) is a non-binary addressing operation. Furthermore, as... Figure 2The 12-column cache shown in the diagram, with this layout, means that two bytes of characters addressable from the memory array cannot be stored in a single row of the cell array. As described in more detail below, each cache group is used to transmit one cache row of data (a 12-bit cache transfer unit) per cycle of the cache clock pulse signal cCLK. For the example using four groups, the configuration causes each group to transmit one cache row, or four cache rows and 48 bits within a single cache clock pulse cycle. These 48 bits consist of three 16-bit words used for transmission between the cache and the interface. However, the data from these three 16-bit words is distributed to memory cells in four different groups of the memory array through this scatter / gather data path configuration of four 12-bit portions.
[0097] like Figure 2 As shown, an array of interconnections is formed over a patterned conductor layer (e.g., a metal layer) covering the page buffer cell array 205 and the cache cell array 210. In parallel operation, the interconnections are used to transfer data from the page buffer cell array 205 to the cache cell array 210. Figure 2 Representative interconnects shown include vertical interconnection 231 on cache row CC2, vertical interconnection 232 on cache row CC8, horizontal interconnection 241 on page buffer row b2, and horizontal interconnection 242 on cache row b2. In the operation of transferring data between page buffer cell array 205 and cache cell array 210, control signals energize horizontal interconnections 241 and 242 to transfer data in selected columns of cells via the vertical interconnection between the page buffer and cache for read operations and write operations between the page buffer and cache. This operation can be performed 12 times (once per column) to transfer the entire contents of page buffer cell array 205 to cache cell array 210, or vice versa.
[0098] Figure 3 Illustration as follows Figure 2 The layout of the page buffer cell array 205 and cache cell array 210 is shown, wherein the same elements have the same element number. For example, in Figure 3In this configuration, the overlapping interconnects formed on the patterned conductor layer include vertical interconnects 331 and 332, and horizontal interconnects 333 corresponding to cell b4. During data transmission between the cache cell array 210 and the cache drivers of column 220, a control signal energizes the vertical interconnect 331 to connect cells in selected column b4 to the horizontal interconnect 333. The horizontal interconnect 333 is connected to the corresponding cache driver of b4 via the vertical interconnect 332. The 12 horizontal interconnects 333 (one per column in the cache) can operate in parallel, thereby converting selected 12 bits from the 16 rows to the corresponding cache driver for communication with the interface circuitry.
[0099] The present invention provides a data path in which binary-addressable memory cells are distributed among transfer memory cells in multiple groups of cells of a memory array, the transfer memory cells having a size that is a non-integer multiple of the size of the binary-addressable memory cell. In this manner, adjacent columns of a cell array in the page buffer and cache store data bits of non-adjacent data in the external addressing mode. According to the data path of the present invention, memory cells (e.g., bytes or words) of the array addressed by the binary address of the external addressing mode are distributed to cells in different groups of the memory array via data lines in different groups, and these memory cells are not aligned with the transfer memory cells.
[0100] Figure 4 The diagram illustrates additional system features with data path configurations as described in this invention. Figure 4 In this array, array 410 is configured as multiple groups GP(0)-GP(3). The data path includes page buffer 415 and cache 420, which are configured as multiple unit arrays (indicated by vertical lines). The respective set of unit arrays of page buffer 415 and cache 420 serve each group GP(0)-GP(3) in the array.
[0101] Page buffer 415 includes X arrays of page buffer cells, each corresponding to X sets of data lines. Each page buffer cell array is used to access an integer P data lines of the multiple data lines and includes at least P page buffer cells arranged in an integer number of N columns and Y (Y = P / N) rows. Furthermore, page buffer 415 includes circuitry for data transfer between the P page buffer cells and the P cache cells of the corresponding cache array.
[0102] The cache 420 is coupled to the page buffer and includes X cache cell arrays (one cache cell array for each of the X data lines), which correspond to X page buffer cell arrays. Each cache cell array accesses an integer P cache cells of the page buffer and includes P cache cells arranged into an integer N cache columns and Y (Y = P / N) cache rows, where N is not a power of 2. Each cache cell array includes a local cache driver for accessing N cache cells of a selected cache row from the Y cache rows of the cell array.
[0103] The data path circuitry is coupled to the cache to connect N cache cells of each of the X cache cell arrays and the selected cache row of the P / N cache row of the selected cache cell array to M interface lines, where M is N multiplied by X. The M interface lines are coupled to an interface that includes circuitry for transmitting data in integer D bit units, where D is a power of 2 and M is a common multiple of D and N.
[0104] Figure 4The data path circuitry shown includes a cache select and support circuitry 425 and a local cache driver 430 for controlling data transfer between the page buffer 415 and the cache 420 within each cell array, and for transferring data from the cache 420 via the local cache driver 430. Furthermore, the data path circuitry includes group cache drivers 432, 434, 436, and 438 for the respective groups GP(0)-GP(3) of the array. In this example, the local cache driver 430 is coupled to the group cache drivers 432, 434, 436, and 438 for the respective groups GP(0)-GP(3) of the array via an N-bit read data path from the cache and an N-bit write data path transferred to the cache. Group cache drivers 432, 434, 436, and 438 are connected to the N-bit read data path for transmission to interface circuitry 440 during read operations, and are also connected to the N-bit write data path for transmission from interface port 442 to the cache during write operations. Interface circuitry 440 includes multiple circuits for transmitting read data RD[N-1:0] from group cache driver 432, read data RD[2N-1:N] from group cache driver 434, read data RD[3N-1:2N] from group cache driver 436, and read data RD[4N-1:3N] from group cache driver 438. Interface circuitry 440 also receives the interface clock pulse signal iCLK on line 490. The multiplexer circuit within interface circuit 440 connects interface port 442 to corresponding portions of the consolidated read data path (read data RD[4N-1:0]). Furthermore, interface circuit 440 includes multiple circuits for transmitting write data WD[N-1:0] to group cache driver 432, for transmitting write data WD[2N-1:N] to group cache driver 434, for transmitting write data WD[3N-1:2N] to group cache driver 436, and for transmitting write data WD[4N-1:3N] to group cache driver 438. The multiplexer circuit within interface circuit 440 connects interface port 442 to corresponding portions of the consolidated read data path and the consolidated write data path (write data WR[4N-1:0]) of read data RD[4N-1:0].
[0105] In the cell array, the location of the cache selection and support circuitry 425 does not necessarily have to be between the cache 420 and the local cache driver 430. The circuitry can be distributed across the cache and page buffer arrays to suit a particular implementation.
[0106] In a read or write operation, a starting address is provided to identify the first smallest addressable unit the operation wishes to access. A sequence of read or write operations across all the smallest addressable units in the entire page is used to support page reads or writes.
[0107] To support non-binary addressing in page buffers and cache arrays, addressable circuitry is provided. For example... Figure 4 As shown, in one example, the row address CA is provided to line 450 to support memory operations using an external binary addressable system. The row address CA on line 450 is applied to address divider 452. As described above, the factor F, which is the divisor of address divider 452, is obtained by dividing the width M = 4N of the selected cell array output in each of the four groups by the size of the smallest addressable cell. In the example described in this invention, M = 48 and the size of the smallest addressable cell is eight bits. The factor F in this example is 48 / 8 = 6. The factor F is not a power of 2, and in this example, the divisor is a non-integer multiple of the size of the smallest addressable cell.
[0108] Address divider 452 outputs column address quotient CA_Q1 to line 460 and column address remainder CA_R1 to line 461. The column address quotient CA_Q1 is applied to Y address generator 454, which outputs a Y address signal on line 471 to cache selection and support circuitry 425 and local cache driver 430. Furthermore, Y address generator 454 exchanges control signals with cache control block 456. A cache clock pulse signal cCLK can be generated by dividing the interface clock pulse signal iCLK at the interface block or by utilizing other clock pulse circuitry or sources, and the cache clock pulse signal cCLK on line 470 can be provided to cache control block 456. The cache clock pulse signal cCLK and other control signals can be provided on line 480 to coordinate the operation of the local cache driver 430 and the group cache drivers 432, 434, 436, 438 for group GP(0)-GP(3).
[0109] In each cache clock pulse cycle, four groups of outputs are used to select a row of a selected cell array. The four groups of outputs form a cache-addressable cell of M bits, where M is a common multiple of the largest memory cell size DMAX output over a single cycle of the interface clock pulse signal iCLK (DMAX is 16 for DDR with 8 port outputs), and the quantity N is the number of bits per unit array output per group (N is 12 in these examples). In this example, the common multiple of 16 and 12 is 48, and M equals 48. Because M / N equals 4, there are four data lines. Because M / 16 equals 3, there are three 16-bit words in the cache-addressable cell.
[0110] To transfer a page, the Y address generator 454 generates a signal to select the starting cell array in each group of cell arrays, and transfers one cache line from each of the four starting cell arrays via the interface. Then, iteratively selecting all cell arrays in the four groups according to the M-bit sequence of cache-addressable units (CA) to output the entire page. In a configuration supporting random start addresses for in-page reads or writes, one of the four starting cell arrays stores the first bit of the smallest addressable unit identified by the row address CA. In each cache clock cycle, a set of M bits is transferred to the interface circuitry 440. The interface uses the remainder signal CA_R1 to generate a control signal for the multiplexer to select the smallest addressable unit identified by the external address from the M bits of the data path, and provides the selected smallest addressable unit to the interface for transfer at interface port 442. As described above, in the double data rate implementation, two smallest addressable units can be transferred within an interface clock cycle.
[0111] The data path requires circuitry to convert binary addressing for external access to the smallest addressable unit into “non-binary” addressing of the page buffer and cache to access cache rows that hold a non-integer multiple of the smallest addressable unit in bits. Similarly, the data path requires circuitry to convert non-binary addressing of the page buffer and cache into binary addressing of the input / output interface. One method of controlling this addressing mode involves an address divider, where the binary address of the interface and memory array is divided by a factor F, which is a function of the cache addressable unit size M and the smallest addressable unit size of the binary addressing mode for the array (e.g., four transfer memory units per N bits). As mentioned above, if the least significant bit of the binary addressing mode corresponds to a byte (8 bits), and the least significant bit is the smallest memory unit externally addressed via the I / O interface, then the factor F is M / 8, which is 48 / 8 = 6 for the example. If the least significant bit of the binary addressing mode corresponds to a word (16 bits), the factor F based on this formula will be M / 16 (48 / 16 = 3).
[0112] The examples above use a cache row size of 12 and a group number equal to 4. These are advantageous examples but not limiting ones. Other sizes can be used as long as the cache row size is a non-integer multiple of the minimum addressable unit size, or a non-integer multiple of the maximum data size output within one interface clock pulse cycle.
[0113] Figure 5 The diagram illustrates a read data path configuration with a page size of 16K bytes of page data lines (2K additional data lines), an 8-bit minimum addressable unit size, a 12-bit cache row size, and a 12x16 cell array size with 16 rows. The write data path is similar and is omitted for clarity.
[0114] Page size (PS): 16KB + 2KB;
[0115] RD bus size: 48 bits = 6 bytes = 3 words;
[0116] Number of groups (NOG): 4;
[0117] Cell array size (UAS): 16 columns x 12 bits;
[0118] Number of cell arrays per group (NOUA);
[0119] PS≤UAS x NOUA x NOG;
[0120] NOUA≥PS / (UAS x NOG)=(18x1024x8) / (16x12x4)=192.
[0121] Therefore, in this configuration, each group has 192 cell arrays, and the page buffer / cache has a 768-cell array of auxiliary memory arrays. Each group of data lines comprises 192x12x16 data lines. To transfer a page, the four groups of cell arrays operate in parallel and access the 192 cell arrays of each group sequentially. In the sequence, a selected cell array (responding to the address signal YAH of the cache cell array) can output a 12-bit selected cache row (responding to the address signal YAL of the cache row within the cache cell array) in each cache clock pulse cycle. This selected cache row, together with the four group outputs, forms an M-bit cache-addressable cell. This will execute 16 address signal YAL cycles to output all rows of the selected cache cell array, and then execute the next address signal YAH cycle, sequentially moving to the next cache cell array until all 192 cell arrays of all groups are accessed.
[0122] Other configurations are available. A generalization supporting the technical configurations described in this invention is as follows.
[0123] A 2i x N cell array is proposed (a 16x12 cell array was shown above for the cases of i=4 and N=12). N can be any non-integer multiple of the smallest addressable cell size.
[0124] Within a CACHE_CLK, the selected cell array for each group reads / writes only N bits of metadata.
[0125] In a CACHE_CLK, M bits of metadata are read / written from the cache, where M is a common multiple of N and DMAX, and DMAX is the maximum number of bits transferred in one interface clock pulse cycle (DMAX = 16 for DDR in the example above). M is not a power of 2 (M = 48 for the example above).
[0126] For each M / 2i output of an N-bit cache line, the YADD counter (YA) is incremented by 1 to select the next line. An address divider is added to convert the external binary address to the page buffer and the YADD system of the cache, as well as the YADD system of the interface circuitry.
[0127] In an N-bit cache column unit write operation, sequentially addressed bytes or other minimum addressable units of the externally addressable system are distributed among the array groups, such that the M bits written to the array group and transmitted per cache clock cycle consist of an integer number of sequential minimum addressable units of the interface. Similarly, in an N-bit cache column unit read operation, sequentially addressable minimum addressable units of the externally addressable system are collected from the array group, such that the M bits read from the array group and transmitted per cache clock cycle consist of an integer number of sequential minimum addressable units of the interface.
[0128] Figure 6 The diagram illustrates another example configuration with a page size of 16K bytes of page data lines with 2K additional data lines, an 8-bit minimum addressable unit size, a 12-bit cache row size, and a 12x12 cell array size with 12 rows. Here, the number of cache rows equals the number of columns, and therefore the number of cache rows is also a non-integer multiple of the minimum addressable unit size.
[0129] Page size (PS): 16KB + 2KB;
[0130] RD bus size: 48 bits = 6 bytes = 3 words;
[0131] Number of groups (NOG): 4;
[0132] Cell array size (UAS): 12 rows x 12 bits
[0133] Number of cell arrays per group (NOUA);
[0134] PS≤UAS x NOUA x NOG;
[0135] NOUA≥PS / (UAS x NOG)=(18x1024x8) / (12x12x4)=256.
[0136] Therefore, in this configuration, each group has 256 cell arrays, and the page buffer / cache has 1024 cell arrays for the assist memory array. To transfer a page, the four groups of cell arrays operate in parallel and access the 256 cell arrays sequentially. In the sequence, the selected cell array can output a 12-bit selected line each cache clock cycle, which is combined with the outputs of the four groups. This will execute for 12 cycles, and then the next cell array will be accessed sequentially until all 256 cell arrays of all groups have been accessed. In this example, since the number of cache lines per cell array is not a power of 2, the cell arrays are not configured according to the probability (2i x N).
[0137] At Figure 6 In the NxN configuration, compared to a configuration aligned to a power of 2 (e.g., 16) boundary, the Y address generator generates addressing signals YAL and YAH, which resolves the complexity of determining when to jump to the next cell array on a 12-cycle boundary.
[0138] The configuration can vary based on the parameters described in this invention, including page size (16K, 32K, etc.), cache row size N, minimum addressable unit size (bytes, words), and maximum data transfer per interface clock cycle (DMAX, for example, two bytes in DDR mode). The cache row size N can be selected based on the physical layout of a particular implementation. The cache row size N used in the above example is equal to 12. The cache row size N can be any other non-integer multiple of the minimum addressable unit size.
[0139] For example, N can be an odd number like 9, although this might result in the remaining number of data lines for the page size not being divisible by 9 bits. In this case, assuming DMAX is 8, then M is a common multiple of 9 and 8, which could be 72. In this case, there would be 72 / 9 = 8 unit arrays.
[0140] While the invention has been disclosed with reference to the above detailed preferred embodiments and examples, it should be understood that these examples are intended to be illustrative rather than limiting. Modifications and combinations will readily occur to those skilled in the art, and such modifications and combinations will be within the spirit of the invention and the scope of the appended claims.
[0141] In summary, although the present invention has been disclosed above with reference to embodiments, it is not intended to limit the invention. Those skilled in the art to which this invention pertains can make various modifications and refinements without departing from the spirit and scope of the invention. Therefore, the scope of protection of this invention should be determined by the appended claims.
Claims
1. A memory that uses an addressing mode for addressing, having a minimum addressable unit, the minimum addressable unit having a size, characterized in that, include: A memory array; One data interface; as well as A data path for transferring data between the memory array and the data interface using multiple transfer storage units with N bits, where N is a non-integer multiple of the size of the smallest addressable unit; The data path includes a cache configured to connect to the data interface. ; The cache consists of X cache unit arrays, each array comprising P cache units arranged in N cache rows and Y cache columns, where P equals Y multiplied by N. ; The data of multiple smallest addressable units is distributed across X groups of cache units in the data transfer operation of N-bit cache units. In one cache clock cycle, a selected column of a selected cache unit array from each of the X groups of cache unit arrays is transferred, and the transferred data of the X groups of cache unit arrays forms an M-bit cache addressable unit, where M is N multiplied by X, and N, X, P, Y, and M are integers greater than 1.
2. The memory according to claim 1, characterized in that, This data path also includes: A page buffer, which is used to connect to the memory array; The memory array includes multiple data lines that are used to connect to the page buffer, which is located in the X group of data lines.
3. The memory of claim 2, wherein the data path includes a plurality of circuits coupled to the cache for connecting N cache cells of a selected Y cache row of a selected cache cell array of each X group of cache cell arrays to M interface lines; and The data interface is coupled to the M interface lines and includes multiple circuits that transmit data in D-bit units, where D is an integer multiple of the size of the smallest addressable unit and M is a common multiple of D and N.
4. The memory according to claim 3, comprising: Multiple address translation circuits generate multiple selection signals for the data path in response to an input address.
5. The memory according to claim 1, wherein, The memory array includes multiple data lines for connection to a page buffer, which is configured on X sets of data lines; as well as The data page based on this addressing method includes multiple minimum addressable cells with multiple sequential addresses, and the data page stored in the memory array is distributed among multiple N-bit memory cells, which are coupled to data lines that are set non-sequentially across X groups.
6. The memory according to claim 1, wherein N is not a power of 2.
7. The memory according to claim 1, wherein: The memory array includes multiple data lines for accessing multiple memory cells of the memory array. These data lines include X groups of data lines, where X is an integer greater than 1. A page buffer for accessing these data lines, the page buffer comprising multiple page buffer units, each page buffer unit for each data line; The cache is coupled to the page buffer. The cache includes multiple cache cell arrays, each of the multiple cache cell arrays is used for each of the X data lines, and each cache cell array is used to access an integer P buffer cells of the page buffer. Multiple data path circuits are coupled to the cache to connect N cache cells of a selected Y cache row of a selected cache cell array of each X group of cache cell arrays to M interface lines; The data interface is coupled to the M interface lines. The data interface includes circuitry for data transmission in D-bit units, where D is an integer multiple of the size of the smallest addressable unit, and M is a common multiple of D and N. as well as Multiple address translation circuits generate multiple selection signals for these data path circuits in response to an input address.
8. The memory of claim 7, wherein the address translation circuitry includes an address divider for dividing the input address by a factor equal to the number M divided by the size of the smallest addressable unit to generate the selection signals.
9. The memory of claim 8, wherein the address divider generates a quotient and a remainder, and the address translation circuitry includes a cache address generator for selecting a plurality of cache rows and a plurality of cache cell arrays, and the interface includes circuitry for generating the selection signal to select a D-bit cell connection of the M interface lines for transmission.
10. The memory of claim 7, wherein the page buffer comprises an X-group page buffer cell array for corresponding data lines of the X-group data lines, each page buffer cell array being used to access an integer P data lines of the data lines and comprising P page buffer cells arranged in an integer number of N columns and Y rows, and each page buffer cell array including circuitry for data transfer between the P page buffer cells and P cache cells of a corresponding cache cell array of the cache.
11. The memory of claim 7, wherein the data lines have an orthogonal pitch, and each cache row of the cache cell array of the cache conforms to the orthogonal pitch of the N data lines of the data lines.
12. The memory of claim 7, wherein the address translation circuits generate a series of selection signals for the data path circuits and the interface to transmit a plurality of D-bit cells having a plurality of sequential addresses on the interface.
13. The memory of claim 7, wherein the data lines are coupled to a memory cell of a page of the memory array, and the address translation circuitry generates a series of selection signals for the data path circuitry and the interface to transmit a data page of multiple D-bit cells having multiple sequential addresses on the interface.
14. The memory of claim 7, wherein the data path circuitry includes X group drivers coupled to the cache, each group driver being coupled to one of the cache cell arrays.
15. The memory of claim 7, wherein N is not a power of 2.
16. The memory of claim 7, wherein D is 16 and N is 12.
17. The memory of claim 7, wherein D is twice the size of the smallest addressable unit.
18. The memory of claim 7, wherein D is equal to the size of the smallest addressable unit.
Citation Information
Patent Citations
Fast page continuous read
US20200125443A1